Agent Skills: ai-post-training
Post-training and alignment: reward modeling, RLHF/PPO, DPO/DAAs, GRPO, RLVR, RLAIF, over-optimization. Use when adapting an SFT model with preference or verifiable-reward signals.
UncategorizedID: vasilyu1983/ai-agents-public/ai-post-training
8719
Install this agent skill to your local
Skill Files
Browse the full folder contents for ai-post-training.
Loading file tree…
Select a file to preview its contents.