Agent Skills: ai-post-training

Post-training and alignment: reward modeling, RLHF/PPO, DPO/DAAs, GRPO, RLVR, RLAIF, over-optimization. Use when adapting an SFT model with preference or verifiable-reward signals.

UncategorizedID: vasilyu1983/ai-agents-public/ai-post-training

Install this agent skill to your local

pnpm dlx add-skill https://github.com/vasilyu1983/ai-agents-public/ai-post-training

Skill Files

Browse the full folder contents for ai-post-training.

Download Skill

Loading file tree…

Select a file to preview its contents.