Agent Skills: unsloth-dpo

Direct Preference Optimization (DPO) for aligning models with preference data without separate reward models. Triggers: dpo, preference optimization, rlhf, ref_model=none, patchdpotrainer, dpotrainer.

UncategorizedID: cuba6112/skillfactory/unsloth-dpo

Install this agent skill to your local

pnpm dlx add-skill https://github.com/cuba6112/skillfactory/unsloth-dpo

Skill Files

Browse the full folder contents for unsloth-dpo.

Download Skill

Loading file tree…

Select a file to preview its contents.