Agent Skills: unsloth-dpo
Direct Preference Optimization (DPO) for aligning models with preference data without separate reward models. Triggers: dpo, preference optimization, rlhf, ref_model=none, patchdpotrainer, dpotrainer.
UncategorizedID: cuba6112/skillfactory/unsloth-dpo
Install this agent skill to your local
Skill Files
Browse the full folder contents for unsloth-dpo.
Loading file tree…
Select a file to preview its contents.