Agent Skills: preference-optimization
Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.
UncategorizedID: wshobson/agents/preference-optimization
39,5404,213
Install this agent skill to your local
Skill Files
Browse the full folder contents for preference-optimization.
Loading file tree…
Select a file to preview its contents.