Agent Skills: preference-optimization

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

UncategorizedID: wshobson/agents/preference-optimization

Repository

wshobsonLicense: MIT
39,5404,213

Install this agent skill to your local

pnpm dlx add-skill https://github.com/wshobson/agents/preference-optimization

Skill Files

Browse the full folder contents for preference-optimization.

Download Skill

Loading file tree…

Select a file to preview its contents.