Agent Skills: unsloth-orpo

One-step preference alignment using Odds Ratio Preference Optimization (ORPO) (triggers: ORPO, preference optimization, alignment, ORPOTrainer, log_odds_ratio, binary preference).

UncategorizedID: cuba6112/skillfactory/unsloth-orpo

Install this agent skill to your local

pnpm dlx add-skill https://github.com/cuba6112/skillfactory/unsloth-orpo

Skill Files

Browse the full folder contents for unsloth-orpo.

Download Skill

Loading file tree…

Select a file to preview its contents.