Agent Skills: unsloth-grpo

Implementation of Group Relative Policy Optimization (GRPO) for training reasoning models, optimized for 8x memory savings (triggers: GRPO, reasoning, DeepSeek-R1, reinforcement learning, RLVR, GRPOTrainer, thinking tokens).

UncategorizedID: cuba6112/skillfactory/unsloth-grpo

Install this agent skill to your local

pnpm dlx add-skill https://github.com/cuba6112/skillfactory/unsloth-grpo

Skill Files

Browse the full folder contents for unsloth-grpo.

Download Skill

Loading file tree…

Select a file to preview its contents.