Agent Skills: unsloth-grpo
Implementation of Group Relative Policy Optimization (GRPO) for training reasoning models, optimized for 8x memory savings (triggers: GRPO, reasoning, DeepSeek-R1, reinforcement learning, RLVR, GRPOTrainer, thinking tokens).
UncategorizedID: cuba6112/skillfactory/unsloth-grpo
Install this agent skill to your local
Skill Files
Browse the full folder contents for unsloth-grpo.
Loading file tree…
Select a file to preview its contents.