Agent Skills: grpo-rlvr-training

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.

UncategorizedID: wshobson/agents/grpo-rlvr-training

Repository

wshobsonLicense: MIT
39,5404,213

Install this agent skill to your local

pnpm dlx add-skill https://github.com/wshobson/agents/grpo-rlvr-training

Skill Files

Browse the full folder contents for grpo-rlvr-training.

Download Skill

Loading file tree…

Select a file to preview its contents.