Agent Skills: grpo-rlvr-training
Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.
UncategorizedID: wshobson/agents/grpo-rlvr-training
39,5404,213
Install this agent skill to your local
Skill Files
Browse the full folder contents for grpo-rlvr-training.
Loading file tree…
Select a file to preview its contents.