Agent skill · wshobson
grpo-rlvr-training
Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically
Reference it in any coding agent with:
@skills wshobson/grpo-rlvr-training