Agent skill · wshobson

grpo-rlvr-training

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically

Reference it in any coding agent with:

@skills wshobson/grpo-rlvr-training

Browse the @skills marketplace