grpo skills

34 tagged grpo, measured the same way as everything else here.

Browse within: Post-Training 27Distributed Training 19Multi-Agent 14Multimodal 14agentic-rl 14DPO 13reinforcement-learning 12RLHF 11PPO 10HuggingFace 7ai-research 7fine-tuning 7DeepSpeed 6TRL 6

debug-hang

02

redai-infra/Relax

Skill Claude CodeCodex

A troubleshooting workflow for Ray, a system that runs machine-learning jobs across multiple computers or GPUs, when a distributed training job stops making progress.

580 4d ago A 67 tokens original Apache-2.0

redai-infra/Relax

Skill Claude CodeCodex

Integrate a new NVIDIA NeMo Gym environment into Relax as a three-step recipe. Use when adding or debugging a recipe under examples/nemogymagentic/recipes; covers data preparation, a local private Gym service, direct Ray training launch, verifier validation, callback networking, lifecycle cleanup, and failure triage.

580 4d ago A 73 tokens original Apache-2.0

verl-to-relax

04

redai-infra/Relax

Skill Claude CodeCodex

Migrate RL training recipes from verl to Relax framework. Use when user wants to port reward functions, tool environments, training scripts, or any recipe code from the verl (volcengine/verl) codebase to Relax. Handles reward, rollout, tool/env, dataset, and launch script conversion. Supports both colocate (default)…

580 4d ago A 78 tokens original Apache-2.0

unsloth-buddy

05

TYH-labs/unsloth-buddy

Skill Claude CodeCodex

This skill should be used when users want to fine-tune language models or perform reinforcement learning (SFT, DPO, GRPO, ORPO, KTO, SimPO) using the highly optimized Unsloth library. Covers environment setup, LoRA patching, VRAM optimization, vision/multimodal fine-tuning, TTS, embedding training, and…

276 2mo ago C 123 tokens original MIT

grpo-rl-training

06

liortesta/ClawdAgent

Skill Claude CodeCodex

Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.

11 4d ago A 26 tokens copy · 100% Apache-2.0