Orchestra-Research/AI-Research-SKILLs
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
34 tagged grpo, measured the same way as everything else here.
Browse within: Post-Training 27Distributed Training 19Multi-Agent 14Multimodal 14agentic-rl 14DPO 13reinforcement-learning 12RLHF 11PPO 10HuggingFace 7ai-research 7fine-tuning 7DeepSpeed 6TRL 6
Orchestra-Research/AI-Research-SKILLs
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
Skill Claude CodeCodex
A troubleshooting workflow for Ray, a system that runs machine-learning jobs across multiple computers or GPUs, when a distributed training job stops making progress.
Skill Claude CodeCodex
Integrate a new NVIDIA NeMo Gym environment into Relax as a three-step recipe. Use when adding or debugging a recipe under examples/nemogymagentic/recipes; covers data preparation, a local private Gym service, direct Ray training launch, verifier validation, callback networking, lifecycle cleanup, and failure triage.
Skill Claude CodeCodex
Migrate RL training recipes from verl to Relax framework. Use when user wants to port reward functions, tool environments, training scripts, or any recipe code from the verl (volcengine/verl) codebase to Relax. Handles reward, rollout, tool/env, dataset, and launch script conversion. Supports both colocate (default)…
Skill Claude CodeCodex
This skill should be used when users want to fine-tune language models or perform reinforcement learning (SFT, DPO, GRPO, ORPO, KTO, SimPO) using the highly optimized Unsloth library. Covers environment setup, LoRA patching, VRAM optimization, vision/multimodal fine-tuning, TTS, embedding training, and…
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.