Orchestra-Research/AI-Research-SKILLs
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
53 tagged post training, measured the same way as everything else here.
Browse within: GRPO 27Distributed Training 19diffusers 17diffusion-models 17distillation 17inference 17video-generation 17Multi-Agent 15Multimodal 14agentic-rl 14RLHF 11PPO 10reinforcement-learning 10DPO 8
Orchestra-Research/AI-Research-SKILLs
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
Skill Claude CodeCodex
Manual /add-model workflow for implementing a FastVideo model or first-class component port after add-model-01-prep has staged reference code and weights. Organizes the port into numbered phases with conversion rules, component policies, parity gates, and handoff checks.
Skill Claude CodeCodex
Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, environment-caused benchmark shift, or reviewed v2 calibration using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed…
Skill Claude CodeCodex
Seed HF reference artefacts for a single newly-added SSIM test (pixel .mp4 for runtexttovideosimilaritytest-style tests, or latent .pt for runtexttolatentsimilaritytest-style tests). Runs the test on Modal L40S, downloads the generated artefacts via modal volume get, pauses for the user to verify (visual eyeball for…
Skill Claude CodeCodex
A troubleshooting workflow for Ray, a system that runs machine-learning jobs across multiple computers or GPUs, when a distributed training job stops making progress.
Skill Claude CodeCodex
Integrate a new NVIDIA NeMo Gym environment into Relax as a three-step recipe. Use when adding or debugging a recipe under examples/nemogymagentic/recipes; covers data preparation, a local private Gym service, direct Ray training launch, verifier validation, callback networking, lifecycle cleanup, and failure triage.
Skill Claude CodeCodex
Migrate RL training recipes from verl to Relax framework. Use when user wants to port reward functions, tool environments, training scripts, or any recipe code from the verl (volcengine/verl) codebase to Relax. Handles reward, rollout, tool/env, dataset, and launch script conversion. Supports both colocate (default)…
sinanuozdemir/advanced-ai-intensive
Skill Claude CodeCodex
You have a shell. You can write Python files and run them. Use that to answer the user's question by searching a pre-built corpus.
Skill Claude CodeCodex
Pi overlay for diagnosing Tinker training, renderer, and deploy issues. Use when training is slow, hung, mismatched vs vLLM/SGLang, or an error is opaque. Canonical triage lives in Tinker Cookbook.
Skill Claude CodeCodex
Sample, evaluate, and post-train Inkling and Inkling-Small from Pi. Use when the user mentions Inkling, tml-renderers, TMLv0, thinking effort, Inkling-Small, or Inkling audio/images. Load before writing Inkling training or eval code.
Skill Claude CodeCodex
Pi overlay for Tinker post-training. Use when the user wants to plan SFT/RL/DPO/distillation, pick a model, or run experiments from Pi. Canonical methodology lives in Tinker Cookbook — do not reinvent it here.
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
Skill Claude CodeCodex
Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.