post training skills

53 tagged post training, measured the same way as everything else here.

Browse within: GRPO 27Distributed Training 19diffusers 17diffusion-models 17distillation 17inference 17video-generation 17Multi-Agent 15Multimodal 14agentic-rl 14RLHF 11PPO 10reinforcement-learning 10DPO 8

add-model

02

hao-ai-lab/FastVideo

Skill Claude CodeCodex

Manual /add-model workflow for implementing a FastVideo model or first-class component port after add-model-01-prep has staged reference code and weights. Organizes the port into numbered phases with conversion rules, component policies, parity gates, and handoff checks.

4.2k 2d ago A 55 tokens original Apache-2.0

hao-ai-lab/FastVideo

Skill Claude CodeCodex

Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, environment-caused benchmark shift, or reviewed v2 calibration using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed…

4.2k 2d ago A 162 tokens original Apache-2.0

hao-ai-lab/FastVideo

Skill Claude CodeCodex

Seed HF reference artefacts for a single newly-added SSIM test (pixel .mp4 for runtexttovideosimilaritytest-style tests, or latent .pt for runtexttolatentsimilaritytest-style tests). Runs the test on Modal L40S, downloads the generated artefacts via modal volume get, pauses for the user to verify (visual eyeball for…

4.2k 2d ago A 146 tokens original Apache-2.0

debug-hang

05

redai-infra/Relax

Skill Claude CodeCodex

A troubleshooting workflow for Ray, a system that runs machine-learning jobs across multiple computers or GPUs, when a distributed training job stops making progress.

580 4d ago A 67 tokens original Apache-2.0

redai-infra/Relax

Skill Claude CodeCodex

Integrate a new NVIDIA NeMo Gym environment into Relax as a three-step recipe. Use when adding or debugging a recipe under examples/nemogymagentic/recipes; covers data preparation, a local private Gym service, direct Ray training launch, verifier validation, callback networking, lifecycle cleanup, and failure triage.

580 4d ago A 73 tokens original Apache-2.0

verl-to-relax

07

redai-infra/Relax

Skill Claude CodeCodex

Migrate RL training recipes from verl to Relax framework. Use when user wants to port reward functions, tool environments, training scripts, or any recipe code from the verl (volcengine/verl) codebase to Relax. Handles reward, rollout, tool/env, dataset, and launch script conversion. Supports both colocate (default)…

580 4d ago A 78 tokens original Apache-2.0

mcp_demo

08

sinanuozdemir/advanced-ai-intensive

Skill Claude CodeCodex

You have a shell. You can write Python files and run them. Use that to answer the user's question by searching a pre-built corpus.

39 1mo ago A 0 tokens

tinker-debug

09

gvkhosla/pi-tinker

Skill Claude CodeCodex

Pi overlay for diagnosing Tinker training, renderer, and deploy issues. Use when training is slow, hung, mismatched vs vLLM/SGLang, or an error is opaque. Canonical triage lives in Tinker Cookbook.

24 4d ago A 52 tokens original Apache-2.0

tinker-inkling

10

gvkhosla/pi-tinker

Skill Claude CodeCodex

Sample, evaluate, and post-train Inkling and Inkling-Small from Pi. Use when the user mentions Inkling, tml-renderers, TMLv0, thinking effort, Inkling-Small, or Inkling audio/images. Load before writing Inkling training or eval code.

24 4d ago A 66 tokens original Apache-2.0

tinker-research

11

gvkhosla/pi-tinker

Skill Claude CodeCodex

Pi overlay for Tinker post-training. Use when the user wants to plan SFT/RL/DPO/distillation, pick a model, or run experiments from Pi. Canonical methodology lives in Tinker Cookbook — do not reinvent it here.

24 4d ago A 54 tokens original Apache-2.0

grpo-rl-training

12

liortesta/ClawdAgent

Skill Claude CodeCodex

Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.

11 5d ago A 26 tokens copy · 100% Apache-2.0