Agent
RL algorithm expert. Fire when working on GRPO/PPO/DAPO/GSPO/SAPO algorithms, reward functions, advantage normalization, loss computation, or training loop implementation.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Agent
RL algorithm expert. Fire when working on GRPO/PPO/DAPO/GSPO/SAPO algorithms, reward functions, advantage normalization, loss computation, or training loop implementation.
Agent
FSDP backend expert. Fire when working on FSDP-based training, parameter sharding, FSDP weight update, CPU offloading, or troubleshooting FSDP-related issues.
Agent
Ray orchestration & service deployment expert. Fire when working on Ray Serve deployment, placement groups, service lifecycle, rollout engine management, health monitoring, or troubleshooting job launch and GPU allocation issues.
Agent
An expert assistant for integrating and configuring Megatron, a system for training very large machine-learning models across multiple processors or machines.
Agent
Ray framework expert. Fire when working on Ray cluster management, ray.init/ray.remote/ray.get patterns, placement groups, scheduling strategies, Ray Serve deployments, ray job submit, runtime environments, or troubleshooting Ray-specific errors (serialization, object store, GCS, scheduling failures).