Agent Claude Code
RL algorithm expert. Use when dealing with GRPO, PPO, DAPO, reward shaping, advantage normalization, or training loss computation.
5.7k
+2 today A 32 tokens
original Apache-2.0
12 tagged llm reasoning, measured the same way as everything else here.
Browse within: llm-agent 12machine-learning-systems 12mlsys 12
Agent Claude Code
RL algorithm expert. Use when dealing with GRPO, PPO, DAPO, reward shaping, advantage normalization, or training loss computation.
Agent Claude Code
Code verification agent. Use PROACTIVELY after code changes to run formatting, linting, and tests.
Agent Claude Code
Expert on cluster launching and resource scheduling (Slurm/Ray/Kubernetes). Use when user modifies launcher/scheduler code, configures cluster resources, or troubleshoots deployment issues.