Variance-aware benchmark for AI coding agents. Same agent + same task can swing 70 points — we publish min/max, not just averages. Claude Code · Gemini CLI · Codex CLI · Aider · 10 tasks · Docker sandbox · MIT.
Latest release v0.2.0 — v0.2.0 — Variance reporting + reframed narrative · 14 May 2026
1 file for Claude Code: AgentBench-Live CLAUDE.md — 536 tokens loaded in every session.
CLAUDE.md A 536 tok These files are jackjin1997/AgentBench-Live's own configuration — they tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.