jackjin1997/AgentBench-Live

Variance-aware benchmark for AI coding agents. Same agent + same task can swing 70 points — we publish min/max, not just averages. Claude Code · Gemini CLI · Codex CLI · Aider · 10 tasks · Docker sandbox · MIT.

Latest release v0.2.0 — v0.2.0 — Variance reporting + reframed narrative · 14 May 2026

1 file for Claude Code: AgentBench-Live CLAUDE.md — 536 tokens loaded in every session.

4Stars on the repository
1Files it configures its agents with
536Tokens loaded in every session
1Agent configured

Instructions

These files are jackjin1997/AgentBench-Live's own configuration — they tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.