Agents' Last Exam
Agents' Last Exam is an evaluation framework and benchmark for testing AI agents on long, economically relevant tasks with outcomes that can be checked. Researchers and developers use its sandboxed tasks, agent runner, and deterministic graders to measure performance across professional work in 55 subdomains and 13 industry groups.
Latest release runtime-qe-bgw-6.7.0-4.0 — QE 6.7MaX / BerkeleyGW 4.0 runtime · 13 Aug 2026
These files are rdi-berkeley/agents-last-exam's own configuration. They tell Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 1,548 tok