Open-source, local-first evaluation infrastructure for applied AI systems, built for developer and agent workflows.
Latest release v0.0.3 — v0.0.3 - 'qt add' to customize your own benchmark from the remote benchmark registry · 11 Aug 2026
2 files for Codex, OpenCode and Claude Code: quantiles AGENTS.md, quantiles CLAUDE.md — 2,673 tokens loaded in every session.
AGENTS.md A 2,640 tok CLAUDE.md A 33 tok These files are quantiles-evals/quantiles's own configuration — they tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.