Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemini CLI, Goose, OpenCode, or any custom agent you plug in; verify behavior with rule checks, workspace diffs, multi-judge LLM consensus; pin reliability with pass^k variance across trials. Git worktrees, optional Docker sandbox.
Latest release v0.2.1 · 30 Jul 2026
1 file for Codex and OpenCode: agent-belt AGENTS.md — 2,594 tokens loaded in every session.
AGENTS.md A 2,594 tok These files are jfrog/agent-belt's own configuration — they tell Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.