Playwright for coding agents. Benchmark Claude Code, Codex, Gemini, and OpenCode on your own tasks - and test that your skills, MCP servers, and CLIs work when an agent uses them. Sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Latest release v0.11.7 · 8 Sept 2026
1 file for Claude Code: coder_eval CLAUDE.md — 11,227 tokens loaded in every session.
CLAUDE.md B 11,227 tok These files are UiPath/coder_eval's own configuration — they tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.