UiPath/coder_eval

Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.

119Stars on the repository
20Mods indexed here, across every type
3d agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

ANTIGRAVITY

01

UiPath/coder_eval

Agent

Run Google Antigravity (Gemini) as the agent under evaluation in Coder Eval — installation, authentication, model and skill configuration, and how its telemetry maps to sandboxed, weighted scoring.

119 3d ago A 40 tokens original Apache-2.0

CLAUDE_CODE

02

UiPath/coder_eval

Agent

Configure and run the default Claude Code agent in Coder Eval — the full agent-config surface, direct vs. Bedrock authentication, permission modes, sandbox isolation, skills/plugins, early stop, and token telemetry.

119 3d ago A 42 tokens original Apache-2.0

CODEX

03

UiPath/coder_eval

Agent

Run OpenAI Codex as the agent under evaluation in Coder Eval — installation, authentication, task configuration, and how Codex telemetry maps to sandboxed, weighted scoring.

119 3d ago A 35 tokens original Apache-2.0

HARNESS_PARITY

04

UiPath/coder_eval

Agent

One task file, run on any harness, must be the same task. runlimits.maxturns was the field that broke that promise hardest: Claude Code enforced it, and Codex and Antigravity accepted it and never read it, so maxturns: 6 ran capped on one backend and unbounded on the other two.

119 3d ago A 0 tokens original Apache-2.0

OPENCODE

05

UiPath/coder_eval

Agent

Run OpenCode, the open-source terminal coding agent, as the agent under evaluation in Coder Eval — installation, provider authentication, model selection, and how its event stream maps to sandboxed, weighted scoring.

119 3d ago A 42 tokens original Apache-2.0