coder-eval
01Plugin Claude Code
Test whether your Claude Code skills actually trigger, and benchmark AI coding agents against your own tasks.
119 3d ago A
tokens not measured
original Apache-2.0
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
Plugin Claude Code
Test whether your Claude Code skills actually trigger, and benchmark AI coding agents against your own tasks.
Plugin Claude Code
Test whether your Claude Code skills actually trigger, and benchmark any coding agent — author, run, and analyze the eval suite that proves it, locally or as a CI gate.