openclaw/shellbench

The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.

These files are openclaw/shellbench's own configuration. They tell Claude Code and Codex how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

139Stars on the repository
2Files it configures its agents with
Tokens loaded in every session
2Agents configured

Skills