The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.
These files are openclaw/shellbench's own configuration. They tell Claude Code and Codex how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
.agents/skills/crabbox/SKILL.md A 36 tok .agents/skills/shellbench-research-runbook/SKILL.md A 76 tok