openclaw/shellbench

The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.

138Stars on the repository
2Mods indexed here, across every type
4d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

crabbox

01

openclaw/shellbench

Skill Claude CodeCodex

Use Crabbox for ClawBench remote Linux validation. Default to Blacksmith Testbox; includes direct Blacksmith and owned AWS fallback notes when Crabbox fails.

138 4d ago A 36 tokens original MIT

openclaw/shellbench

Skill Claude CodeCodex

Plan, smoke-test, execute, checkpoint, publish, audit, and reproduce full ShellBench native benchmark campaigns across OpenClaw, Hermes, Codex, and Claude Code, including model and reasoning identity, pinned harness versions, n=3 qualification through n=6 research runs, S3 trace retention, and…

138 4d ago A 76 tokens original MIT