Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/randomittin/heimdall/benchgit clone --depth 1 https://github.com/randomittin/heimdallWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00101 | $0.00994 |
| Opus 5 | $0.00051 | $0.00497 |
| Sonnet 5 | $0.00020 | $0.00199 |
| Haiku 4.5 | $0.00010 | $0.00099 |
Grade A, and why
bench scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 89 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/hmd:bench — reproduce the public benchmark table
Runs bin/heimdall-bench — the documented entry door over the measurement
harness (bin/benchmark). It does NOT invent numbers: it produces them by
running a fixed suite of representative coding tasks under two arms (raw Claude
Code vs Claude Code driven through Heimdall) and emitting the exact table format
the published evals/ table uses. The number in any Heimdall claim is one a
stranger can reproduce here.
When to use
- Someone challenges a benchmark number — point them at
heimdall benchso they reproduce it on their own machine instead of trusting the README. - Before a launch, to regenerate the table on a pinned model (the flagship).
- To see, with zero API spend, exactly what a real run measures.
Instructions
-
Always start dry (the default). Validates the suite and prints the capture plan with NO API calls and NO agents spawned:
heimdall-bench heimdall-bench --dry # explicit, identicalThis lists every task, its category, and its
verify[]steps, then prints the exact arm commands a live run would issue and what each metric is parsed from. A cold stranger can run this without burning a token. -
Run live only when the user opts in. A
--liverun invokes theclaudeCLI for every task × arm and spends real API tokens. Pin the model for a reproducible, publishable table:heimdall-bench --live --model <model-id>It writes
evals/benchmark/results.jsonl(one machine-readable line per task × arm) andevals/benchmark/results.md(the published-format markdown table). Numbers come from the CLI's own usage accounting and from actually running each task'sverify[]— nothing is hand-tuned. -
Reprint the last live table without re-running:
heimdall-bench --table -
Honesty principle. Publish everything measured, including tasks where Heimdall loses on tokens, wall time, or cost. A pricier arm that ships
7/7beats a cheaper arm that ships3/7; hiding the trade-off would make the whole table untrustworthy.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 89 lines · 0 tokens per session scan A 6ec5ee615455
bench is a command published in the GitHub repository randomittin/heimdall (5 stars, last pushed 11d ago), licensed MIT. It adds 101 tokens to every session and 994 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
autocode
Fully autonomous development pipeline. Give it a feature description or PRD file and it will plan, write tests, implement code, run tests, and verify — all automatically with zero human intervention.
autocode-new
Scan current project and generate a project-specific autocode development pipeline (agents, commands, rules, hooks). Interactive setup wizard.
dispatcher
Pick the next-best repo to work on across the portfolio — rank free repos, recommend one, claim its lease atomically, and route to the entry command.
standup
Show a daily standup summary with completed, in-progress, and blocked tasks across all active epics.
dev-planner
Generate or update DEV-PLAN.md with phased development plan from Product-Spec.md.
iterate
Iteration delivery — delivery check + archive + commit & tag + bump version.