Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/homenshum/nodebenchai/nodebench-qagit clone --depth 1 https://github.com/HomenShum/NodeBenchAIWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00301 |
| Opus 5 | $0.00000 | $0.00151 |
| Sonnet 5 | $0.00000 | $0.00060 |
| Haiku 4.5 | $0.00000 | $0.00030 |
Grade A, and why
nodebench-qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
NodeBench QA Loop
Full QA loop: crawl → findings → fix suggestions → re-crawl → savings report.
Steps
-
Crawl the site Call
site_map({ url: '<target-url>' })to crawl all pages and get a session ID. -
Review findings Call
site_map({ session_id: '<id>', action: 'findings' })to see all QA issues. -
Generate test suggestions Call
suggest_tests({ session_id: '<id>' })to get scenario-based test cases. -
Fix issues For each finding, trace the root cause (5-whys) and apply minimal targeted fixes.
-
Re-crawl and diff Call
diff_crawl({ url: '<target-url>', baseline_id: '<original-session-id>' })to compare before/after. -
Report savings Call
compare_savings()to show token usage, time saved, and cost estimates.
Quick one-liner
If you have a running dev server:
site_map({ url: 'http://localhost:5191' })
Then follow the next suggestions in the response to drill down.
When to use
- After completing any implementation sprint
- Before any deploy or PR
- When the user says "QA", "check the site", "run quality checks"
- As part of the continuous flywheel loop
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 40 lines · 0 tokens per session scan A cd68d5615ffd
nodebench-qa is a command published in the GitHub repository HomenShum/NodeBenchAI (14 stars, last pushed 19d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 301 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
triage-ci
Triage E2E failures via local reruns or CI artifacts, classify flaky vs real-bug, present findings report.
implement
E2E-first Test driven development — unit tests written last.
write-tests
Create E2E agent runner (no unit tests).
debug
Deep-dive artifact analysis for diagnosing E2E test failures.
web-checklist
머지 후 웹 테스트 체크리스트 + 완료 추적 (v6).
verify
Grade work that already exists and decide whether it can merge. Runs the project's current unit, integration, and E2E suites plus security scanning and type checking, scores every dimension 0-10, and returns a merge verdict with a VERIFIED-vs-CLAIMED evidence manifest. Writes no test files and edits no source. Use…