Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/homenshum/nodebenchai/scenario-testinggit clone --depth 1 https://github.com/HomenShum/NodeBenchAIWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00457 |
| Opus 5 | $0.00000 | $0.00229 |
| Sonnet 5 | $0.00000 | $0.00091 |
| Haiku 4.5 | $0.00000 | $0.00046 |
Grade A, and why
scenario-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Scenario Testing Review
Review the tests in the current file or feature against the scenario-based testing mandate.
What to check
For every test in scope, verify it meets ALL six criteria:
- Who — Is a specific user persona defined? (not just "a user")
- What — Is the test starting from a user goal, not a function signature?
- How — Are action sequence, timing, and concurrency explicitly specified?
- Scale — Does the test define behavior at 1 user AND 10+ concurrent? If not, flag it.
- Duration — Does the test cover both short-running (burst) AND long-running (sustained) scenarios?
- Failure modes — Are edge cases, degraded conditions, and adversarial inputs covered?
Output format
For each test found, output:
Test: <test name>
Persona: <defined / MISSING>
Goal orientation: <user goal / function-signature-oriented — REWRITE>
Scale coverage: <1x only / 10x / 100x>
Duration: <short-only / long-only / BOTH>
Failure modes covered: <list or NONE>
Verdict: PASS / NEEDS REWORK
Then for each NEEDS REWORK:
- Explain what's missing
- Provide a rewritten scenario docstring using the anatomy template:
Scenario: <name>
User: <persona>
Goal: <user goal>
Prior state: <system state before scenario>
Actions: <sequence with timing>
Scale: <1x / 10x / 100x>
Duration: <single request / session / multi-day>
Expected: <state + side effects + UI>
Edge cases: <degraded / adversarial / partial>
Anti-patterns to catch and flag
- Test has no persona ("user" with no context)
- Clean DB every time — no accumulated state
- No concurrency
- All assertions on return values only, no side effect checks
mockImplementation(() => ({}))covering every dependency- Happy path only — no sad paths
- "It passes in CI" with no production-scale verification plan
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 54 lines · 0 tokens per session scan A afc0fe51ca05
scenario-testing is a command published in the GitHub repository HomenShum/NodeBenchAI (14 stars, last pushed 19d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 457 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
template
Manage issue templates for streamlined issue creation.
design-review
Workflow recipe — review a design end-to-end, ending in measured numbers rather than adjectives, by chaining 4 skills.
setup-pm-skills
Onboard a new user — find out what they do, recommend the right bundles & top skills, and set up a project CONTEXT.md so every skill is tailored to them.
security-review
CWE 기반 보안 검토 + STRIDE 위협 모델링 (v6 - effort:max 강제).
auto-browse
Auto-browse — learn, optimize, and graduate browser operations or web data-mining workflows.
release-harn
Run the tag-first Harn release workflow.