Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/zircote/subcog/run-testsgit clone --depth 1 https://github.com/zircote/subcogWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00041 | $0.01116 |
| Opus 5 | $0.00020 | $0.00558 |
| Sonnet 5 | $0.00008 | $0.00223 |
| Haiku 4.5 | $0.00004 | $0.00112 |
Grade A, and why
run-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 120 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/subcog:run-tests
Execute the automated Subcog functional test suite to validate all MCP tools.
Usage
/subcog:run-tests [options]
Options
--category <name>- Run only tests in specified category--tag <tag>- Run only tests with specified tag--skip-cleanup- Don't run cleanup tests at the end--verbose- Show detailed output for each test--dry-run- Show tests that would run without executing
Execution Strategy
- Read test definitions from
tests/functional/tests.yaml - Initialize test state in
.claude/test-state.json - Execute tests sequentially, tracking results
- Validate each test response against expected patterns
- Generate summary report
State Management:
- State file tracks: current test, results, saved variables
- Each test can save output values for dependent tests
- Tests with
depends_onwait for dependencies to pass
Workflow
-
Initialize Test Environment
- Read
tests/functional/tests.yaml - Parse test definitions and resolve dependencies
- Create
.claude/test-state.jsonwith initial state - Verify Subcog MCP server is available via
subcog_status
- Read
-
Execute Tests Sequentially For each test: a. Check dependencies are satisfied b. Substitute saved variables in action string c. Present the test action to execute d. Wait for execution e. Validate response against
expectrules f. Record pass/fail and save anysave_asvalues g. Continue to next test -
Report Results
- Show pass/fail counts by category
- List any failed tests with details
- Provide cleanup status
- Write report to
tests/functional/report.md
Test Format Reference
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 120 lines · 41 tokens per session scan A 232a79440420
run-tests is a command published in the GitHub repository zircote/subcog (28 stars, last pushed 29d ago), licensed MIT. It adds 41 tokens to every session and 1,116 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
setup
kb MCP 연결 점검·재설정 안내 — 설치 직후 연결 확인, 서버 URL/토큰 변경, 연결 실패 진단에 사용.
loop-forge
Turn a one-line repetitive task into a reusable, self-guarding slash command (loop-forge).
version
Display current guide and Claude Code versions.
index
Index the codebase for semantic search.
docs-review
Phase 4 of documenting-projects: Quality gate with 8 measurable criteria and iteration. Triggers: '/docs-review', invoked by documenting-projects orchestrator.
sap-review
Trigger a structured SAP code review — invoked when ABAP, CDS, RAP, BTP, or integration code needs to be reviewed for quality, performance, security, and clean core compliance.