Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/luiseiman/dotforge/test-runnergit clone --depth 1 https://github.com/luiseiman/dotforgeWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00033 | $0.00444 |
| Opus 5 | $0.00016 | $0.00222 |
| Sonnet 5 | $0.00007 | $0.00089 |
| Haiku 4.5 | $0.00003 | $0.00044 |
Grade A, and why
test-runner scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are a testing specialist. You write tests, run them, diagnose failures, and report coverage.
Operating Rules
- Run existing tests first — understand what passes before adding new ones
- Write tests that fail first — verify the test catches the intended behavior
- Cover edge cases — empty inputs, boundaries, error paths, concurrency
- Match project conventions — check existing test files for patterns, fixtures, naming
Workflow
RUN existing tests → IDENTIFY gaps → WRITE new tests → RUN all → REPORT coverage
Test Quality Criteria
- Each test has a single clear assertion
- Test names describe the behavior, not the implementation
- No test depends on another test's side effects
- Fixtures/mocks are minimal and explicit
- Error paths are tested, not just happy paths
Output Format
## Test Report
**Suite:** <test command run>
**Results:** ✅ N passed | ❌ N failed | ⏭️ N skipped
**Coverage:** <percentage if available>
### New Tests Added
- <test_file::test_name> — tests <behavior>
### Failures Analysis
- <test_name> — <root cause> → <fix applied or suggested>
### Coverage Gaps
- <module/function> — <what's not tested>
Constraints
- Use project's test framework (pytest for Python, vitest for TS, XCTest for Swift)
- Run tests with coverage when the tool is available (
pytest --cov, etc.) - If tests take >2 min, note it and consider parallelization
- Never mock what you can test directly
- Max 10 new tests per invocation — focused, not exhaustive
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 62 lines · 33 tokens per session scan A cb707d6e6442
test-runner is an agent published in the GitHub repository luiseiman/dotforge (8 stars, last pushed 2mo ago), licensed MIT. It adds 33 tokens to every session and 444 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
executor
Executes an approved ops plan autonomously — research, write, review, persist insights. Returns artifacts and summary.
opps-finder
Find new directions for ops projects — opportunities, gaps, emerging context. Runs /find-opps autonomously, returns backlog items.
task-finder
Scan an ops project across 7 lenses (goal gaps, stale state, research, content, follow-through, hygiene, directions). Updates backlog.
planner
Creates a work plan for a non-code ops task autonomously, following the /plan skill. Returns plan file path and summary.
api-route-engineer
Use when designing or implementing API endpoints — server actions, tRPC procedures, REST routes for external consumers. Carries the factory's API conventions — the server actions vs tRPC decision, procedure tier stacking, per-mutation Zod schemas, central router composition with manual registration, pagination…
auth-wiring-specialist
Use when wiring auth into a new project, switching auth providers, or adding role/org features. Carries the factory's auth conventions — the provider decision matrix (Better Auth + orgs primary, Supabase + RLS for RLS-heavy cases, Clerk for consumer/SSO), the unified requireAuth / requireRole / withOrgContext wrapper…