Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add sd0xdev/sd0x-harness --skill test-healthgit clone --depth 1 https://github.com/sd0xdev/sd0x-harnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/sd0xdev/sd0x-harness/test-health)<a href="https://agentmods.dev/skills/sd0xdev/sd0x-harness/test-health"><img src="https://agentmods.dev/badge/skills/sd0xdev/sd0x-harness/test-health.svg" alt="Measured on agentmods" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00070 | $0.02630 |
| Opus 5 | $0.00035 | $0.01315 |
| Sonnet 5 | $0.00014 | $0.00526 |
| Haiku 4.5 | $0.00007 | $0.00263 |
Grade A, and why
test-health scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 240 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Test Health — Holistic Coverage Measurement
Trigger
- Keywords: test health, coverage measurement, test metrics, coverage trend, test inventory, holistic test audit
When NOT to Use
| Scenario | Alternative |
|---|---|
| Run tests | /verify |
| Review test sufficiency only | /codex-test-review |
| Generate unit tests | /codex-test-gen |
| Feature-doc coverage only | /check-coverage |
| Context-aware test execution + triage | /test-deep |
Workflow
flowchart TD
U[User: /test-health] --> M{Mode?}
M --> |quick| Q[Quick Mode]
M --> |--full| F[Full Mode]
Q --> Q1[Test Inventory]
Q1 --> Q2[Consume Coverage Artifacts]
Q2 --> Q3[Trend Delta]
Q3 --> QR[Quick Dashboard]
F --> A[Phase A: /check-coverage]
A --> B[Phase B: Coverage Collection]
B --> C[Phase C: /codex-test-review]
C --> D[Phase D: Aggregate Dashboard]
D --> T[Trend Snapshot]
T --> FR[Full Dashboard]
Modes
| Mode | Trigger | Content | Duration |
|---|---|---|---|
quick (default) |
/test-health |
Test inventory + consume artifacts + trend delta | <15s |
full |
/test-health --full |
Phase A→B→C→D (feature coverage + instrumentation + qualitative + aggregation) | 2-5min |
Quick Mode Workflow
- Test Inventory: Count test files by layer using Glob (see
references/test-count-parsers.mdfor layer classification). If--scope <path>specified, limit Glob to that directory. If verify-runner cache exists (.claude/cache/verify/), read historical logs for test counts. - Coverage Artifacts: Scan for existing coverage artifacts (see
references/artifact-formats.md). If--scopespecified, scan within scope only. Never execute project commands in quick mode. - Trend Delta: Read previous snapshot, compute delta (see
references/trend-schema.md). Skip if--no-trendflag is set. - Output: Quick Dashboard.
Full Mode Workflow
Phase A: Feature Coverage
What ships with it
6 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 240 lines · 70 tokens per session scan A b76de35a3afa
test-health is a skill published in the GitHub repository sd0xdev/sd0x-harness (188 stars, last pushed 2d ago), licensed MIT. It adds 70 tokens to every session and 2,630 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
building
Implementation skill for writing production code with TDD. Covers the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation patterns. Loaded by component-builder and bug-investigator.
evaluator-write-qa-parallel
Internal Auto-Harness evaluator skill for parallel sprint QA and QA report writing. Use only inside the Evaluator subagent during evaluatorqaparallel.
evaluator-write-qa
Internal Auto-Harness evaluator skill for sprint QA and QA report writing. Use only inside the Evaluator subagent during qa mode.
evaluator-write-final-parallel
Internal Auto-Harness evaluator skill for parallel final QA report aggregation. Use only inside the Evaluator subagent during evaluatorfinalparallel.
evaluator-write-retest-parallel
Internal Auto-Harness evaluator skill for parallel sprint retest and retest report writing. Use only inside the Evaluator subagent during evaluatorretestparallel.
evaluator-write-retest
Internal Auto-Harness evaluator skill for sprint retest and retest report writing. Use only inside the Evaluator subagent during retest mode.