Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/nirecom/agents/review-testsnpx skills add nirecom/agents --skill review-testsgit clone --depth 1 https://github.com/nirecom/agentsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/nirecom/agents/review-tests)<a href="https://agentmods.dev/skills/nirecom/agents/review-tests"><img src="https://agentmods.dev/badge/skills/nirecom/agents/review-tests.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00011 | $0.01604 |
| Opus 5 | $0.00005 | $0.00802 |
| Sonnet 5 | $0.00002 | $0.00321 |
| Haiku 4.5 | $0.00001 | $0.00160 |
Grade A, and why
review-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 59 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Review test case completeness against source code via Codex (single round, no re-loop).
Procedure
Note: the Stop-guard silence during dispatch is automatic (PostToolUse marks the step in_progress). Do not emit NEXT_STEP_PAUSE.
Read rules/shell-commands.md before the first Bash command, or before writing a file — defensive measure: RT-2's incident showed the rule content was not effectively available at Bash-issuance time in this context: fork execution.
RT-0. Resolve the session-bound linked worktree path: run "$AGENTS_CONFIG_DIR/bin/resolve-worktree-path" (Bash, as a single standalone command — no variable-capture syntax on the Bash tool's own command line, per rules/shell-commands.md); its stdout is WORKTREE for later steps.
If WORKTREE == "NOSTATE", treat WORKTREE as empty — the internal scripts handle the CWD-fallback path for that case.
RT-0a. Read:
rules/core-principles.mdrules/test.md— on-demand-only; never auto-injected, so this Read is mandatoryskills/_shared/test-design.mdskills/_shared/test-design/protection-fix-tests.md— additionally, for security / guard / classifier fix targetsskills/_shared/test-design/parser-regex-tests.md— additionally, for parser / regex / allowlist targets RT-1. Identify staged test file(s) and source file(s):- Run
"$AGENTS_CONFIG_DIR/skills/review-tests/scripts/select-staged-files.sh"(Bash, single standalone command); its stdout isSTAGED. - If exit 3 (linked worktree unresolvable): do NOT fall back to cwd;
present "Could not identify the linked worktree. Re-run
/review-testsfrom the linked worktree, or specify the test and source files manually." and ask the user for the files. - Select test file(s) and source file(s) from
$STAGEDor from the user's manual input. RT-2. Assemble review input via the Write tool only — concatenate test file(s) and source file(s) contents into<PLANS_DIR>/<session-id>-test-review.md. Do not substitute Bash-based assembly for the Write tool call in this step — seerules/shell-commands.mdTool Selection Priority for what counts as shell-based writing. Resolve<PLANS_DIR>viaskills/_shared/resolve-plans-dir.md. InitializeEXTENSIONS_USED=0. RT-3. Invoke"$AGENTS_CONFIG_DIR/skills/review-tests/scripts/run-codex-review-loop.sh"(Bash), exportingAGENTS_CONFIG_DIR,SESSION_ID,PLANS_DIR,EXTENSIONS_USED. The wrapper auto-adds--context test-design.md. Exit-code handling (SSOT:skills/_shared/codex-review-loop.md; single-round — no re-loop): - exit 0 APPROVED → RT-5 COMPLETE.
- exit 1 NEEDS_REVISION → terminal; save stdout to
<PLANS_DIR>/<session-id>-test-review-codex-round-<N>-raw.md(<N>from<PLANS_DIR>/<session-id>-test-review-last-round.txt); present gaps; suggest specific test cases → RT-5 WARNINGS (no re-loop). - exit 2 ESCALATE → run
review-loop-summarize-concerns --budget-remaining 0; present summary → RT-5 WARNINGS. - exit 6 HIGH_UNRESOLVED → save stdout to
<PLANS_DIR>/<session-id>-test-review-codex-round-<N>-raw.md(<N>from<PLANS_DIR>/<session-id>-test-review-last-round.txt); runreview-loop-summarize-concerns --budget-remaining 0; present unresolved HIGH concerns → RT-5 WARNINGS (do not emitWORKFLOW_REVIEW_TESTS_COMPLETE). - exit 8 → terminal guard was previously armed (test files unchanged since last terminal); HALT.
- exit 3 → silently launch
test-reviewersubagent; APPROVED → RT-5 COMPLETE; NEEDS_REVISION → RT-5 WARNINGS. - exit 4 → HALT with blocking error; do NOT launch fallback; do NOT emit sentinel.
- exit 5 → does not occur (MAX_EXTENSIONS=0); treat as exit 4 HALT if received.
- exit 7 FINALIZE_FAILED →
<PLANS_DIR>/<session-id>-test-review-unresolved-concerns.jsoncould not be written; HALT, surface the## Concern Ledger: FINALIZE-FAILEDline, launch no fallback, emit no sentinel. After an ESCALATE, confirm the artifact with"$AGENTS_CONFIG_DIR/bin/concern-ledger" check-finalized --plans-dir <PLANS_DIR> --session-id <session-id> --format test-reviewbefore RT-5. RT-4. Triage the concerns againstskills/_shared/priority-hierarchy.mdbefore emitting the sentinel: a concern that contradicts an approved intent.md / outline.md / detail.md decision — including a documented TL3 gap or a deferral to manual verification — is rejected, not a gap. State each rejection and the decision it rests on, and exclude it from the RT-5c warnings count. Skip on exit 0 (no concerns). RT-5. Emit workflow sentinel — two separate Bash calls, not chained: - RT-5a. Run
node "$AGENTS_CONFIG_DIR/bin/compute-staged-tests-token.js" "<WORKTREE-or-empty>"(Bash, single standalone command,<WORKTREE-or-empty>substituted with RT-0's resolved value); its stdout isTOKEN. - RT-5b. (adequate)
echo "<<WORKFLOW_REVIEW_TESTS_COMPLETE: token=${TOKEN}>>" - RT-5c. (gaps/warnings)
echo "<<WORKFLOW_REVIEW_TESTS_WARNINGS: token=${TOKEN} warnings=N — blocking: /write-code stays blocked until the gaps are addressed and /review-tests is re-run>>" - RT-5d. Skip when
WORKFLOW_WRITE_TESTS_NOT_NEEDEDwas emitted (propagated skip).
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today Changed e5c821e29706
- 4d ago First seen · 59 lines · 11 tokens per session scan A b6b27e04603d
review-tests is a skill published in the GitHub repository nirecom/agents (3 stars, last pushed today), licensed MIT. It adds 11 tokens to every session and 1,604 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…