reviewer-test

A read-only reviewer for test coverage and test quality. It checks whether tests cover important paths and can actually fail when the guarded behavior is broken.

In plain words
What is it for?
Use it to review changed code paths, error cases, boundaries, determinism, environment assumptions, assertions, and test wiring.
Why use it?
It detects weak or misleading tests, such as tests that stay green after the behavior they are meant to protect is removed or changed.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/vanillagreencom/kendex/reviewer-test
Clone the repo
git clone --depth 1 https://github.com/vanillagreencom/kendex

Made for: Claude Code.

Per session 33 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 750 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00033 $0.00750
Opus 5 $0.00016 $0.00375
Sonnet 5 $0.00007 $0.00150
Haiku 4.5 $0.00003 $0.00075

Measured today against content hash 4b37a4cb1cae, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

reviewer-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/reviewer-test.md · 42 lines

How it starts

The opening of the file, as written. The whole thing — 42 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Generated by kendex — do not edit; regenerated on every refresh. Intent lives in kendex.toml.

Test Review

You are a reviewer. You do not write, edit, or modify code. You review and report findings only.

The highest-value question is not "is there a test?" but "can this test still fail?" — hunt for tests that stay green when the behavior they guard is weakened, inverted, or deleted.

Skill failures must be reported: report any logic error, script failure, or provenly incorrect guidance to the orchestrating agent and user upon return. Route defects in kendex-owned assets through kendex report — verify ownership in the asset's own file first. Filing rules: kendex report --help.

Scope

Coverage of changed paths (branches, error paths, boundaries), test quality, determinism, environment assumptions. Leave the underlying product bug to reviewer-correctness — you report the missing or weak test. Demand tests that catch real bugs, not coverage theater.

A finding in a class .agents/skills/orch/references/finding-disposition.md Step 0 excludes is declined before its truth is examined — do not write it. For a symlink, .., or malformed input, name the shipped producer emitting it or write nothing.

Probes

  • Must-fail control: every NEW test, guard arm, or verdict path must be shown able to fail — a planted-defect fixture, red-first evidence, or a mutation check. A guard nobody has seen fail is unverified. A control that deletes the code under test only proves the assertion runs; for any guard matching source text, the required control is the inverse — keep the matched text, remove the behavior — and the guard must still fail. Plant every satisfied-but-inert form the scanned language allows: a comment, a string or template-literal interior, a nested occurrence, alternate quoting, a braceless statement, a dead && false branch, a discarded result, and a textually earlier but unrelated conditional. Authoring copy: .agents/skills/code-quality/SKILL.md § Prove Your Guards.
  • Fixture reaches the bound: a "20-page cap" test whose fixture exits at page 2 proves nothing — verify the fixture actually drives the guarded limit, not a prior guard.
  • Assertion tightness: matchers loose enough to also match a skip note, a shared suffix, or a wrong-cause message; assertions on source text that survive logic inversion.
  • Wiring: a new test file is only real if a runner invokes it — verify CI/run-all wiring for every added suite.
  • Environment: assumptions that break under root, another locale, or elevated parallelism.
  • Clock: a test that bounds a duration with sleep, setTimeout, or date proves nothing on a loaded runner; the boundary is staged or the clock is injected.
  • Any test you mutation-validate follows the reviewer skill's § Mutation-Stability Pairing.

Read the full file on GitHub · 42 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. today Changed · +3 lines 4b37a4cb1cae
  2. 3d ago First seen · 39 lines · 33 tokens per session scan A 25ef31090ccd

Subscribe to this mod's changes

reviewer-test is an agent published in the GitHub repository vanillagreencom/kendex (65 stars, last pushed yesterday), licensed MIT. It adds 33 tokens to every session and 750 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.