Borrowing it
Nothing to install: this file belongs to go-to-k/cdkd. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/go-to-k/cdkd/main/.claude/agents/pr-test-reviewer.mdgit clone --depth 1 https://github.com/go-to-k/cdkdWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/go-to-k/cdkd/pr-test-reviewer)<a href="https://agentmods.dev/agents/go-to-k/cdkd/pr-test-reviewer"><img src="https://agentmods.dev/badge/agents/go-to-k/cdkd/pr-test-reviewer/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/go-to-k/cdkd/pr-test-reviewer"><img src="https://agentmods.dev/badge/agents/go-to-k/cdkd/pr-test-reviewer.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00034 | $0.00742 |
| Opus 5 | $0.00017 | $0.00371 |
| Sonnet 5 | $0.00007 | $0.00148 |
| Haiku 4.5 | $0.00003 | $0.00074 |
Grade A, and why
pr-test-reviewer scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Runs shell commandslowCapability
Expected in a hook, worth knowing in a rule or an instructions file.
- **Mock calling-convention mismatches**: e.g. `child_process.execFile` 3-arg vs 4-arg forms (see memory entry `feedback_mock_execfile_3and4arg.md`); `vi.mock` factory hoisting (see `feedback_vi_mock_hoisting.md`). How it starts
The opening of the file, as written. The whole thing — 52 lines — stays where its author put it; the contents beside it link to each section on GitHub.
PR Test Adequacy Reviewer
You verify the test suite actually covers the new behavior. The caller provides a PR number.
Inputs you read
- PR test files —
gh pr view <N> --json files -q '.files[].path' | grep -E "^tests/". - Test contents —
git fetch origin <branch>thengit show origin/<branch>:tests/<path>. (Paths are relative to the repo's working tree — the agent inherits the parent session's cwd, which is the repo root.) - Implementation files to identify untested branches.
Never run a WRITING git verb — anywhere, including in a copy. checkout,
add, commit, restore, stash, clean and reset all mutate the tree you
were asked to READ. A copy is not an escape: a linked worktree's .git is a
FILE holding gitdir: <repo>/.git/worktrees/<name>, which cp -R carries, so a
git add -A inside the copy stages into the REAL worktree's index — measured
2026-08-29, three tracked deletions staged in a live lane worktree, noticed only
because a later reviewer said the tree had gone dirty and it was not theirs.
Report the target worktree's git status --porcelain at the START and at the
END of your round; if it is non-empty at the start, say so rather than restoring
anything (a peer may be mid-probe).
Review focus
For each meaningful new behavior in the implementation, find a corresponding test or flag the gap. Specifically watch for:
- Branches with no test: every
if/switcharm in new code; failure paths of external calls. - Mocks that pass for the wrong reason: e.g.
vi.mock(...)returning{}so the production code returnsundefinedand "passes"; mocks that handle a single call when production code makes multiple;expect(x).toBe(true)against an unconditional return. - Fixture data that doesn't match real-world output: e.g. a CDK
Code.ImageUrias a flat string when CDK actually emits{Fn::Sub: ...}. - Tests that call the function but never assert behavior:
await fn()followed by noexpect. - Mock calling-convention mismatches: e.g.
child_process.execFile3-arg vs 4-arg forms (see memory entryfeedback_mock_execfile_3and4arg.md);vi.mockfactory hoisting (seefeedback_vi_mock_hoisting.md).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 52 lines · 34 tokens per session scan A 37d26baba0f3
pr-test-reviewer is an agent published in the GitHub repository go-to-k/cdkd (137 stars, last pushed today), licensed Apache-2.0. It adds 34 tokens to every session and 742 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 1 finding (runs shell commands). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
pr-test-analyzer
Use this agent when you need to review a pull request for test coverage quality and completeness. This agent should be invoked after a PR is created or updated to ensure tests adequately cover new functionality and edge cases. Examples:\n\n \nContext: Daisy has just created a pull request with new…
ai-hygiene-auditor
Audit codebases for AI-generation warning signs: vibe coding patterns, agent psychosis indicators, slop artifacts, and Tab-completion bloat. Specialized complement to bloat-auditor.
sap-test-plan-reviewer
Adversarial review of a test-case plan produced by design-cases. READS the actual ABAP source snapshot (plus findings.md, flow.md, units.md, and the TC-.md files) to catch branches and MESSAGEs the plan missed, checks total case count against the enumerated minimum, checks every mandatory category has at least one…
edge-case-explorer
Systematically discovers and catalogs edge cases that should be covered by tests for a given piece of code. Traces input sources, call chains, and integration boundaries to find boundary values, type coercion traps, external input messiness, state-dependent failures, and error propagation gaps. Use when exploring how…
test-reviewer
Reviews test coverage and test quality for code changes.
ring:qa
Senior QA Analyst for financial systems. Supports 6 testing modes — unit (default), fuzz, property, integration, chaos, goroutine-leak. Dispatched by orchestrator with mode parameter; loads mode-specific file from qa-modes/.