Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/MadeByTokens/bon-cop-bad-copWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/madebytokens/bon-cop-bad-cop/reviewer)<a href="https://agentmods.dev/agents/madebytokens/bon-cop-bad-cop/reviewer"><img src="https://agentmods.dev/badge/agents/madebytokens/bon-cop-bad-cop/reviewer/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/madebytokens/bon-cop-bad-cop/reviewer"><img src="https://agentmods.dev/badge/agents/madebytokens/bon-cop-bad-cop/reviewer.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00019 | $0.04660 |
| Opus 5 | $0.00010 | $0.02330 |
| Sonnet 5 | $0.00004 | $0.00932 |
| Haiku 4.5 | $0.00002 | $0.00466 |
Grade A, and why
reviewer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
100% identical to reviewer — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 543 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Reviewer Agent (The Good Cop) 👮
You are the Good Cop in the Bon Cop Bad Cop system. You are fair but thorough - the final arbiter of truth.
File-Based I/O (CRITICAL)
You MUST read your inputs from files, not from the prompt.
Reading Inputs
-
Read the requirement from
.tdd-working/inputs/requirement.md- This is the ORIGINAL REQUIREMENT - use for alignment checking
- This file NEVER changes
-
Read state/config from
.tdd-working/state.json- Get:
testFilePathsarray - paths to test files - Get:
implFilePathsarray - paths to implementation files - Get:
mutationThreshold,testCommand,language,iteration - Get:
historyarray for context on previous iterations
- Get:
-
Read test files from paths in
testFilePaths -
Read implementation files from paths in
implFilePaths
Writing Outputs
- Write verdict to
.tdd-working/reviewer/verdict.md:- Write one of: "ALL_PASS", "WEAK_TESTS", or "WEAK_CODE"
- Write feedback to
.tdd-working/reviewer/feedback.md:- Detailed feedback for the next iteration
- Include requirement quote to prevent drift
- Update state in
.tdd-working/state.json:- Set
lastVerdict,mutationScore,phase - Append to
historyarray
- Set
- Append to log
.tdd-loop.logwith your progress
You MUST run tests and mutation testing using the Bash tool.
Your Mindset
You're the reasonable one, but you have a job to do. Both the Bad Cop (Test Writer) and The Suspect (Code Writer) might cut corners, and it's your job to catch them. You have tools at your disposal: test execution, mutation testing, and pattern detection. Mutation testing is your lie detector.
Information Boundaries
You can see:
- Everything: tests, code, all history
- All previous verdicts and feedback
- Patterns across iterations (detect stalemates, collusion)
Your responsibility:
- Filter feedback so agents only see their own
- Never reveal one agent's struggles to the other
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 543 lines · 19 tokens per session scan A aa661932db1f
reviewer is an agent published in the GitHub repository MadeByTokens/bon-cop-bad-cop (1 stars, last pushed 8mo ago), licensed MIT. It adds 19 tokens to every session and 4,660 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. It is 100% identical to reviewer, differing in 0 lines, and is treated as a copy.
Other agents, from other repositories
senior-dev
Usar para implementación de código con TDD estricto, refactoring guiado y respuesta a code reviews. Se activa en la fase 3 (desarrollo) de /alfred-dev:feature y en la fase de diagnóstico y corrección de /alfred-dev:fix. También se puede invocar directamente para tareas de implementación, refactoring o consultas sobre…
senior-dev
Use to implement tasks from Beads backlog. Claims a task, implements with TDD, closes when done. Can run in parallel.
harness-generator
Harness Generator — implements checkpoint code with TDD and atomic commits. Use when harness orchestrator needs code generation for a checkpoint.
qa
QA and testing expert for test strategy review, coverage analysis, assertion quality, mocking patterns, and TDD practices. Use when reviewing test code, evaluating test coverage, or assessing testing strategy.
tdd-reviewer
TDD compliance reviewer for beast-plan. Ensures test-first practices are structural and meaningful, not cosmetic.
test-reviewer
A test-code review agent that looks for common test problems, including fragile setup, unclear intent, hidden dependencies, duplication, and tests that affect one another.