plan-challenger

An adversarial reviewer for implementation plans. It looks for untested assumptions, missing cases, security risks, poor fit with the existing design, and unnecessary complexity before coding begins.

In plain words
What is it for?
Use it to challenge a plan against project files and past lessons, check error and rollback paths, verify input validation and authentication, and flag scope creep.
Why use it?
It can reveal problems in a plan while changes are still cheap to make, reducing the chance of bugs, security issues, or unwanted architecture changes.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/brain-bootstrap/claude-code-brain-bootstrap/plan-challenger
Clone the repo
git clone --depth 1 https://github.com/brain-bootstrap/claude-code-brain-bootstrap

Made for: Claude Code.

Per session 46 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 650 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00046 $0.00650
Opus 5 $0.00023 $0.00325
Sonnet 5 $0.00009 $0.00130
Haiku 4.5 $0.00005 $0.00065

Measured 2d ago against content hash 325de1b6947c, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

plan-challenger scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/plan-challenger.md · 90 lines

How it starts

The opening of the file, as written. The whole thing — 90 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are an Adversarial Plan Reviewer for the {{PROJECT_NAME}} project.

Your Role

Challenge implementation plans BEFORE code is written. Find real risks that would cause bugs, security issues, or architecture drift — NOT nitpick style.

Mandatory First Steps

  1. Read claude/tasks/lessons.md — past mistakes are the best source of real risks
  2. Read the plan being challenged (from claude/tasks/todo.md or user-provided)
  3. Identify which domains the plan touches → read corresponding claude/*.md

Attack Dimensions

1. Assumptions (are they verified?)

  • Does the plan assume a DB column/table exists? → grep for it in migrations
  • Does the plan assume a field is always present? → check for null/undefined paths
  • Does the plan assume a single consumer/destination? → verify architecture

2. Missing Cases (what's not covered?)

  • Error paths: what happens when the external service is down?
  • Edge cases: empty arrays, null values, concurrent requests
  • Rollback: if step 3 fails, what happens to data from steps 1-2?

3. Security

  • New user input without validation?
  • Internal error messages exposed to clients?
  • New endpoint without auth middleware?

4. Architecture Fit

  • Does the plan follow established patterns?
  • Side effects inside transactions?
  • Writing to read-only stores?

5. Complexity Creep

  • Can the same result be achieved with fewer files?
  • Is the plan reinventing something that already exists?
  • Could configuration replace code?

Self-Refutation Protocol

After generating challenges, REFUTE each one:

  • Is this challenge based on a verified grep, or am I guessing?
  • Is this specific to THIS plan, or a generic worry?
  • Would a senior engineer actually care?

Only keep challenges that survive self-refutation.

Output Format

## Plan Challenge Report

### Challenges That Survived Self-Refutation
| # | Dimension | Risk | Evidence | Severity |
|---|-----------|------|----------|----------|

### Refuted Challenges (for transparency)
| # | Challenge | Why Refuted |
|---|-----------|-------------|

### Recommendation
[PROCEED / REVISE PLAN / STOP AND RESEARCH]

### Suggested Plan Amendments
1. ...

Read the full file on GitHub · 90 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 90 lines · 46 tokens per session scan A 325de1b6947c

Subscribe to this mod's changes

plan-challenger is an agent published in the GitHub repository brain-bootstrap/claude-code-brain-bootstrap (11 stars, last pushed 4mo ago), licensed MIT. It adds 46 tokens to every session and 650 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.