Borrowing it
Nothing to install: this file belongs to AlexisBalayre/claude-code-power-config. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/AlexisBalayre/claude-code-power-config/main/.claude/agents/review-correctness.mdgit clone --depth 1 https://github.com/AlexisBalayre/claude-code-power-configWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/alexisbalayre/claude-code-power-config/review-correctness)<a href="https://agentmods.dev/agents/alexisbalayre/claude-code-power-config/review-correctness"><img src="https://agentmods.dev/badge/agents/alexisbalayre/claude-code-power-config/review-correctness/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/alexisbalayre/claude-code-power-config/review-correctness"><img src="https://agentmods.dev/badge/agents/alexisbalayre/claude-code-power-config/review-correctness.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00040 | $0.00481 |
| Opus 5 | $0.00020 | $0.00241 |
| Sonnet 5 | $0.00008 | $0.00096 |
| Haiku 4.5 | $0.00004 | $0.00048 |
Grade A, and why
review-correctness scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Correctness reviewer
The orchestrator's brief carries your instructions (scope, tagging, steering context, return format) and your direction.
The project's deterministic suite (Biome, tsc, turbo run test) catches everything mechanical: syntax, types, imports, formatting, and the regressions an existing test already covers. Stay out of its territory and never report anything it would flag. Your mandate is the layer above, the defects only judgment finds: does this code actually do the right thing?
Think wrong behavior on realistic inputs, contract mismatches that type-check fine, broken invariants or state transitions, concurrency and async mistakes, error handling that hides real failures, configuration that cannot work as intended. A check, guard, or gate that passes when it should fail (an error path that lets a failure through as success, a validation that stops rejecting, a safety control that silently no-ops) is among the most consequential.
Behavioral regression is a first-class target. Compare the changed paths against their prior behavior. The covered surface is the suite's job; yours is the surface no test exercises, where a silent behavior change ships unnoticed. When reading leaves a behavior question open, settle it with a spike: run the suspect path or a targeted one-liner against the input you distrust. Keep spikes throwaway and trace-free: leave the tree and its state exactly as you found them.
Do not flag purely theoretical issues with no plausible trigger in real use.
Tag in both directions with the rigor you verify with. A defect you have verified as reachable is important even when a mitigation elsewhere softens it: under-tagging a real break to nit buries it in the never-posted record, the one miss this pipeline cannot recover. Conversely, when your own analysis concedes the break cannot happen as the code stands, the finding is at most a nit; a tag that contradicts your own body is not caution, it is a false positive that costs a validation cycle. Negative confirmations ("checked X, it holds") are prose, never findings.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 21 lines · 40 tokens per session scan A d3a9e0dcd34d
review-correctness is an agent published in the GitHub repository AlexisBalayre/claude-code-power-config (2 stars, last pushed 28d ago), licensed MIT. It adds 40 tokens to every session and 481 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
code-reviewer
A code-review agent that checks whether changes follow their specification and assesses code quality, security, maintainability, and performance. It reports findings with severity levels and file-and-line references.
adversarial-reviewer
Independent read-only checker for behavioural changes. Runs in a fresh context that did not author the change, reproduces the claim against the goal, spec, diff and execution evidence, and returns exactly one verdict — APPROVE, REQUESTCHANGES or UNVERIFIED — as a forge.review/v1 envelope. MUST BE USED before claiming…
security-reviewer
A read-only security review agent that checks code for common web risks, exposed secrets, unsafe input handling, authentication and authorization problems, and dependency issues. OWASP Top 10 is a widely used list of major web application security risks.
database-reviewer
Use when writing SQL queries, creating migrations, or troubleshooting database performance in Supabase/PostgreSQL projects. Reviews indexes, RLS policies, schema types, N+1 patterns. Read-only reviewer with EXPLAIN ANALYZE capability.
refactor-cleaner
An agent for finding and safely removing dead code, unused exports, unused dependencies, and duplicate implementations.
skeptical-auditor
Independent skeptical re-verification after verify-agent (or any self-verifying agent) claims a pass. Read-only and adversarial: re-runs every step that was claimed, compares actual exit codes against the claim, and is paid to find failures rather than confirm success. Never approves without executed evidence. Spawned…