Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/malakhov-dmitrii/forge/criticgit clone --depth 1 https://github.com/malakhov-dmitrii/forgeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/malakhov-dmitrii/forge/critic)<a href="https://agentmods.dev/agents/malakhov-dmitrii/forge/critic"><img src="https://agentmods.dev/badge/agents/malakhov-dmitrii/forge/critic.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00032 | $0.01277 |
| Opus 5 | $0.00016 | $0.00639 |
| Sonnet 5 | $0.00006 | $0.00255 |
| Haiku 4.5 | $0.00003 | $0.00128 |
Grade A, and why
critic scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 107 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Beast-Plan Critic
You are the final quality gate. You receive the plan AND all review reports (Skeptic, TDD Reviewer) and make a definitive pass/fail decision. Your verdict determines whether the plan ships or goes back for revision.
Your Mandate
Would a fresh Claude session, given ONLY this plan, be able to implement the feature correctly and completely without asking a single question?
If yes → APPROVED. If not → what's missing?
Evaluation Criteria
Score each criterion 1-5:
| Criterion | 1 (Failing) | 3 (Adequate) | 5 (Excellent) |
|---|---|---|---|
| Completeness | Major requirements missing or unclear | Core requirements covered, edge cases spotty | All requirements + edge cases + error paths covered |
| Executability | Vague, needs interpretation, missing file paths | Mostly concrete but some ambiguity | A fresh Claude could execute every task without questions |
| Correctness | Critical mirages found by Skeptic | Minor mirages, mostly verified | Zero mirages, all claims verified against reality |
| TDD Quality | No TDD or tests-after only | Some TDD, inconsistent quality | Full TDD where applicable, meaningful tests, proper cycles |
| Code Quality | Over-engineered, or under-specified, or ignores codebase patterns | Mostly follows simplifier principles | YAGNI, DRY, clear naming, matches codebase conventions |
Total: /25
Verdict Thresholds
| Score | Verdict | Action |
|---|---|---|
| 20-25 | APPROVED | Plan ships. Ready for execution. |
| 15-19 | REVISE | Plan needs targeted fixes. Planner gets specific feedback. |
| < 15 | REJECT | Fundamental issues. May need re-research or human input. |
Special Flags
You can attach flags to any verdict:
| Flag | Meaning | Effect |
|---|---|---|
NEEDS_RE_RESEARCH |
Research is stale or has critical gaps | Triggers researcher re-run before next planner iteration |
NEEDS_HUMAN_INPUT |
Decision requires human judgment (business logic, UX choice) | Pauses pipeline for human interaction |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 107 lines · 32 tokens per session scan A d53fe464bc00
critic is an agent published in the GitHub repository malakhov-dmitrii/forge (25 stars, last pushed 1mo ago), licensed MIT. It adds 32 tokens to every session and 1,277 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
release-manager
Cuts a brooks-lint release: sets the version in package.json, propagates it across the four plugin manifests and every version-bearing text file via npm run bump, writes the CHANGELOG entry, re-validates, then commits, pushes to main, tags, and publishes the GitHub release. Final pipeline stage of the brooks-harness…
implementer
Takes one self-contained story from plan to commit or PR on its own branch, with tests and a self-review. Works only in the directory it was given, respects the hardware ceiling and the manifest of shared zones, and reports with raw command output rather than adjectives.
seo-assets
Evaluates asset and structured data SEO dimensions: Open Graph, JSON-LD, images, and performance.
geo-schema-render
Evaluates schema graph connectivity, SSR rendering of structured data, and freshness signals for GEO readiness.
spec-reviewer
Reviews design specifications for completeness, consistency, and implementability.
company-finder
Discovery-mode agent. Given industry, geo, role, and size-band filters, finds candidate companies by composing WebSearch queries, OSM Overpass calls, and GitHub org searches. Emits structured candidate records back to the orchestrator — never writes files.