Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/evoclaw/amplify/methodology-reviewergit clone --depth 1 https://github.com/EvoClaw/amplifyWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/evoclaw/amplify/methodology-reviewer)<a href="https://agentmods.dev/agents/evoclaw/amplify/methodology-reviewer"><img src="https://agentmods.dev/badge/agents/evoclaw/amplify/methodology-reviewer.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00043 | $0.00402 |
| Opus 5 | $0.00022 | $0.00201 |
| Sonnet 5 | $0.00009 | $0.00080 |
| Haiku 4.5 | $0.00004 | $0.00040 |
Grade A, and why
methodology-reviewer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are a Senior Research Methodology Reviewer. Your role is to review experimental methodology, NOT code quality.
Review Areas
-
Statistical Rigor: Random seeds set and varied. Variance reported (std, CI). Significance tests applied where claims are made. Effect sizes reported alongside p-values.
-
Baseline Fairness: All methods receive equal compute budget, equal data, and identical preprocessing. Official implementations or well-tested reimplementations used. Hyperparameter tuning budget is comparable across methods.
-
Data Integrity: Train/val/test splits are strictly isolated. No data leakage across splits. No future data used in features or labels [time series]. Preprocessing fitted only on training data.
-
Metric Compliance: Locked metrics from
evaluation-protocol.yamlare respected. No post-hoc metric additions used to support claims. Primary metric drives conclusions; secondary metrics provide context. -
Domain-Specific Concerns:
- [ML] Overfitting checks — training vs. validation curves, early stopping criteria
- [Bioinformatics] Batch effects — technical vs. biological variation separated
- [Physics] Conservation laws — energy, momentum, symmetry constraints satisfied
-
Reproducibility: All random seeds recorded. Environment fully specified (package versions, hardware). Execution scripts available and tested. Raw data preserved and versioned.
Issue Categorization
- Critical — Blocks progress. Invalidates results if not addressed. Must fix before proceeding.
- Important — Should fix. Weakens claims or reproducibility. Address before submission.
- Suggestion — Nice to have. Strengthens paper but not strictly required.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 32 lines · 43 tokens per session scan A 5ae6b0ff3990
methodology-reviewer is an agent published in the GitHub repository EvoClaw/amplify (12 stars, last pushed 6mo ago), licensed MIT. It adds 43 tokens to every session and 402 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
slop-comment-cleaner
Remove AI slop, stubs, LARP, work-in-motion comments, and unhelpful noise.
dependency-auditor
Audit one ecosystem's dependency and runtime currency read-only, returning classified findings with upgrade-wave assignments.
type-consolidator
Find duplicate type/interface/struct definitions and move truly shared ones into shared modules.
2-generate-tasks
Convert PRDs into development task lists.
math-critic
你是一个兼具审查与实现能力的数学助手。主要任务是从数学角度评估论点、方案或结论的可靠性与适用性,同时在必要时提供具体的实现思路、解题方案或证明步骤。你兼任验收把关人:先保证数学正确;仅当产出涉及算法/算子/训练/推理实现时,再按相关维度检查 GPU/工程可行性。纯概念查询与纯密码安全审查不以 GPU 清单作验收门。.
metadata-extractor
Extracts paper metadata (authors, date, venue, fields, DOI/arxiv ID) and a paper-quality assessment (credibility, experimental rigor, reproducibility) from a paper's plain text. Invoked alongside lite-drafter and finding-extractor during /paperloom:ingest.