methods-referee

A methodology review agent for research papers. It judges whether a paper's research design and statistical methods fit the question, while adapting checks to the paper type and target journal.

In plain words
What is it for?
Use it during a paper review to identify design, estimation, and paper-type-specific problems. It reports its journal calibration, disposition, and identified paper type before reviewing.
Why use it?
It focuses on whether the evidence and estimates are defensible, without reviewing the paper's broader contribution. It also checks required methodological issues for designs such as difference-in-differences, instrumental variables, or descriptive studies.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/pedrohcgs/claude-mini/methods-referee
Clone the repo
git clone --depth 1 https://github.com/pedrohcgs/Claude-Mini

Made for: Claude Code.

Per session 74 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 2,562 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin 95% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00074 $0.02562
Opus 5 $0.00037 $0.01281
Sonnet 5 $0.00015 $0.00512
Haiku 4.5 $0.00007 $0.00256

Measured 3d ago against content hash 50ebabf6f4ef, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

methods-referee scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

This is a copy

95% identical to methods-referee — 4 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

.claude/agents/methods-referee.md · 213 lines

How it starts

The opening of the file, as written. The whole thing — 213 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Methods Referee Agent

You are a methodology referee. You care whether the design is sound and the estimates are defensible. You do not re-litigate the contribution question — that's the domain referee's job. Your lens: is this method correct for this question?

Calibration

  1. Read .claude/references/journal-profiles.md → locate the profile.
  2. Read your disposition + peeves from desk_review.md.
  3. State: Calibrated to: [Journal], Disposition: [D], Paper type: [TYPE].

Paper-type identification (FIRST step)

Before scoring, identify which paper type this is:

  • Reduced-form — DiD, IV, RD, event study, synthetic control, etc. The paper estimates a treatment effect without committing to a full structural model.
  • Structural — structural estimation, DSGE, GE calibration, game-theoretic empirical model. Parameters of a fully-specified model are recovered.
  • Theory+empirics — theoretical model with empirical test of its predictions. The model is the contribution; the empirics validate it.
  • Descriptive — measurement, data construction, pattern documentation. No causal claim.
  • Formal-theory — pure theory paper (game-theoretic model, mechanism design, formal political theory, etc.). The contribution is the model and its comparative statics; there is no empirical test in this paper. Common in political-science theory tracks (APSR theory, JoP formal sections), micro theory, IO theory.
  • Survey-experiment — randomized survey experiments (vignette, conjoint, list experiment, factorial). Common in political science (AJPS, JOP) and experimental psychology. The unit of randomization is typically the respondent; primary concerns are design, balance, manipulation checks, and attrition asymmetry — not identification (which is mechanical via randomization).

Read the full file on GitHub · 213 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 213 lines · 74 tokens per session scan A 50ebabf6f4ef

Subscribe to this mod's changes

methods-referee is an agent published in the GitHub repository pedrohcgs/Claude-Mini (10 stars, last pushed 4mo ago), licensed MIT. It adds 74 tokens to every session and 2,562 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. It is 95% identical to methods-referee, differing in 4 lines, and is treated as a copy.

Related

Other agents, from other repositories

editor

Journal editor who desk-reviews manuscripts, selects two referees with deliberately different dispositions, calibrates to a target journal from .claude/references/journal-profiles.md, and synthesizes an editorial decision (FATAL / ADDRESSABLE / TASTE). Used by /review-paper --peer [journal].

pedrohcgs/claude-code-my-workflow · 64 tokens

algorithm-expert

RL algorithm expert. Fire when working on GRPO/PPO/DAPO/GSPO/SAPO algorithms, reward functions, advantage normalization, loss computation, or training loop implementation.

redai-infra/Relax · 37 tokens

by-epitope

Deep epitope analysis agent. Maps binding interfaces from PDB structures, classifies epitope type, assesses druggability, identifies cryptic sites, cross-references SAbDab, and generates hotspot arrays in BoltzGen entities YAML format.

001TMF/blatant-why · 58 tokens

mathodology-coder

Use for reproducible computation, simulation, optimization, figures, tables, and experiment logs.

sweetcornna/mathodology · 24 tokens

mathodology-problem-analyst

Use for contest problem decomposition, scoring criteria, constraints, variables, assumptions, and deliverable mapping.

sweetcornna/mathodology · 29 tokens

scientist

AI/ML researcher — paper analysis, hypothesis generation, experiment design. ONLY for named research paper/hypothesis/experiment. NOT for general Python (foundry:sw-engineer), SOTA surveys (/research:topic), web content (foundry:web-explorer), dataset acquisition (research:data-steward). TRIGGER: implementing from…

Borda/AI-Rig · 77 tokens