paper-autoraters

paper-autoraters is a skill for Claude Code, Codex from woodfishhhh/EZ_math_model. It costs 121 tokens per session (1,773 once invoked), scanned A, a copy of paper-autoraters, MIT.

A set of four automated graders for research papers, based on PaperOrchestra, a research-paper generation system. They score citations and literature reviews, or compare two papers side by side.

In plain words
What is it for?
Scoring a generated paper against a reference paper, comparing two paper-writing systems, or checking whether a paper-orchestration workflow produces acceptable results.
Why use it?
It provides consistent checks for citation coverage, literature-review quality, and overall paper quality instead of relying only on informal judgment.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Scoring a generated paper against a reference paper, comparing two paper-writing systems, or checking whether a paper-orchestration workflow produces acceptable results.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/woodfishhhh/ez_math_model/paper-autoraters
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add woodfishhhh/EZ_math_model --skill paper-autoraters
Clone the repo
git clone --depth 1 https://github.com/woodfishhhh/EZ_math_model

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for paper-autoraters

README.md
[![agentmods](https://agentmods.dev/badge/skills/woodfishhhh/ez_math_model/paper-autoraters/github.svg)](https://agentmods.dev/skills/woodfishhhh/ez_math_model/paper-autoraters)
Your own site
<a href="https://agentmods.dev/skills/woodfishhhh/ez_math_model/paper-autoraters"><img src="https://agentmods.dev/badge/skills/woodfishhhh/ez_math_model/paper-autoraters/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for paper-autoraters

Your own site · 80×15
<a href="https://agentmods.dev/skills/woodfishhhh/ez_math_model/paper-autoraters"><img src="https://agentmods.dev/badge/skills/woodfishhhh/ez_math_model/paper-autoraters.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 121 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,773 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin 100% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00121 $0.01773
Opus 5 $0.00060 $0.00886
Sonnet 5 $0.00024 $0.00355
Haiku 4.5 $0.00012 $0.00177

Measured 11d ago against content hash 6f0f5604aa96, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

paper-autoraters scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (scripts/compute_f1.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

This is a copy

100% identical to paper-autoraters — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

skills/ez-math-model/external/paper-orchestra/skills/paper-autoraters/SKILL.md · 154 lines

How it starts

The opening of the file, as written. The whole thing — 154 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Paper Autoraters (App. F.3)

Faithful implementation of the four LLM-as-judge autoraters used in PaperOrchestra (Song et al., 2026, arXiv:2604.05018, §5 and App. F.3).

These are the metrics the paper uses to demonstrate that PaperOrchestra beats single-agent and AI-Scientist-v2 baselines. Use them to:

  1. Score a generated paper against a ground-truth paper.
  2. Compare two paper-writing pipelines side-by-side.
  3. Validate your own host-agent execution of the paper-orchestra pipeline.

The four autoraters

Autorater What it does Inputs Output
Citation F1 — P0/P1 partition Partitions reference list into P0 (must-cite) and P1 (good-to-cite) given the paper text one paper text + its references list JSON {ref_num: "P0"|"P1"}
Literature Review Quality 6-axis 0-100 score for Intro+Related Work, with anti-inflation hard caps one paper PDF/text + reference avg citation count JSON with axis_scores, penalties, summary, overall_score
SxS Overall Paper Quality Holistic side-by-side preference judgment two papers (PDF or text) JSON with winner ∈ {paper_1, paper_2, tie}
SxS Literature Review Quality Side-by-side preference, Intro+Related Work only two papers JSON with winner ∈ {paper_1, paper_2, tie}

The paper uses Gemini-3.1-Pro and GPT-5 as judges, set to temperature 0.0 (Gemini) or default 1.0 (GPT-5, which doesn't allow temperature adjustment). Use whatever your host LLM is.

Workflow

Citation F1 (compute Precision / Recall / F1 vs ground truth)

This is a two-step procedure:

Step 1: Partition the reference lists into P0 / P1

For both the ground-truth paper AND the generated paper, run the LLM with references/citation-f1-prompt.md:

inputs:
  paper_text:    full paper LaTeX or markdown
  references_str: numbered reference list (e.g., "1. Vaswani et al. (2017)
                  Attention Is All You Need. NeurIPS. 2. He et al. (2016)
                  Deep Residual Learning for Image Recognition. CVPR. ...")

output: JSON {"1": "P0", "2": "P1", "3": "P0", ...}

Read the full file on GitHub · 154 lines

Files

What ships with it

5 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 154 lines · 121 tokens per session scan A 6f0f5604aa96

Subscribe to this mod's changes

paper-autoraters is a skill published in the GitHub repository woodfishhhh/EZ_math_model (41 stars, last pushed 1mo ago), licensed MIT. It adds 121 tokens to every session and 1,773 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. It is 100% identical to paper-autoraters, differing in 0 lines, and is treated as a copy.