wanshuiyin/Anti-Autoresearch

Don't trust an autoresearch paper at face value. Reviewer-side integrity forensics (self-consistency + fabrication), deterministic verdict. 61 signals: 46 integrity hack-patterns (families A–H, verdict-bearing) + 13 zero-weight AI writing-style impressions (AIS) + 2 advisory. Not an opaque AI-text classifier. The dual of ARIS.

153Stars on the repository
12Mods indexed here, across every type
3d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Synthesize the single strongest EVIDENCE-BOUND reviewer case to reject a paper, built ONLY from the evidence ledger (claims.json) + the other auditors' confirmed findings — never free-floating LLM critique. Two fresh cross-model codex threads: an attack writes the 200-word rejection paragraph (every accusation tagged…

not rated 153 +4 3d ago A SkillSpector: pass 194 tokens original MIT

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Transparent, itemized impressions of AI-generated WRITING STYLE — the repo's ONLY non-integrity track. Two passes: a deterministic defensive-hedge density screen (tools/checkaistyle.py, AIS-DEFENSIVE-HEDGE) plus a fresh cross-model GROSS-cases-only semantic pass over the 13 AIS- style tells (broken narrative arc, LLM…

not rated 153 +4 3d ago A SkillSpector: pass 284 tokens original MIT

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Audit whether a paper's baseline comparisons are COMPLETE, FAIR, and SIGNIFICANT: a required recent SOTA baseline is missing while 'best/SOTA' is claimed (HP-MISSING-BASELINE); a baseline is undertuned / given less compute-tuning-data, run at a mismatched config, or the equal-budget ablation-as-baseline is absent…

not rated 153 +4 3d ago A SkillSpector: warn 310 tokens original MIT

citation-forensics

04

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Citation-integrity forensics: is every reference real, correctly attributed, and used in a context the cited work actually supports? Catches hallucinated references (no paper at the claimed arXiv id/DOI/venue, fabricated authors/year), metadata drift (wrong year/venue/version), and wrong-context citations (a real…

not rated 153 +4 3d ago A SkillSpector: warn 193 tokens original MIT

consistency-audit

05

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Flagship intra-paper self-consistency forensics: does the paper contradict ITSELF across abstract/intro/tables/body/appendix, and does the method DESCRIBED match the method EVALUATED? Needs no external ground truth — works PDF-only (L0). Runs a deterministic arithmetic pass + a fresh cross-model semantic pass, every…

not rated 153 +4 3d ago A SkillSpector: warn 133 tokens original MIT

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Audit whether a paper's EVALUATION DESIGN actually measures what it claims and whether its reporting is complete — the validity layer family D (experiment-forensics) cannot reach. Three patterns: train/test leakage means the reported score may not measure generalization (HP-EVAL-LEAKAGE — adopts the Kapoor & Narayanan…

not rated 153 +4 3d ago A SkillSpector: warn 458 tokens original MIT

evidence-ledger

07

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Build the deterministic evidence ledger (artifactmanifest.json + claims.json) that every other Anti-Autoresearch auditor reads. One pass inventories artifacts, derives the observability level (L0 PDF-only / L1 +LaTeX / L2 +repo+results) by fixed rule, and extracts span-anchored, hashed, checkable claims (numbers…

not rated 153 +4 3d ago A SkillSpector: warn 221 tokens original MIT

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Audit experiment integrity against the evidence ledger. At L2 (repo + result files present) a fresh cross-model reviewer reads the eval code line-by-line for fake/derived ground truth, score self-normalization, phantom results (a paper number with no backing file/key), dead/uncalled metric code, verified-scope…

not rated 153 +4 3d ago A SkillSpector: warn 237 tokens original MIT

wanshuiyin/Anti-Autoresearch

Skill Claude Code

MEMO-ONLY prior-work overlap advisory: surfaces the two ADVISORY taxonomy signals neither a tool nor a model can decide from the paper alone — ADV-TRIVIAL-COMBINATION (standard A+B+C / 缝合 stapling) and ADV-DUPLICATE-PUBLICATION (repackaged / duplicate submission). The executor RETRIEVES candidate prior work (DBLP…

not rated 153 +4 3d ago A SkillSpector: pass 284 tokens original MIT

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Checkable-ish surface presentation signals a reviewer notices first — duplicate/near-identical tables, leftover pipeline/template strings, too-few or LLM-looking figures, and page-padding. AUXILIARY ONLY and weak by design: a deterministic pass (tools/checkpresentation.py — dup-table + pipeline-artifact) plus a fresh…

not rated 153 +4 3d ago A SkillSpector: warn 246 tokens original MIT

wanshuiyin/Anti-Autoresearch

Skill Claude Code

Family-G proof & derivation integrity forensics: does a THIRD PARTY's written proof/derivation actually establish its theorem, or does it skip an obligation, assume its own conclusion, take an invalid step, drift a symbol's meaning, or smuggle an unstated assumption? Decides from the WRITTEN proof/derivation …

not rated 153 +4 3d ago A SkillSpector: warn 222 tokens original MIT

anti-autoresearch

12

wanshuiyin/Anti-Autoresearch

Skill Claude Code

End-to-end substantive-integrity forensic sweep of a research paper (especially autoresearch / AI-Scientist-style output). Orchestrates the whole pipeline: ingest (arxiv-id | pdf | dir → working dir + pdftotext for L0) → /evidence-ledger (artifact manifest + observability level L0/L1/L2 + span-anchored claims.json) →…

not rated 153 +4 3d ago A SkillSpector: warn 275 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: