evidence-preserving-research

evidence-preserving-research is a skill for Claude Code, Codex from AutoConference/AutoConference-skill. It costs 63 tokens per session (7,620 once invoked), scanned A, original, Apache-2.0.

A research workflow for running comparable experiments on a computational method and preserving the evidence behind the results. It tests individual parts of the method, checks where claims came from, and passes measured findings to paper writing.

In plain words
What is it for?
Use it to run a defined study, compare variations, investigate an open research question, decide which experiments to repeat, and prepare verified aggregates for a paper.
Why use it?
It reduces unsupported conclusions and makes it easier to tell whether a result came from the method, a changed component, or an experimental mistake.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to run a defined study, compare variations, investigate an open research question, decide which experiments to repeat, and prepare verified aggregates for a paper.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/autoconference/autoconference-skill/research
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add AutoConference/AutoConference-skill --skill research
Clone the repo
git clone --depth 1 https://github.com/AutoConference/AutoConference-skill

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for evidence-preserving-research

README.md
[![agentmods](https://agentmods.dev/badge/skills/autoconference/autoconference-skill/research/github.svg)](https://agentmods.dev/skills/autoconference/autoconference-skill/research)
Your own site
<a href="https://agentmods.dev/skills/autoconference/autoconference-skill/research"><img src="https://agentmods.dev/badge/skills/autoconference/autoconference-skill/research/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for evidence-preserving-research

Your own site · 80×15
<a href="https://agentmods.dev/skills/autoconference/autoconference-skill/research"><img src="https://agentmods.dev/badge/skills/autoconference/autoconference-skill/research.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 63 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 7,620 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00063 $0.07620
Opus 5.5 $0.00025 $0.03048
Sonnet 5 $0.00013 $0.01524
Haiku 4.5 $0.00006 $0.00762

Measured 3d ago against content hash 326868c21bae, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-28, from the pricing page.

Security

Grade A, and why

evidence-preserving-research scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

The scan reads SKILL.md. This mod also ships 7 executable files (scripts/check_calibration.py, scripts/check_design_validity.py, scripts/check_manuscript_readiness.py, …), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

research/SKILL.md · 599 lines

How it starts

The opening of the file, as written. The whole thing — 599 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Autonomous research on a computational method

You are a researcher. You have a machine, a published method, and a question its authors left open. Your job is to answer it with experiments, then write up what you found — including if what you found is "the idea did not work".

You run unattended. Nobody is going to answer a question you ask, so do not ask one. If something breaks, fix it and continue.

Start with the study brief

This skill is the discipline. It does not know which study you are on. That lives in STUDY.md at the root of the repository you are working in, and it is the first thing you read.

references/study-brief.md states what a brief must contain and why each part is load-bearing. In short, STUDY.md tells you:

The method what it is, in a paragraph, and the paper it comes from
The open question what you are investigating, and any entry points already known
The metric its name, whether lower or higher is better, whether it can be satisfied degenerately, whether it embeds a horizon, and whether it is constrained — nothing here can infer any of that
Frozen the evaluation semantics you may not touch, and why changing one makes every number before and after it incomparable
The method surface what you may change; this is also the module list your ablation-plan.tsv must cover
How to run one experiment the exact command, where results land, and how to read the metric without pulling a training log into your context
The budget how long one experiment may take before you kill it
Known breakages traps already hit and diagnosed, so you do not rediscover them

If STUDY.md is absent, stop and say so. Do not infer the study from the code and start running: a brief you wrote yourself is not a brief, and the frozen list in particular is not derivable from reading a repository — it is a decision about what a number in this venue means.

A worked instance is in references/study-bound-to-disagree.md; copy it to the study repository as STUDY.md if that is the study you are on.

Read the full file on GitHub · 599 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago Changed · -7 lines 326868c21bae
  2. 4d ago First seen · 606 lines · 63 tokens per session scan A 870197b027ea

Subscribe to this mod's changes

evidence-preserving-research is a skill published in the GitHub repository AutoConference/AutoConference-skill (6 stars, last pushed yesterday), licensed Apache-2.0. It adds 63 tokens to every session and 7,620 once invoked, about $0.0003 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-24.

Related

Other skills, from other repositories

instrument-data-to-allotrope

Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV. Use this skill when scientists need to standardize instrument data for LIMS systems, data lakes, or downstream analysis. Supports auto-detection of instrument types. Outputs include full…

anthropics/knowledge-work-plugins · 123 tokens

matlab

Build, review, migrate, and safely plan MATLAB or GNU Octave numerical workflows, including arrays, tabular/time data, tests, projects, graphics, MAT files, and explicit Python interoperability.

K-Dense-AI/scientific-agent-skills · 42 tokens

exploratory-data-analysis

Perform bounded, local exploratory analysis of explicitly supported scientific files. Use for redacted CSV/TSV/JSON profiles; optional NumPy, HDF5, FASTA/FASTQ, and basic image metadata inspection; missingness/leakage audits; outlier and transformation sensitivity; and rigorous EDA report scaffolds. Other domain…

K-Dense-AI/scientific-agent-skills · 83 tokens

phylogenetics

Build and analyze phylogenetic trees using MAFFT (multiple alignment), IQ-TREE 2 (maximum likelihood), and FastTree (fast NJ/ML). Visualize with ETE3 or FigTree. For evolutionary analysis, microbial genomics, viral phylodynamics, protein family analysis, and molecular clock studies.

K-Dense-AI/scientific-agent-skills · 68 tokens

research-engineer

An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.

davila7/claude-code-templates · 43 tokens

mapping-to-snomed

Maps clinical concept spans extracted by OpenMed to SNOMED CT concepts through a USER-SUPPLIED terminology server (the user's own Ontoserver, Snowstorm, or UMLS/UTS), never a bundled vocabulary. Use when the user wants to code findings, disorders, procedures, body structures, or substances to SNOMED CT, run an ECL…

maziyarpanahi/openmed · 205 tokens