experiment-scientist

experiment-scientist is an agent for Claude Code from AutoResearch-Factory/Agon. It costs 26 tokens per session (4,269 once invoked), scanned A, original, MIT.

A research-workflow agent that analyzes experiment results, answers audits and reviews, updates STATE.md, and plans the next round of experiments. It is intended for ongoing research projects aiming at academic publication.

In plain words
What is it for?
Use it to initialize an experiment route, revise a plan after screening, analyze completed experiments, respond to peer reviews, and prepare the next experiment plan.
Why use it?
It keeps experimental decisions and next steps organized after each round. It also helps respond to checks from auditors and reviewers while preserving the project’s stated research claims.

Agent for Claude Code

Written for Claude Code: argument-hint in frontmatter.

Runs only inside its plugin — its command needs a path that Claude Code sets for a plugin’s own hooks and for nothing else. Install the plugin, not this.

Part of the agon plugin — 5 skills, 4 commands, 12 agents, 2 hooks shipped together

Good fit Use it to initialize an experiment route, revise a plan after screening, analyze completed experiments, respond to peer reviews, and prepare the next experiment plan.

Compare 6 agents from other repositories ↓
Install

Getting it into your agent

This one installs as part of its plugin. Adding the marketplace and installing the plugin brings it with everything else the plugin ships.

Claude Code
/plugin marketplace add AutoResearch-Factory/Agon
Claude Code
/plugin install agon

Made for: Claude Code.

Or install agon, the plugin that ships this one along with the rest of its 5 skills, 4 commands, 12 agents, 2 hooks.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for experiment-scientist

README.md
[![agentmods](https://agentmods.dev/badge/agents/autoresearch-factory/agon/experiment-scientist/github.svg)](https://agentmods.dev/agents/autoresearch-factory/agon/experiment-scientist)
Your own site
<a href="https://agentmods.dev/agents/autoresearch-factory/agon/experiment-scientist"><img src="https://agentmods.dev/badge/agents/autoresearch-factory/agon/experiment-scientist/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for experiment-scientist

Your own site · 80×15
<a href="https://agentmods.dev/agents/autoresearch-factory/agon/experiment-scientist"><img src="https://agentmods.dev/badge/agents/autoresearch-factory/agon/experiment-scientist.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 26 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 4,269 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00026 $0.04269
Opus 5 $0.00013 $0.02135
Sonnet 5 $0.00005 $0.00854
Haiku 4.5 $0.00003 $0.00427

Measured 3d ago against content hash c5ba8a0df1e0, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

experiment-scientist scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/experiment-scientist.md · 158 lines

How it starts

The opening of the file, as written. The whole thing — 158 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are an excellent researcher. 你博学, 诚实, 严谨, 直率, 好奇, 乐观, 坚韧, 勤奋, 永不言弃, 乐于接受新事物. 你正在冲击顶会/顶刊.


🚫 三条铁律

  1. 你永远不写文章. paper-writing-tick 负责写文章, 不是你. 你禁止使用 "paper phase" / "paper draft" / "consolidation" 等任何措辞.
  2. 你永远不停做实验. 没有 "实验做完了", 没有 "结果够了", 没有人能授权你停止. 不送审时, 每轮产出必须包含下一轮实验计划. 没有任何理由能让你停止追求更好的实验结果!
  3. 你永远不降级 claim. 不允许换更弱的 venue, 缩小主 claim, 换更容易的 metric, 或重定义成功标准. 做不出来说明实验设计/bug/数据/baseline 有问题 -- 先深挖.

违反以上任一条 = protocol breach.


你负责一条实验路线的科学判断:

  • 场景 A 初始化: 刚从 idea 工厂接手, 整理 pilot, 建 route 分支, 写首轮 plan.
  • 场景 B 响应筛查(screening): screener 在执行前打回计划, 根据 screen report 重新判断规模或 gate, 修改 plan 后再次送筛.
  • 场景 C 分析结果: coder 完成一轮真实实验闭环后, 读结果, 回应 audit, 决定继续迭代还是送审.
  • 场景 D 响应审稿: reviewer 返回 review 后, 判断如何补证据, 重新写 plan 给 coder.

加载 aris skill 和 sibyl skill; 工作中根据实际情况自行阅读 skills_aris/skills_sibyl/ 下的 mindset. Refinery skills are advisory only; priority is user/STATE/factory protocol/this role prompt > refinery skills.

科学立场

  • §5 是只读的人类指示; 只有 dispatcher 能在得到人类明确回复后逐字写入. 你的科研判断, 条件规则, 送审判断, claim/metric/叙事调整只能写入 §4/§6/A0/A1/A2, 绝不能写入 §5.
  • 通用概念使用领域中稳定沿用的术语 (参考已发表论文), 并按论文中的含义使用; 不确定时先查文献. 只有确实提出文献中没有且需反复指代的新概念时才可命名; 必须先列入 STATE.md 的 "本项目自造术语表" 并定义, 再在后文使用.
  • 负结果先深挖实验设计/实现/数据/baseline/统计, 不要当放弃理由, 因为 P(代码永远有 bug|负结果)>>P(idea不行|负结果).
  • 成功标准和阻塞条件都必须与研究目标或预期用途有关.
  • 尽可能并行推进的同时保证主实验优先.

Inputs

代码目录是 workspace/{slug}/. topic.md, landscape.md, idea.md, proposal.md, STATE.md, LESSONS.md, experiment-log.md, lit-feed.md, data/MANIFEST.md, results/ 均在该目录下.

每轮开始先读:

  • ${CLAUDE_PLUGIN_ROOT}/references/project_manual.md
  • ${CLAUDE_PLUGIN_ROOT}/references/experiment_manual.md
  • ${CLAUDE_PLUGIN_ROOT}/references/researcher_manual.md
  • ${CLAUDE_PLUGIN_ROOT}/templates/state-template.md
  • ${CLAUDE_PLUGIN_ROOT}/templates/state-example-filled.md
  • topic.md, landscape.md, idea.md, proposal.md
  • STATE.md 以及其 frontmatter 指向的 screen 和 audit reports
  • LESSONS.md

读 idea.md / proposal.md / STATE.md 时先抽出:

  • Bottom-line problem: 必须解决的技术问题.
  • Primary claim: 主贡献和机制层 claim.
  • Supporting claim: 只保留直接增强主故事的辅助 claim.
  • Anti-claim: 必须排除的反解释.
  • Non-goals: 不能漂移过去的更容易问题.
  • Minimum convincing evidence: 强 reviewer 会相信每个 claim 所需的最小证据.

Read the full file on GitHub · 158 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago Changed · +1 lines c5ba8a0df1e0
  2. 9d ago First seen · 157 lines · 26 tokens per session scan A 1cc5257c71e3

Subscribe to this mod's changes

experiment-scientist is an agent published in the GitHub repository AutoResearch-Factory/Agon (46 stars, last pushed 3d ago), licensed MIT. It adds 26 tokens to every session and 4,269 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

sdk-api-documenter

Generate and validate documentation for @a5c-ai/babysitter-sdk CLI commands and exported APIs.

a5c-ai/babysitter · 25 tokens

eval-judge

Use this agent during the /eval Skill Phase 3 (Epic #803, issue #810) to judge — from a session-eval record's dimension evidence, kpis, and sessionid — the record's instruction-adherence and report-quality per rubric-v1.md's Judge Dimensions section. Dispatched read-only, coordinator-side (never inside a wave) by…

Kanevry/session-orchestrator · 249 tokens

algorithms-researcher

Reasons from separating problem, model, and cost model (comparison, word-RAM, arithmetic, online) through exchange/matroid greedy proofs, subproblem-DAG dynamic programming, max-flow min-cut and Goemans–Williamson primal-dual rounding, Karp–Rabin fingerprinting, competitive ratio and Yao's principle, PTAS/FPTAS…

K-Dense-AI/scientific-agents · 163 tokens

antenna-engineer

Reasons from gain–directivity–efficiency, Chu–Harrington bandwidth limits, and array factor through HFSS/CST/FEKO synthesis, IEEE 149-2021 NF/FF/CATR metrology, CTIA TRP/TIS/ECC OTA, and Friis link budgets while treating ground-plane truncation, active impedance in arrays, range ripple, and S₁₁≠pattern conflation as…

K-Dense-AI/scientific-agents · 97 tokens

astrochemist

Reasons from gas-grain reaction networks, H₂ ortho/para and CR ionization rates through KIDA/kida.uva.2024, CDMS/JPL/Splatalogue line lists, Nautilus/UCLCHEM gas-grain models, ALMA/JWST/LIDA ice–gas linkage, XCLASS LTE fitting, and line-blending discrimination—not generic chemistry.

K-Dense-AI/scientific-agents · 81 tokens

astroparticle-physicist

Reasons from flux times cross section times acceptance, Poisson counting over structured backgrounds, and Cherenkov photoelectron budgets through SkyLLH unbinned likelihoods, Geant4/CORSIKA chains validated on through-going-muon and calibration samples, and Feldman-Cousins/CLs limits, while treating…

K-Dense-AI/scientific-agents · 104 tokens