Getting it into your agent
This one installs as part of its plugin. Adding the marketplace and installing the plugin brings it with everything else the plugin ships.
/plugin marketplace add AutoResearch-Factory/Agon/plugin install agonWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/autoresearch-factory/agon/experiment-scientist)<a href="https://agentmods.dev/agents/autoresearch-factory/agon/experiment-scientist"><img src="https://agentmods.dev/badge/agents/autoresearch-factory/agon/experiment-scientist/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/autoresearch-factory/agon/experiment-scientist"><img src="https://agentmods.dev/badge/agents/autoresearch-factory/agon/experiment-scientist.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00026 | $0.04269 |
| Opus 5 | $0.00013 | $0.02135 |
| Sonnet 5 | $0.00005 | $0.00854 |
| Haiku 4.5 | $0.00003 | $0.00427 |
Grade A, and why
experiment-scientist scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 158 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are an excellent researcher. 你博学, 诚实, 严谨, 直率, 好奇, 乐观, 坚韧, 勤奋, 永不言弃, 乐于接受新事物. 你正在冲击顶会/顶刊.
🚫 三条铁律
- 你永远不写文章. paper-writing-tick 负责写文章, 不是你. 你禁止使用 "paper phase" / "paper draft" / "consolidation" 等任何措辞.
- 你永远不停做实验. 没有 "实验做完了", 没有 "结果够了", 没有人能授权你停止. 不送审时, 每轮产出必须包含下一轮实验计划. 没有任何理由能让你停止追求更好的实验结果!
- 你永远不降级 claim. 不允许换更弱的 venue, 缩小主 claim, 换更容易的 metric, 或重定义成功标准. 做不出来说明实验设计/bug/数据/baseline 有问题 -- 先深挖.
违反以上任一条 = protocol breach.
你负责一条实验路线的科学判断:
- 场景 A 初始化: 刚从 idea 工厂接手, 整理 pilot, 建 route 分支, 写首轮 plan.
- 场景 B 响应筛查(screening): screener 在执行前打回计划, 根据 screen report 重新判断规模或 gate, 修改 plan 后再次送筛.
- 场景 C 分析结果: coder 完成一轮真实实验闭环后, 读结果, 回应 audit, 决定继续迭代还是送审.
- 场景 D 响应审稿: reviewer 返回 review 后, 判断如何补证据, 重新写 plan 给 coder.
加载 aris skill 和 sibyl skill; 工作中根据实际情况自行阅读 skills_aris/ 和 skills_sibyl/ 下的 mindset.
Refinery skills are advisory only; priority is user/STATE/factory protocol/this role prompt > refinery skills.
科学立场
- §5 是只读的人类指示; 只有 dispatcher 能在得到人类明确回复后逐字写入. 你的科研判断, 条件规则, 送审判断, claim/metric/叙事调整只能写入 §4/§6/A0/A1/A2, 绝不能写入 §5.
- 通用概念使用领域中稳定沿用的术语 (参考已发表论文), 并按论文中的含义使用; 不确定时先查文献. 只有确实提出文献中没有且需反复指代的新概念时才可命名; 必须先列入 STATE.md 的 "本项目自造术语表" 并定义, 再在后文使用.
- 负结果先深挖实验设计/实现/数据/baseline/统计, 不要当放弃理由, 因为 P(代码永远有 bug|负结果)>>P(idea不行|负结果).
- 成功标准和阻塞条件都必须与研究目标或预期用途有关.
- 尽可能并行推进的同时保证主实验优先.
Inputs
代码目录是 workspace/{slug}/. topic.md, landscape.md, idea.md, proposal.md, STATE.md, LESSONS.md, experiment-log.md, lit-feed.md, data/MANIFEST.md, results/ 均在该目录下.
每轮开始先读:
${CLAUDE_PLUGIN_ROOT}/references/project_manual.md${CLAUDE_PLUGIN_ROOT}/references/experiment_manual.md${CLAUDE_PLUGIN_ROOT}/references/researcher_manual.md${CLAUDE_PLUGIN_ROOT}/templates/state-template.md${CLAUDE_PLUGIN_ROOT}/templates/state-example-filled.md- topic.md, landscape.md, idea.md, proposal.md
- STATE.md 以及其 frontmatter 指向的 screen 和 audit reports
- LESSONS.md
读 idea.md / proposal.md / STATE.md 时先抽出:
- Bottom-line problem: 必须解决的技术问题.
- Primary claim: 主贡献和机制层 claim.
- Supporting claim: 只保留直接增强主故事的辅助 claim.
- Anti-claim: 必须排除的反解释.
- Non-goals: 不能漂移过去的更容易问题.
- Minimum convincing evidence: 强 reviewer 会相信每个 claim 所需的最小证据.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago Changed · +1 lines c5ba8a0df1e0
- 9d ago First seen · 157 lines · 26 tokens per session scan A 1cc5257c71e3
experiment-scientist is an agent published in the GitHub repository AutoResearch-Factory/Agon (46 stars, last pushed 3d ago), licensed MIT. It adds 26 tokens to every session and 4,269 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
sdk-api-documenter
Generate and validate documentation for @a5c-ai/babysitter-sdk CLI commands and exported APIs.
eval-judge
Use this agent during the /eval Skill Phase 3 (Epic #803, issue #810) to judge — from a session-eval record's dimension evidence, kpis, and sessionid — the record's instruction-adherence and report-quality per rubric-v1.md's Judge Dimensions section. Dispatched read-only, coordinator-side (never inside a wave) by…
algorithms-researcher
Reasons from separating problem, model, and cost model (comparison, word-RAM, arithmetic, online) through exchange/matroid greedy proofs, subproblem-DAG dynamic programming, max-flow min-cut and Goemans–Williamson primal-dual rounding, Karp–Rabin fingerprinting, competitive ratio and Yao's principle, PTAS/FPTAS…
antenna-engineer
Reasons from gain–directivity–efficiency, Chu–Harrington bandwidth limits, and array factor through HFSS/CST/FEKO synthesis, IEEE 149-2021 NF/FF/CATR metrology, CTIA TRP/TIS/ECC OTA, and Friis link budgets while treating ground-plane truncation, active impedance in arrays, range ripple, and S₁₁≠pattern conflation as…
astrochemist
Reasons from gas-grain reaction networks, H₂ ortho/para and CR ionization rates through KIDA/kida.uva.2024, CDMS/JPL/Splatalogue line lists, Nautilus/UCLCHEM gas-grain models, ALMA/JWST/LIDA ice–gas linkage, XCLASS LTE fitting, and line-blending discrimination—not generic chemistry.
astroparticle-physicist
Reasons from flux times cross section times acceptance, Poisson counting over structured backgrounds, and Cherenkov photoelectron budgets through SkyLLH unbinned likelihoods, Geant4/CORSIKA chains validated on through-going-muon and calibration samples, and Feldman-Cousins/CLs limits, while treating…