experiment-designer

experiment-designer is an agent for coding agents from ai-analyst-lab/ai-analyst-plugin. It costs 31 tokens per session (3,783 once invoked), scanned A, a copy of experiment-designer, MIT.

An experiment-design agent that plans tests or comparison-based analyses for causal questions. It also estimates how much data is needed and defines decisions before results are seen.

In plain words
What is it for?
Use it to design randomized experiments or quasi-experiments, estimate statistical power, choose safety metrics, and prepare a pre-registered analysis plan.
Why use it?
It turns a vague idea into a testable plan with a metric, expected direction, guardrails, feasibility assessment, and rules for interpreting outcomes.

Agent

Part of the ai-analyst-plus plugin — 44 skills, 1 command, 13 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/ai-analyst-lab/ai-analyst-plugin/experiment-designer
Clone the repo
git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst-plugin

Or install ai-analyst-plus, the plugin that ships this one along with the rest of its 44 skills, 1 command, 13 agents.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for experiment-designer

README.md
[![agentmods](https://agentmods.dev/badge/agents/ai-analyst-lab/ai-analyst-plugin/experiment-designer.svg)](https://agentmods.dev/agents/ai-analyst-lab/ai-analyst-plugin/experiment-designer)
Your own site
<a href="https://agentmods.dev/agents/ai-analyst-lab/ai-analyst-plugin/experiment-designer"><img src="https://agentmods.dev/badge/agents/ai-analyst-lab/ai-analyst-plugin/experiment-designer.svg" alt="Measured on agentmods" height="20"></a>
Per session 31 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 3,783 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin 91% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00031 $0.03783
Opus 5 $0.00015 $0.01892
Sonnet 5 $0.00006 $0.00757
Haiku 4.5 $0.00003 $0.00378

Measured 5d ago against content hash 20bd55db02ad, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

experiment-designer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

This is a copy

91% identical to experiment-designer — 88 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

ai-analyst-plus/agents/experiment-designer.md · 357 lines

How it starts

The opening of the file, as written. The whole thing — 357 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Agent: Experiment Designer

Purpose

Design experiments or quasi-experimental analyses to test causal hypotheses. Handles the full path from feasibility assessment through test design, power estimation, guardrail selection, and pre-registration of decision rules — so the team knows what they'll do with every possible outcome before seeing results.

Inputs

  • {{HYPOTHESIS}}: The testable hypothesis to evaluate (from the Hypothesis agent or user). Must include the specific metric, expected direction, and mechanism. If vague ("Feature X improves retention"), prompt the user to specify a metric, threshold, and time window.
  • {{DATASET}}: Data source for computing baseline metrics, variance, and sample sizes needed for power estimation.
  • {{CONSTRAINTS}}: What type of experiment is feasible? One of:
    • full_ab — can randomize users into treatment and control
    • limited_traffic — can randomize but traffic/sample is small
    • no_randomization — already shipped or can't randomize, but have a comparison group
    • post_hoc — already shipped, no comparison group, need observational analysis
    • unknown — the agent will assess feasibility in Step 1

Workflow

Step 1: Assess Feasibility

Determine which experimental path is appropriate.

1a. Feasibility decision tree:

Can we randomize users?
├── YES: Is traffic sufficient for statistical power?
│   ├── YES → Full A/B test (Step 2)
│   └── NO  → Limited-traffic design (Step 2, with adjustments)
└── NO: Has the change already shipped?
    ├── NO: Do we have a natural comparison group?
    │   ├── YES → Diff-in-diff design (Step 3)
    │   └── NO  → Pre-post design (Step 3)
    └── YES: Is there a natural comparison group?
        ├── YES → Diff-in-diff or matching (Step 3)
        └── NO  → Pre-post with caveats (Step 3)

1b. If {{CONSTRAINTS}} is unknown, determine feasibility by asking:

  • Is the change something we can gate by user ID or session? (→ randomization possible)
  • What is the current traffic/user volume for the affected flow? (→ power feasibility)
  • Has the change already been shipped? (→ post-hoc only)
  • Is there a group that was NOT affected? (→ comparison group exists)

Read the full file on GitHub · 357 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 357 lines · 31 tokens per session scan A 20bd55db02ad

Subscribe to this mod's changes

experiment-designer is an agent published in the GitHub repository ai-analyst-lab/ai-analyst-plugin (32 stars, last pushed 9d ago), licensed MIT. It adds 31 tokens to every session and 3,783 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. It is 91% identical to experiment-designer, differing in 88 lines, and is treated as a copy.

Related

Other agents, from other repositories

editor

Journal editor who desk-reviews manuscripts, selects two referees with deliberately different dispositions, calibrates to a target journal from .claude/references/journal-profiles.md, and synthesizes an editorial decision (FATAL / ADDRESSABLE / TASTE). Used by /review-paper --peer [journal].

pedrohcgs/claude-code-my-workflow · 64 tokens

Geoprocessing Specialist

ArcPy and Python toolbox expert who automates spatial workflows — builds .pyt toolboxes, Model Builder processes, batch geoprocessing automation, and custom analysis scripts for ArcGIS Pro.

SHAdd0WTAka/Zen-Ai-Pentest · 45 tokens

research-scout

Scans the NeqSim codebase to discover scientific paper opportunities that will drive code improvement. Every paper must improve NeqSim — adding tests, validating models against data, hardening algorithms, or implementing new capabilities. Produces ranked, actionable topics that feed into the planner agent.

equinor/neqsim · 61 tokens

algorithm-expert

RL algorithm expert. Fire when working on GRPO/PPO/DAPO/GSPO/SAPO algorithms, reward functions, advantage normalization, loss computation, or training loop implementation.

redai-infra/Relax · 37 tokens

mathodology-problem-analyst

Use for contest problem decomposition, scoring criteria, constraints, variables, assumptions, and deliverable mapping.

sweetcornna/mathodology · 29 tokens

astronomical-instrumentation-scientist

Reasons from system-level error budgets, the diffraction limit and Strehl ratio, detector figures of merit, and resolving power through Zemax/Code V tolerancing, ETC radiometry, AO modeling, and on-sky standard-star commissioning while treating flexure drift, IR persistence, ghosts, and quasi-static speckles as…

K-Dense-AI/scientific-agents · 78 tokens