experiment-next

experiment-next is a command for Claude Code from with-geun/alive-analysis. It costs 0 tokens per session (3,479 once invoked), scanned A, original, MIT.

A command for advancing a structured experiment to its next stage. The experiment stages are design, validation, analysis, decision, and learning, with the current stage recorded in project files.

In plain words
What is it for?
Use it to move a full experiment forward, review its stage checklist, and identify missing or blocked work before continuing.
Why use it?
It removes the need to work out which experiment and stage are active or which file should come next. It also flags stop signals before advancing.

Command for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/with-geun/alive-analysis/experiment-next
Clone the repo
git clone --depth 1 https://github.com/with-geun/alive-analysis

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for experiment-next

README.md
[![agentmods](https://agentmods.dev/badge/commands/with-geun/alive-analysis/experiment-next.svg)](https://agentmods.dev/commands/with-geun/alive-analysis/experiment-next)
Your own site
<a href="https://agentmods.dev/commands/with-geun/alive-analysis/experiment-next"><img src="https://agentmods.dev/badge/commands/with-geun/alive-analysis/experiment-next.svg" alt="Measured on agentmods" height="20"></a>
Per session 0 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 3,479 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00000 $0.03479
Opus 5 $0.00000 $0.01740
Sonnet 5 $0.00000 $0.00696
Haiku 4.5 $0.00000 $0.00348

Measured 5d ago against content hash 3f52d09f4ed1, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

experiment-next scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/commands/experiment-next.md · 409 lines

How it starts

The opening of the file, as written. The whole thing — 409 lines — stays where its author put it; the contents beside it link to each section on GitHub.

/experiment next

Advance the current experiment to the next stage.

Instructions

Step 1: Identify current experiment

Read .analysis/status.md to find active experiments (Type: 🧪 Experiment).

  • If only 1 active Full experiment → select it automatically
  • If multiple active → ask user which experiment to advance (show ID + title)
  • Quick experiments don't use /experiment next (all sections are in one file)
    • If user selects a Quick, remind them to fill sections in order within the file

Step 2: Determine current stage

Read the experiment folder to see which stage files exist:

  • Only 01_design.md → current stage is DESIGN
  • Up to 02_validate.md → current stage is VALIDATE
  • Up to 03_analyze.md → current stage is ANALYZE
  • Up to 04_decide.md → current stage is DECIDE
  • All 5 files exist → current stage is LEARN (experiment is complete)

Step 3: Review current stage checklist

Read the current stage's checklist (embedded at the bottom of the current stage file).

Check if there are any 🔴 (stop) items:

  • If 🔴 items exist → warn the user: "There are stop signals in your {stage} checklist. Review before proceeding."
  • If all items are unchecked → remind: "Don't forget to review the checklist before moving on."
  • If items are checked → proceed

Special gate: DESIGN → VALIDATE Before advancing from DESIGN, verify:

  • Pre-registration section is filled (this MUST be locked before seeing any data)
  • If empty, warn: "Your analysis plan should be pre-registered before starting the experiment. Fill the Pre-registration section in 01_design.md first."

Step 4: Generate next stage file

Based on current stage, generate the next file.

DESIGN → VALIDATE (generate 02_validate.md):

# VALIDATE: {title}
> ID: {ID} | Type: 🧪 Experiment | Stage: ✅ VALIDATE | Updated: {YYYY-MM-DD}

## Pre-Experiment Checks

### AA Test / Historical Balance
- Method: (AA test on pre-period data / historical segment comparison)
- Result: Groups are balanced / not balanced
- Key metrics checked:
  | Metric | Control | Treatment | Δ | Balanced? |
  |--------|---------|-----------|---|-----------|
  | | | | | |

### Sample Ratio Mismatch (SRM)
- Expected ratio: {from 01_design.md}
- Actual ratio: {observed}
- Chi-square test p-value: (p < 0.001 indicates SRM)
- SRM detected? No / Yes → ⚠️ investigate before proceeding

> 💡 SRM indicates a problem with randomization. Common causes:
> - Bot filtering differences between variants
> - Redirect timing differences
> - Initialization bias (treatment takes longer to load)
> If SRM is detected, DO NOT proceed to Analyze. Fix the root cause first.

### Segment Balance
| Segment | Control % | Treatment % | Balanced? |
|---------|-----------|-------------|-----------|
| New users | | | |
| Returning users | | | |
| Platform (iOS/Android/Web) | | | |
| {custom segment} | | | |

### Instrumentation Check
- [ ] Primary metric events logging correctly
- [ ] Secondary metric events logging correctly
- [ ] Guardrail metric events logging correctly
- [ ] No duplicate counting
- [ ] Attribution window correctly set

## External Factors
- Holidays/events during experiment period:
- Planned releases/deployments:
- Marketing campaigns:
- Other active experiments on same population:

## Ramp-Up Log
| Date | Traffic % | Observation | Action |
|------|-----------|-------------|--------|
| {start} | {initial}% | | |

## Validation Summary
- Ready to proceed? ✅ Yes / 🔴 No — {reason}
- Issues found:
- Mitigation:

---
{Insert VALIDATE checklist from ab-tests/checklists/validate.md}

Read the full file on GitHub · 409 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 409 lines · 0 tokens per session scan A 3f52d09f4ed1

Subscribe to this mod's changes

experiment-next is a command published in the GitHub repository with-geun/alive-analysis (41 stars, last pushed 3mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 3,479 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.