monitor

A command for checking whether a running experiment is still safe to continue before there is enough data to judge its result. An experiment compares different versions or treatments to measure their effects.

In plain words
What is it for?
Use it to check exposure pace, detect sample-ratio mismatches, and identify conditions that may require pausing or ending an experiment.
Why use it?
It separates early safety checks from the later question of whether the experiment succeeded, helping surface tracking or allocation problems during the run.

Command

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/mixpanel/ai-plugins/monitor
Clone the repo
git clone --depth 1 https://github.com/mixpanel/ai-plugins
Per session 0 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,431 The whole file, excluding the scripts and references it only reads on demand.
Security scan B 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.01431
Opus 5 $0.00000 $0.00715
Sonnet 5 $0.00000 $0.00286
Haiku 4.5 $0.00000 $0.00143

Measured 2d ago against content hash dbdff2d6ef86, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade B, and why

monitor scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Strips warnings and disclaimersmediumAnti-refusal

Omitting safety caveats hides risk from the user and is a common jailbreak preamble.

- Don't moralise about peeking — explain the math once, then route the user to safe signals.
plugins/mixpanel/skills/manage-experiment/commands/monitor.md · 107 lines

How it starts

The opening of the file, as written. The whole thing — 107 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Command: monitor

Mid-flight safety checks on a running experiment. This command answers "is it safe to keep this experiment running?" — distinct from interpret, which answers "did the experiment work?" Monitor is for the middle of the experiment, before there's enough signal to interpret. Peek only at what's safe to peek at; surface anything that warrants pause or termination.

The umbrella SKILL.md defines the shared glossary. Phase-specific terms below.


Glossary (monitor-specific)

  • Sample pace. The ratio of actual exposures accumulated to expected exposures at this point in the experiment's planned duration. A pace below 0.7 (≥30% slower than projected) suggests the experiment is underpowered relative to its design, or that something is wrong with exposure tracking.
  • Mid-flight SRM. A Sample Ratio Mismatch detected during the experiment, before exposures are mature. Distinct from the SRM check at interpretation time — mid-flight SRM is a bucketing-bug early-warning, not a verdict on the result.

The peeking trap and the peek-safety table (what's safe to look at mid-flight, what isn't) live in the umbrella's Cross-command policies — this command applies them, doesn't re-derive them.


Components (monitor-specific)

For the peek-safety table (what's safe to look at mid-flight, what isn't), see the umbrella's Cross-command policies. For the guardrail hard-gate threshold, same place.

Terminate-early decision rules

Three situations that justify ending a running experiment before its planned end:

  1. Trustworthiness failure. SRM fails mid-flight, or a misconfiguration is discovered that invalidates the design. Terminate, fix, restart. The accumulated exposures are not salvageable.
  2. Guardrail regression beyond the hard-gate threshold (defined in the umbrella). The guardrail regresses by more than the threshold, with a tight CI. Continuing exposes more users to a measurable harm. Terminate and route to interpret for the ship/iterate verdict.
  3. Sequential stopping boundary crossed (Sequential tests only). The platform's sequential boundary fires. This is the by-design early stop — terminate and route to interpret.

Read the full file on GitHub · 107 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 107 lines · 0 tokens per session scan B dbdff2d6ef86

Subscribe to this mod's changes

monitor is a command published in the GitHub repository mixpanel/ai-plugins (15 stars, last pushed 8d ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 1,431 tokens. A static security scan graded it B with 1 finding (strips warnings and disclaimers). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.