ab-test

An experiment-planning command for comparing two versions of a page or user flow. It defines the hypothesis, success measures, required sample size, test length, and analysis plan.

In plain words
What is it for?
Use it to plan tests of page designs, layouts, text, pricing pages, signup flows, checkout steps, or calls to action.
Why use it?
It prevents teams from running tests without knowing what result would count as success or whether enough users will take part. This helps avoid spending traffic and development time on unreliable conclusions.

Command

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/brainbytes-dev/everything-claude-marketing/ab-test
Clone the repo
git clone --depth 1 https://github.com/brainbytes-dev/everything-claude-marketing
Per session 19 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,667 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00019 $0.01667
Opus 5 $0.00010 $0.00834
Sonnet 5 $0.00004 $0.00333
Haiku 4.5 $0.00002 $0.00167

Measured 2d ago against content hash 87e10ffdbdeb, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

ab-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

commands/ab-test.md · 194 lines

How it starts

The opening of the file, as written. The whole thing — 194 lines — stays where its author put it; the contents beside it link to each section on GitHub.

/ab-test

Design rigorous A/B tests with structured hypotheses, defined success metrics, calculated sample size requirements, estimated test durations, and analysis plans.

What This Command Does

The /ab-test command takes your testing idea and transforms it into a methodologically sound experiment design. It ensures your test has a clear hypothesis, measurable outcomes, sufficient statistical power, and a defined analysis plan before you invest development and traffic resources. The output is a complete test brief that your engineering, design, and analytics teams can execute against.

The command delegates to the cro-specialist agent, which applies conversion rate optimization expertise and statistical testing methodology to design experiments that produce reliable, actionable results.

When to Use

  • You want to test a new page design, layout, or content variation against the current version
  • You are optimizing conversion rates on landing pages, signup flows, or checkout processes
  • You need to validate a hypothesis about user behavior before committing to a full redesign
  • You want to test pricing page layouts, CTA copy, or value proposition framing
  • You are running email subject line tests and need proper methodology
  • You want to test changes to onboarding flows or feature adoption prompts
  • You need to justify a proposed change to stakeholders with data

How It Works

  1. Hypothesis Formulation — Structures your test idea into a formal hypothesis with independent variable, dependent variable, and expected outcome
  2. Metric Definition — Defines primary and secondary success metrics with clear measurement methods
  3. Variant Design — Specifies exactly what changes between control and treatment, and why those changes are expected to impact the metric
  4. Sample Size Calculation — Determines how many visitors or users you need for statistically significant results at your desired confidence level
  5. Duration Estimation — Estimates how long the test needs to run based on your traffic volume and required sample size
  6. Segmentation Plan — Identifies audience segments to analyze separately for heterogeneous treatment effects
  7. Analysis Plan — Defines how results will be evaluated, including what constitutes a winner and when to stop the test

Read the full file on GitHub · 194 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 194 lines · 19 tokens per session scan A 87e10ffdbdeb

Subscribe to this mod's changes

ab-test is a command published in the GitHub repository brainbytes-dev/everything-claude-marketing (5 stars, last pushed 5mo ago), licensed MIT. It adds 19 tokens to every session and 1,667 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other commands, from other repositories

feature

Post a feature to the Features Board on production and write tweets.

samuelclay/NewsBlur · 12 tokens

seo-geo

SEO/GEO end-to-end along the SITE loop: survey demand and competitors, implement content, tune quality/tech/on-page, and evaluate authority/rankings/reports/memory (--phase survey|implement|tune|evaluate). Not sure? Use /aaron-marketing:auto.

aaron-he-zhu/aaron-marketing-skills · 58 tokens

email-audit

Full-spectrum email audit — technical rendering issues (Phase 1, Email Designer) and copy/strategy critique (Phase 2, Email Copywriter). Accepts pasted HTML or a description of an existing email.

Adityaraj0421/naksha-studio · 46 tokens

competitive-landscape-analyzer

Comprehensive competitive analysis combining deep research with ad platform data. Includes competitor profiling, positioning maps, spend estimation, creative analysis, and strategic recommendations. Integrates Google Ads auction insights and Meta Ad Library research. (requires Pro subscription).

Ad-Superpowers/ad-superpowers-plugin · 46 tokens

review-video

Make a " Reviews" video — a fast, faceless VO montage of REAL, verified competitor reviews that names the recurring complaints and positions YOUR business as the alternative, then hands off to your own customer testimonials.

Bomx/super-video-maker-skill · 45 tokens

client-report

Produce the full client-facing first-scan pitch report for a prospect, end to end — live scan, research pass, branded PDF.

prashishh/seo-geo-report-engine · 27 tokens