Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add matteotitta/genesys-skills --skill ab-testinggit clone --depth 1 https://github.com/matteotitta/genesys-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/matteotitta/genesys-skills/ab-testing)<a href="https://agentmods.dev/skills/matteotitta/genesys-skills/ab-testing"><img src="https://agentmods.dev/badge/skills/matteotitta/genesys-skills/ab-testing/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/matteotitta/genesys-skills/ab-testing"><img src="https://agentmods.dev/badge/skills/matteotitta/genesys-skills/ab-testing.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00155 | $0.02340 |
| Opus 5 | $0.00077 | $0.01170 |
| Sonnet 5 | $0.00031 | $0.00468 |
| Haiku 4.5 | $0.00015 | $0.00234 |
Grade A, and why
ab-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 214 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/ab-testing — statistically valid test design
Design A/B tests that produce trustworthy winners. The bar: pre-commit to sample size, hold the line on duration, only call results at 95% confidence.
Without this discipline, tests yield false-positive "winners" that don't replicate in production — and the team loses trust in the experimentation program within 2 quarters.
Doctrine inherited (Step 7 — 0626 rollout, locked 2026-06-04)
Output complies with output-tenets.md, output-simplicity.md. Step 6 calibration: see [[feedback_execution_doctrine_refinements_step6]].
Refinements applied: R1 (test brief is client-team review surface — cleaned cites in appendix), R3 (test-result reports operator-direct), R6 (variant CTA hierarchy explicit per stage), R9 (verb-led test-design section names).
When to invoke
- Landing-page test on hero copy / CTA / pricing variant.
- Signup-flow optimization (paired with
/signup-onboarding-auditfindings). - Email subject-line test in lifecycle sequence.
- In-product popup variant (per
/in-app-popupsA/B hypothesis output). - Pricing-page test on tier presentation / CTA copy.
Do NOT invoke when:
- Traffic / volume is too low for statistical power (see Step 3 — sample-size table).
- The hypothesis is "let's try this" with no observation-based reasoning.
- Multi-variable changes simultaneously — confounded; can't isolate cause.
- The decision is reversible at near-zero cost (just ship and observe).
Workflow
Step 1 — Hypothesis in canonical form
Use the format:
"Because {observation}, we believe {change} will cause {outcome} measured by {metric} within {timeframe}."
Example:
"Because 47% of users drop off at the company-size field, we believe removing the field will increase signup completion by 25% measured by signup-completion-rate within 2 weeks."
Each component is mandatory:
- Observation: evidence-based, not gut.
- Change: single-variable. Don't bundle.
- Outcome: direction + magnitude prediction.
- Metric: specific event name (per
/analytics-tracking-plan). - Timeframe: pre-committed; matches sample-size duration.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 214 lines · 155 tokens per session scan A ebf1da72a57f
ab-testing is a skill published in the GitHub repository matteotitta/genesys-skills (36 stars, last pushed 1mo ago), licensed MIT. It adds 155 tokens to every session and 2,340 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
gingiris-b2b-growth
🇺🇸 B2B SaaS Growth — PLG vs SLG Playbook — Diagnose whether your problem is distribution, pricing, or PMF. PLG/SLG selection by ACV and sales cycle, the 5-stage path from $0 to $10M ARR, NRR discipline, affiliate & channel motion, enterprise tiering. Built from HeyGen, Deel, Vercel, Supabase, Snowflake patterns.…
gr-b2b-growth
A guide to growing a business-to-business software product from early user research to large-scale sales. B2B software is sold to companies rather than individual consumers.
go-to-market-playbook
A reusable Go-to-Market strategy template for both B2B and B2C launches. Covers positioning, messaging, ICP definition, channel selection, and competitive analysis frameworks. By @WeiYipei.
gingiris-go-global
🇺🇸 AI Product / SaaS Go-Global Complete SOP — From competitor research to launch to monetization. A full-cycle playbook covering Phase 0-5 (market validation, positioning, first 100 users, user interviews, beta-to-growth) plus open-source launch, Product Hunt, Reddit, SEO/GEO, conversion, and org principles.…
gr-competitor-research
Your competitor just launched. You have no idea how they grew so fast. Should you reverse-engineer their website? Track their social media? Map their growth flywheel? This gives you the complete SOP — from Wayback Machine snapshots to X/Twitter propagation analysis to growth flywheel scoring. Built from 150+ AI…
ai-launch-playbook
Launch your AI product to global attention — the playbook behind Manus, Devin, and AFFiNE's breakout launches. Covers AI-specific GTM strategy, hype cycle management, waitlist tactics, and multi-market rollout for maximum day-one impact.