Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Uxcel-Lab/product-skills --skill experimentation-abgit clone --depth 1 https://github.com/Uxcel-Lab/product-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/uxcel-lab/product-skills/experimentation-ab)<a href="https://agentmods.dev/skills/uxcel-lab/product-skills/experimentation-ab"><img src="https://agentmods.dev/badge/skills/uxcel-lab/product-skills/experimentation-ab/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/uxcel-lab/product-skills/experimentation-ab"><img src="https://agentmods.dev/badge/skills/uxcel-lab/product-skills/experimentation-ab.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00134 | $0.02935 |
| Opus 5 | $0.00067 | $0.01468 |
| Sonnet 5 | $0.00027 | $0.00587 |
| Haiku 4.5 | $0.00013 | $0.00294 |
Grade A, and why
pm-experimentation-ab scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 120 lines — stays where its author put it; the contents beside it link to each section on GitHub.
A/B Testing & Experimentation Skill
How this skill behaves (read first)
This is a generative skill (it designs or revises an experiment plan, and can review an existing test setup). Experimentation is where an AI assistant's defaults are quietly wrong in ways that produce confident, false conclusions. Claude will happily "set up an A/B test" that declares a winner on a handful of users, peeks and stops the moment p < 0.05, optimizes one metric with no guardrail for the damage it does elsewhere, runs for three weekdays, and ends at "B won" with no segments, no long-term read, and nothing written down. A bad experiment is worse than none — it launders a guess into "data."
So this skill gates, and the gate is statistical discipline:
- Establish context — is a controlled experiment even the right tool here, what single decision it informs, and the risk level (these set confidence, split, and duration).
- Apply the always-true core — pre-register the hypothesis and success criteria, size the test before running it, run a full cycle, use one primary metric + guardrails, interpret past "it won," and close the learning loop.
- Surface the context-dependent decisions (confidence threshold, traffic split, A/B vs. multivariate, metric type, leading/long-term indicators, test prioritization) with trade-offs; let the user choose.
Then it hands off to pm-okr-metric-validity-audit (are the chosen metrics valid, not vanity?) and pm-assumption-rigor-audit (is the hypothesis the riskiest, falsifiable, pre-committed?).
Scope & pairing: this skill owns the controlled-experiment method and its statistics. It pairs with pm-assumption-testing, which owns choosing what to test and the cheapest way to test it — when the question is early-stage, low-traffic, or better answered qualitatively, defer to assumption-testing rather than forcing an A/B test. It defers broad what-to-build-next ranking to pm-prioritization, metric definitions/OKRs to pm-okrs-kpis, and the high-stakes roll-out go/no-go to pm-decision-quality-audit.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 120 lines · 134 tokens per session scan A 1441f18209ca
pm-experimentation-ab is a skill published in the GitHub repository Uxcel-Lab/product-skills (12 stars, last pushed 2mo ago), licensed MIT. It adds 134 tokens to every session and 2,935 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
apple-container
Apple's open-source container CLI to build, run, and manage OCI/Linux containers as lightweight per-container VMs on Apple-silicon macOS — no Docker daemon required. Use when the user mentions the container CLI, "apple container", running or building containers on macOS without Docker/Podman, container run, container…
atlassian
Manage Jira issues and Confluence wiki pages in Atlassian Cloud. Use when: (1) searching/creating/updating Jira issues with JQL, (2) searching/reading/creating Confluence pages with CQL, (3) managing Jira workflows, transitions, and comments, (4) browsing Confluence spaces and page hierarchies. Supports OAuth 2.1 via…
elevenlabs
Convert documents and text to audio using ElevenLabs text-to-speech. Use this skill when the user wants to create a podcast, narrate a document, read aloud text, generate audio from a file, or convert text to speech.
google-drive
Interact with Google Drive - search files, find folders, list contents, download files, upload files, create folders, move, copy, rename, and trash files. Use when user asks to: search Google Drive, find a file/folder, list Drive contents, download or upload files, create folders, move files, or organize Drive…
playwright-cli
Automates browser interactions for testing and validating your own web applications using playwright-cli. Use when you need terminal-first browser control for navigation, form filling, screenshots, tracing, bound browser sessions, debugging, or generating Playwright test code. Only use against applications you own or…
manus
Delegate complex, long-running tasks to Manus AI agent for autonomous execution. Use when user says 'use manus', 'delegate to manus', 'send to manus', 'have manus do', 'ask manus', 'check manus sessions', or when tasks require deep web research, market analysis, product comparisons, stock analysis, competitive…