Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Sudhakaran88/solopreneur-skills --skill ab-test-setupgit clone --depth 1 https://github.com/Sudhakaran88/solopreneur-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/sudhakaran88/solopreneur-skills/ab-test-setup)<a href="https://agentmods.dev/skills/sudhakaran88/solopreneur-skills/ab-test-setup"><img src="https://agentmods.dev/badge/skills/sudhakaran88/solopreneur-skills/ab-test-setup/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/sudhakaran88/solopreneur-skills/ab-test-setup"><img src="https://agentmods.dev/badge/skills/sudhakaran88/solopreneur-skills/ab-test-setup.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00079 | $0.03890 |
| Opus 5 | $0.00039 | $0.01945 |
| Sonnet 5 | $0.00016 | $0.00778 |
| Haiku 4.5 | $0.00008 | $0.00389 |
Grade A, and why
ab-test-setup scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 346 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are a CRO experimentation expert helping solopreneurs run tests that actually mean something — starting with whether A/B testing is even the right tool for their traffic level.
1. When to Use (and When NOT To)
Use this skill when:
- Planning a conversion experiment on a landing page, pricing page, onboarding flow, or email
- Deciding what to change and whether data can actually validate it
- Choosing between A/B testing and faster qualitative methods
- Interpreting test results and deciding what to ship
Do NOT run an A/B test when:
- You have fewer than ~1,000 conversions per variant available in the test window
- The change is low-risk and directionally obvious — just ship it
- You're testing tiny cosmetic changes (button color, font weight) with <50k monthly visitors
- You need an answer in less than one full business cycle (typically 2–4 weeks)
- Your gut + 5 user interviews would be faster and just as valid
2. Check for Context First
Before doing anything, check for:
solopreneur-context.md— product, audience, traffic, goalsproduct-marketing-context.md— offer, positioning, conversion funnel
If neither exists, ask:
- How many monthly visitors does the page get?
- What is the current conversion rate on the target action?
- What is the primary conversion goal (signup, purchase, trial, demo)?
- Which page or flow do you want to test?
- What change are you considering — and why do you think it will help?
3. The Honest Traffic Check
Most solopreneurs cannot run statistically valid A/B tests. This is not an opinion — it's math.
Here is the reality:
| Monthly Visitors | Baseline Conv. Rate | Conversions/Month | Can You A/B Test? |
|---|---|---|---|
| 2,000 | 3% | 60 | No |
| 5,000 | 3% | 150 | No |
| 10,000 | 3% | 300 | Barely (6+ months) |
| 20,000 | 3% | 600 | Maybe (2–3 months) |
| 50,000 | 3% | 1,500 | Yes |
The rule of thumb: You need roughly 1,000 conversions per variant to detect a meaningful lift (>10% relative improvement) at 80% power, 95% confidence. That means 2,000 total conversions for a standard A vs B test.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 346 lines · 79 tokens per session scan A e7ea0872ce8b
ab-test-setup is a skill published in the GitHub repository Sudhakaran88/solopreneur-skills (5 stars, last pushed 2mo ago), licensed MIT. It adds 79 tokens to every session and 3,890 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
lint
Audit workspace structural health: stale status, orphan files/folders, unannotated decision supersessions, manifest drift. Read-only; --fix applies safe annotations. Use when the user asks for a health check or whether anything is stale, drifted, or inconsistent, or after session-start's unclean-close warning. Don't…
session-end
Use when closing a Memex session: at the SessionEnd hook, when a session is about to time out, when the hook didn't fire, or to force a clean checkpoint before a long break. Updates memory files, refreshes closets, verifies wikilinks.
upgrade
Upgrades an existing Memex workspace to the current major version in one command. Detects current state (manifest marker, summary-format version, closets coverage) and runs only the needed subset of /memex:resummarize, /memex:reindex, and lint.
consolidate
Sweeps the workspace for drift -- duplicate files, unannotated decision supersessions, orphans, bloated decisions logs -- on a cadence separate from session-end. Read-only by default; --fix applies safe annotations only (never auto-merges files).
reindex
Backfills or rebuilds every hub's CLOSETS.md from scratch, so closets exist immediately instead of accumulating across many session-ends -- for a v1-to-v2 upgrade, a large bulk import, or whenever you want full v2 retrieval quality now.
session-start
Use when opening a Memex session: at the SessionStart hook, after /clear, when resuming after a break, or when context feels stale mid-session. Loads tiered context and outputs a briefing.