posthog_experiment

A guide for setting up a PostHog experiment or feature-flag rollout. An A/B experiment compares two versions, while a feature flag controls who can see a change.

In plain words
What is it for?
Use it to choose primary and safety metrics, set exposure events, calculate test duration, verify both variants, and ramp a rollout safely.
Why use it?
It helps define a measurable hypothesis, enroll users correctly, estimate the needed sample, and decide in advance when to ship, stop, or revise the change.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/tarekkharsa/agentstack/experiment
Any agent
npx skills add Tarekkharsa/agentstack --skill experiment
Clone the repo
git clone --depth 1 https://github.com/Tarekkharsa/agentstack

Made for: Claude Code, Codex.

Per session 35 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 470 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00035 $0.00470
Opus 5 $0.00017 $0.00235
Sonnet 5 $0.00007 $0.00094
Haiku 4.5 $0.00003 $0.00047

Measured yesterday against content hash 0fcb96a1f9fd, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

posthog_experiment scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

crates/cli/catalog/skills/posthog/experiment/SKILL.md · 45 lines

What it actually says

PostHog Experiment

Unofficial, agentstack-authored. Not affiliated with or endorsed by PostHog.

Use this skill when a team wants to test a change with a PostHog experiment or roll a feature out behind a flag, and needs the setup to actually yield a trustworthy result.

Workflow

  1. Write the hypothesis as one falsifiable sentence: "Changing X will move [primary metric] by roughly Y for [population]." If you cannot state it this way, the experiment is not ready.
  2. Choose one primary metric that maps directly to the goal (e.g. signup conversion), plus a small set of guardrail metrics that must not regress.
  3. Define the exposure point — the feature flag and the event that marks a user as enrolled. Users must be counted from the moment they could see the change, not from when they convert.
  4. Estimate the needed sample size and runtime from the baseline rate and the minimum effect worth detecting. State how long the test must run; resist calling it early.
  5. Set the rollout: start the flag at a safe percentage, confirm assignment is stable per user, and verify both variants render before ramping.
  6. Pre-commit to the decision rule (ship / kill / iterate) and the metrics that decide it, in writing, before launch.

Conventions

  • One primary metric per experiment. Multiple primaries invite cherry-picking.
  • Do not peek-and-stop: respect the planned runtime and sample size.
  • Keep the feature flag key descriptive and consistent with the experiment name.
  • Roll out gradually (e.g. 5% → 25% → 100%) and watch guardrails at each step.

Boundaries

  • Never declare a winner before the experiment reaches its planned sample size or runtime — early results are noise.
  • Do not change the primary metric or population mid-flight to chase significance.
  • Do not flip a flag to 100% for all users without an explicit go-ahead.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 45 lines · 35 tokens per session scan A 0fcb96a1f9fd

Subscribe to this mod's changes

posthog_experiment is a skill published in the GitHub repository Tarekkharsa/agentstack (3 stars, last pushed 18d ago), licensed Apache-2.0. It adds 35 tokens to every session and 470 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.