experiment-results-interpreter

experiment-results-interpreter is a skill for Claude Code, Codex from KirKruglov/claude-skills-kit. It costs 93 tokens per session (1,567 once invoked), scanned A, original, MIT.

A skill that explains A/B test results in everyday language and recommends whether to release, undo, or continue testing a change. A/B testing compares two versions using a chosen measure, such as sign-ups or purchases.

In plain words
What is it for?
Reviewing experiment results, checking significance and guardrail measures, choosing a next step, and drafting a summary for stakeholders.
Why use it?
It turns test numbers into a clear decision without requiring statistics expertise or database access.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Reviewing experiment results, checking significance and guardrail measures, choosing a next step, and drafting a summary for stakeholders.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/kirkruglov/claude-skills-kit/experiment-results-interpreter
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add KirKruglov/claude-skills-kit --skill experiment-results-interpreter
Clone the repo
git clone --depth 1 https://github.com/KirKruglov/claude-skills-kit

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for experiment-results-interpreter

README.md
[![agentmods](https://agentmods.dev/badge/skills/kirkruglov/claude-skills-kit/experiment-results-interpreter/github.svg)](https://agentmods.dev/skills/kirkruglov/claude-skills-kit/experiment-results-interpreter)
Your own site
<a href="https://agentmods.dev/skills/kirkruglov/claude-skills-kit/experiment-results-interpreter"><img src="https://agentmods.dev/badge/skills/kirkruglov/claude-skills-kit/experiment-results-interpreter/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for experiment-results-interpreter

Your own site · 80×15
<a href="https://agentmods.dev/skills/kirkruglov/claude-skills-kit/experiment-results-interpreter"><img src="https://agentmods.dev/badge/skills/kirkruglov/claude-skills-kit/experiment-results-interpreter.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 93 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,567 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00093 $0.01567
Opus 5 $0.00046 $0.00783
Sonnet 5 $0.00019 $0.00313
Haiku 4.5 $0.00009 $0.00157

Measured 12d ago against content hash 5c21ec014b8f, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

experiment-results-interpreter scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/data-analysis/experiment-results-interpreter/SKILL.md · 151 lines

How it starts

The opening of the file, as written. The whole thing — 151 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Experiment Results Interpreter

This skill takes A/B test results — pasted from an analytics dashboard or described in plain text — and returns a plain-language significance assessment, a ship/rollback/extend recommendation with a documented rationale, and a ready-to-paste stakeholder summary. No statistics background or database access required.

Input:

  • Test description: hypothesis, variant names, primary metric, test duration
  • Results: pre-computed (p-value or confidence interval + lift) or raw numbers (visitors and conversions per variant)
  • Optional: guardrail metrics (secondary metrics to protect)

Output:

  • Test Summary, Results Interpretation, Recommendation with rationale, Draft Stakeholder Summary

Language Detection

Detect the user's language from their message:

  • If Russian (or contains Cyrillic): respond in Russian
  • If English (or other Latin-script language): respond in English
  • If ambiguous: respond in the language of the trigger phrase used

Instructions

Step 1: Validate Input

  1. Check that the user has provided at minimum:

    • A primary metric (what was being measured)
    • At least one result value (conversion rate, lift, p-value, or raw visitor/conversion counts)
  2. If the primary metric is missing: ask "What metric was this experiment measuring? (e.g., signup rate, checkout conversion, 7-day retention)"

    • Exception: if the user refers to "primary metric" or "main metric" without naming it but does provide result values (lift %, p-value, or counts) — proceed using "primary metric" as the metric name placeholder rather than blocking. Name it "primary metric" in the output.
  3. If no results data at all: ask for one of:

    • p-value or confidence interval from their analytics tool
    • Control and treatment: visitors and conversions (to compute significance here)
  4. If statistical data is present but no test description: proceed — infer variant names as "Control" and "Treatment" if not specified.

  5. Do not ask more than one clarifying question at a time. Prioritise the most critical missing piece.

Read the full file on GitHub · 151 lines

Files

What ships with it

4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 151 lines · 93 tokens per session scan A 5c21ec014b8f

Subscribe to this mod's changes

experiment-results-interpreter is a skill published in the GitHub repository KirKruglov/claude-skills-kit (18 stars, last pushed 1mo ago), licensed MIT. It adds 93 tokens to every session and 1,567 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

spark-video-director

Translate a screenplay (one scene at a time) into a provider-agnostic storyboard fragment for the spark-video pipeline. Wraps Shanyin Super Director Master when available — the upstream Shanyin SKILL is the single source of truth for craft when present.

modelstudioai/skills · 58 tokens

spark-video-episode

One-shot autopilot orchestrator — runs the full spark-video pipeline (screenwriter ↔ director per-scene parallel → render chain-DAG parallel + per-clip review → stitch). User confirms at 4 gates (+ 1 mode gate at start + 1 BGM gate when bgm/ folder detected). Use when the user wants "make me an episode" in one command.

modelstudioai/skills · 83 tokens

vox-video-director

Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end with Aliyun Bailian CLI + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all automated. Use this whenever the user wants a "Vox style" video, a paper/torn-paper collage animation, a "motion collage"…

modelstudioai/skills · 238 tokens

bailian-train-deploy

A workflow for using Alibaba Cloud’s Bailian command-line tool to fine-tune or directly deploy AI models as callable services. It covers text, speech-synthesis, image-generation, and video-generation models.

modelstudioai/skills · 321 tokens

spark-video-cast

Scaffold and generate reference assets for characters (cast), locations (movie-set / set dressing), and key props — the three pillars of visual consistency in spark-video. Wraps bl image generate / edit for portrait creation. Use when adding new characters/locations/props or when costume/state changes are needed.

modelstudioai/skills · 66 tokens

spark-video-screenwriter

Turn a user's premise into a structured screenplay (one scene at a time) for the spark-video pipeline. Wraps Shanyin Super Screenwriting Master when available — that upstream Shanyin SKILL is the single source of truth for craft when present.

modelstudioai/skills · 56 tokens