Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add hardiktiwari/PM-operating-OS --skill experiment-writeupgit clone --depth 1 https://github.com/hardiktiwari/PM-operating-OSWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/hardiktiwari/pm-operating-os/experiment-writeup)<a href="https://agentmods.dev/skills/hardiktiwari/pm-operating-os/experiment-writeup"><img src="https://agentmods.dev/badge/skills/hardiktiwari/pm-operating-os/experiment-writeup.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00031 | $0.00443 |
| Opus 5 | $0.00015 | $0.00221 |
| Sonnet 5 | $0.00006 | $0.00089 |
| Haiku 4.5 | $0.00003 | $0.00044 |
Grade A, and why
experiment-writeup scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Experiment Writeup
Help PMs document experiment results in a structured format. Turns raw data into clear narratives with hypothesis, methodology, results, learnings, and decision.
When to Use
- Documenting A/B test outcomes
- Writing up feature experiment results
- Sharing experiment learnings with stakeholders
- Creating a record for future reference
- When asked "help me write up the results"
Process / Template
1. Gather the Data
- Hypothesis (original statement)
- Experiment design (variants, duration, sample size)
- Primary, secondary, and guardrail metric results
- Statistical significance (p-values, confidence intervals)
- Any qualitative feedback or observations
2. Structure the Writeup
Hypothesis
- Restate the original hypothesis
- Brief context on why we ran this
Methodology
- Variants tested (control vs. treatment)
- Duration and sample size
- Target segment
- Any caveats (traffic issues, external events)
Results
- Primary metric — direction, magnitude, significance
- Secondary metrics — supporting or conflicting signals
- Guardrail metrics — did anything regress?
Learnings
- What did we learn? (beyond the numbers)
- Surprises or unexpected findings
- Implications for future work
Decision
- Ship / Iterate / Kill
- Rationale for the decision
- Next steps (if iterating)
3. Write Clearly
- Lead with the decision and key takeaway
- Use plain language; avoid jargon
- Include numbers with context (e.g., "+12% vs. control")
- Call out statistical significance explicitly
Output
A structured Experiment Writeup suitable for:
- Stakeholder sharing (Slack, email)
- Internal documentation
- Experimentation platform notes
- Retrospectives and planning
Format: concise, scannable, decision-oriented. Typically 1–2 pages or equivalent in markdown.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 71 lines · 31 tokens per session scan A 8d758b5180c0
experiment-writeup is a skill published in the GitHub repository hardiktiwari/PM-operating-OS (5 stars, last pushed 4mo ago), licensed MIT. It adds 31 tokens to every session and 443 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
firebase-messaging
Use when setting up Firebase Cloud Messaging, managing permissions and tokens, handling background/foreground notification taps, or dispatching messages server-side (HTTP v1).
firebase-cloud-functions
Use when calling callable functions (httpsCallable), passing data to server-side logic, handling function errors/timeouts, configuring regions, or testing with the Emulator Suite.
karpathy-llm-wiki
Use when building or maintaining a personal LLM-powered knowledge base. Triggers: ingesting sources into a wiki, querying wiki knowledge, linting wiki quality, 'add to wiki', 'what do I know about', or any mention of 'LLM wiki' or 'Karpathy wiki'.
excalidraw-architect
Choose and compose the right Excalidraw diagram - architecture, flowchart, sequence, state, ER, swimlane, process, timeline, quadrant, pyramid, venn, loop, gantt, bar, line, scatter, and more - using the excalidraw-architect-mcp server. Use whenever a reader would learn more from a picture than from prose, or when…
sdd
Execute the Liatrio Spec-Driven Development (SDD) workflow when explicitly invoked by the user. NOTE: this skill is NOT intended to be dynamically loaded or automatically triggered; it should only ever be explicitly called by the user.
bcc-throughline
BCC global progress cockpit (plans.md/progress.md/findings.md). Slash: /bcc-throughline · chat: bcc:throughline · "where are we" · reprioritize · resume after /clear. Not for coding or full PLAN grill.