Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add emichy/product-judgment --skill review-interviewgit clone --depth 1 https://github.com/emichy/product-judgmentWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/emichy/product-judgment/review-interview)<a href="https://agentmods.dev/skills/emichy/product-judgment/review-interview"><img src="https://agentmods.dev/badge/skills/emichy/product-judgment/review-interview/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/emichy/product-judgment/review-interview"><img src="https://agentmods.dev/badge/skills/emichy/product-judgment/review-interview.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00119 | $0.00781 |
| Opus 5 | $0.00060 | $0.00391 |
| Sonnet 5 | $0.00024 | $0.00156 |
| Haiku 4.5 | $0.00012 | $0.00078 |
Grade A, and why
review-interview scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 67 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Review an interview
Code review of the call. Transcript only — never invent quotes. Direct, no
applause — same voice as pressure-test-a-bet.
House style
Answer first, in one sentence. Then only the evidence that carries it. Evidence is selected, not gathered — the two quotes that carry it, not the six that mention it. Every section earns its place by changing what they'd do this week; if it doesn't, cut it. Slack, not email. No preamble, no applause, and none of "the move" / "the play" / "the tell" / "the key insight here" / "Here's my take."
Output contract
| Rule | |
|---|---|
| Line 1 | Fail/pass on evidence — not vibe |
| Body | 2–3 moments, then stop |
| Stop | No features. No "what we learned." One line → extract-customer-insights |
Line 1 examples: "You pitched. They said cool. That's not evidence." / "You got a reaction to the thing. You didn't get how they work." / "You got the workflow. You didn't get the cost." / "The call was over when they said Excel. You wrapped."
The bar
| Layer | Got it? |
|---|---|
| Behavior | Workflow reconstructed — tools, sequence, a real moment |
| Cost | Consequence — time, money, risk — not "annoying" |
| Tradeoff | Sacrifice — what they'd cut, delay, or pay |
Most calls get behavior. Few get cost. Almost none get tradeoff.
Separate the tape from your read. Quotes and behavior are observed; what they imply is inference. A strong intuition is worth naming — don't upgrade it into evidence.
Format (2–3 moments max)
[timestamp if the tape has one] "Exact quote" ↳ Asked: [their question] ↳ Instead: [what to say or show] ↳ Would reveal: [specific insight]
Did well: one concrete thing — not rapport.
Misses to flag:
- Asked instead of showed — had a prototype/PR/screen and ran a survey
- Never widened out — every question tested the anchor, nothing came back they didn't already suspect
- Signal was late — gold quote landed, then thank-you out. Name it: "The call was over when they said [X]. You wrapped."
- Would-you-pay — always yes. Instead: what they already spend (hours, tool, invoice, who signs). Not a better hypothetical. See
prep-customer-interviewStakes. - Story or summary? "We usually" without a moment = summary.
- They talked more than ~30% — say so.
- Prep miss — one line if searchable context existed before the call
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 67 lines · 119 tokens per session scan A 83c8554ffffe
review-interview is a skill published in the GitHub repository emichy/product-judgment (4 stars, last pushed 23d ago), licensed MIT. It adds 119 tokens to every session and 781 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
first-customer-finder
Find and qualify evidence-backed potential first customers, early adopters, design partners, or beta users for a startup using recent public pain and buying signals. Use when Codex needs to analyze a product URL or idea, define an ideal customer profile, research public discussions and business pages, identify…
gingiris-user-interview
A playbook for interviewing users and operating an early product launch, including screening participants, running interviews, testing beta versions, and reviewing churn. It is based on a described process for finding product-market fit, meaning a product consistently meeting a real user need.
gr-user-interview
A framework for interviewing users to discover product-market fit, meaning a strong match between a product and a real user need.
run-customer-discovery
Use when validating a startup idea, product hypothesis, or new feature by conducting structured interviews with target customers before building.
jtbd-extractor
Turn raw research into Jobs-to-be-Done statements showing what users are really trying to accomplish. Use when: extract jobs, jtbd analysis, jobs to be done, what job is the user hiring, underlying user needs.
stage-map
This skill should be used when a founder wants to know which startup stage they're actually in, what the goal and exit criteria of that stage are, what failure modes to watch for, or which tool/module to use next — e.g. "what stage am I at", "what should I focus on now", "am I ready for an MVP", "do I have…