reality-testing-decisions

reality-testing-decisions is a skill for Claude Code, Codex from apple-ouyang/book-to-skill. It costs 47 tokens per session (1,590 once invoked), scanned A, original, MIT.

A method for making decisions by running small, real-world experiments instead of trying to predict the outcome from opinions or intuition.

In plain words
What is it for?
Use it to test products, hiring choices, business ideas, career plans, and other decisions before committing significant time or money.
Why use it?
It replaces uncertain debate with evidence gathered from a low-cost test.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to test products, hiring choices, business ideas, career plans, and other decisions before committing significant time or money.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/apple-ouyang/book-to-skill/reality-testing-decisions
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add apple-ouyang/book-to-skill --skill reality-testing-decisions
Clone the repo
git clone --depth 1 https://github.com/apple-ouyang/book-to-skill

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for reality-testing-decisions

README.md
[![agentmods](https://agentmods.dev/badge/skills/apple-ouyang/book-to-skill/reality-testing-decisions/github.svg)](https://agentmods.dev/skills/apple-ouyang/book-to-skill/reality-testing-decisions)
Your own site
<a href="https://agentmods.dev/skills/apple-ouyang/book-to-skill/reality-testing-decisions"><img src="https://agentmods.dev/badge/skills/apple-ouyang/book-to-skill/reality-testing-decisions/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for reality-testing-decisions

Your own site · 80×15
<a href="https://agentmods.dev/skills/apple-ouyang/book-to-skill/reality-testing-decisions"><img src="https://agentmods.dev/badge/skills/apple-ouyang/book-to-skill/reality-testing-decisions.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 47 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,590 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00047 $0.01590
Opus 5 $0.00023 $0.00795
Sonnet 5 $0.00009 $0.00318
Haiku 4.5 $0.00005 $0.00159

Measured 10d ago against content hash 7089102d0fce, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

reality-testing-decisions scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/reality-testing-decisions/SKILL.md · 117 lines

How it starts

The opening of the file, as written. The whole thing — 117 lines — stays where its author put it; the contents beside it link to each section on GitHub.

用实验替代预测

任务目标

把"这个会不会成功?"的预测问题,转化为"怎么用最小成本验证?"的实验问题。

核心前提:我们对自己预测未来的能力严重高估。不要猜,不要感觉,去测试。

为什么预测不可靠

Tetlock 收集了 82361 个专家预测,结论:专家预测不如基本比率(用历史平均值外推),而基本比率又不如小规模实验。

额外的教育、20年经验、博士学位——都不能提高预测精度。反而,媒体曝光度越高的专家,预测越不准。

基本比率的力量:共同基金 vs 指数基金——基本比率数据如此清晰,以至于选共同基金几乎必然让你退休时更穷。当基本比率足够明确时,直接用它,不需要再做实验。

结论:只要有可能,就应该完全避免预测,用实验代替。

操作步骤

第一步:识别你在预测什么

把当前的问题写出来。如果它的形式是:

  • "我觉得这个产品会有市场"
  • "我感觉这个人能胜任"
  • "我认为这个方向是对的"

这就是在预测。继续下一步。

第二步:设计最小实验

问:用什么最小行动,可以在现实中得到反馈?

不同场景的实验设计:

产品/项目验证

  • CarsDirect 案例:不争论"网上卖车有没有人买",直接建个假网站,第一天卖出 3 辆车
  • 亦仁原则:号是消耗品,先发垃圾内容让市场反馈,不要等到完美再发
  • 默认项目都是通的,默认数据都是假的——对项目乐观去了解细节,对收入数据悲观谨慎投入

招聘/合作

  • 面试表现和工作表现几乎不相关(医学院研究:面试排名第 700 和第 100 的学生,入学后表现无差异)
  • 工作样本 > 面试:让候选人做一个真实任务,用结果评判,不用印象评判
  • 希望实验室:给潜在员工 3 周咨询合约,"面试表现最佳的人常常工作表现最差"
  • Steve Cole / HopeLab:$100k 设计项目,不赌一家公司,同时雇 5 家只做第一阶段($20k),用实际表现而非提案选人。从「OR 思维」转向「AND 思维」

职业/方向选择

  • 想读药学院?先去药房工作几周,哪怕无薪
  • NI 无线传感器:副总裁在投入 200-300 万前,先接了一个大学教授的小项目,走通了再加大投入
  • 绝对不投入做任何无法亲自体验和感受的项目

大家都觉得不靠谱的想法

  • 印度农民 app:库克和团队不看好,但让团队试了,农民收入提高 20%,最终 32.5 万农民使用
  • 不试试怎么知道靠不靠谱?

创业验证

  • 烘焙创业(Heath Brothers 案例):想开烘焙店,不要直接辞职全职投入。先在当地农贸市场摆一个月摊位,观察:有没有回头客?有没有盈利?盲测中你的产品排名如何?用真实市场反馈替代"邻居都夸我的布朗尼"式的自我确认。
  • 药学院实习(Heath Brothers 案例):想读药学院的学生,先去 3 家不同药房各实习 1 个月,而不是直接报名。实习成本:时间。替代的错误成本:4 年学费 + 发现不喜欢这份工作。

第三步:渐进式实验(当恐惧阻止你测试时)

强迫症律师佩吉的 7 步实验——当你觉得"万一出错怎么办"时,把实验拆得更小:

  1. 把法律摘要带回家,做三次修改
  2. 做两次修改
  3. 做一次修改
  4. 晚下班一小时,把摘要留在办公室
  5. 按时回家,不做多余修改
  6. 刻意留下一个标点错误
  7. 刻意留下一个语法错误

结果:没有公司败诉,没有人被解雇,甚至没有人注意到错误。

原理:每完成一步,你就获得了真实数据,而不是继续在脑子里预测"天会不会塌"。

第四步:判断何时停止实验,直接跳入

实验不是拖延的借口。两种情况要区分:

  • 杰森(应该实验):对海洋生物学感兴趣但不了解,先跟随一周,旁听几节课,确认后全力投入
  • 马歇尔(不应该实验):已经确定需要学位,用"先上一节课试试"来拖延,这是逃避
  • 丈夫想辞职(DecisiveWorkbook 案例):如果他在幻想另一份职业,先 ooch——赛车手梦想不像他想象的那么容易实现。但如果他已经确定要换工作,就设定触发条件("找到下一份工作再辞"),而不是无限期实验。

如果你已经确定了方向,小步尝试就是拖延。

注意事项

  • 企业家和公司高管最大的区别:高管相信"预测未来才能控制未来",企业家相信"控制未来就不需要预测它"
  • 60% 的世界 500 强 CEO 创业前没写过商业计划书——他们直接去卖
  • 实验的成本要足够低:亦仁标准是验证一个项目总花费控制在一个月工资以内,优先找只需要时间不需要钱的实验
  • 寻找反馈周期在三个月以内的实验;反馈周期太长,大多数人坚持不下去

Read the full file on GitHub · 117 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 10d ago First seen · 117 lines · 47 tokens per session scan A 7089102d0fce

Subscribe to this mod's changes

reality-testing-decisions is a skill published in the GitHub repository apple-ouyang/book-to-skill (130 stars, last pushed 6mo ago), licensed MIT. It adds 47 tokens to every session and 1,590 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

local-ai-agents

Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…

microsoft/ai-agents-for-beginners · 200 tokens

next-cache-components-adoption

Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…

vercel/next.js · 95 tokens

insight-error-page

Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…

vercel/next.js · 83 tokens

next-cache-components-optimizer

Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…

vercel/next.js · 170 tokens

next-partial-prefetching-adoption

Turn on Partial Prefetching in a Next.js app and work through the insights it surfaces. Use when the user wants to enable or adopt Partial Prefetching, flip the partialPrefetching flag, opt routes in with export const prefetch = 'partial', audit Link prefetch={true} behavior, preserve existing prefetched UI with…

vercel/next.js · 103 tokens