reality-checker

reality-checker is a skill for Claude Code, Codex from RaNDoM6913/claude-code-superkit. It costs 30 tokens per session (1,174 once invoked), scanned A, original, MIT.

Evidence-based readiness assessor — defaults to NEEDS WORK, refuses fantasy A+ ratings, demands overwhelming proof before declaring anything production-ready.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/random6913/claude-code-superkit/reality-checker
Any agent
npx skills add RaNDoM6913/claude-code-superkit --skill reality-checker
Clone the repo
git clone --depth 1 https://github.com/RaNDoM6913/claude-code-superkit

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for reality-checker

README.md
[![agentmods](https://agentmods.dev/badge/skills/random6913/claude-code-superkit/reality-checker.svg)](https://agentmods.dev/skills/random6913/claude-code-superkit/reality-checker)
Your own site
<a href="https://agentmods.dev/skills/random6913/claude-code-superkit/reality-checker"><img src="https://agentmods.dev/badge/skills/random6913/claude-code-superkit/reality-checker.svg" alt="Measured on agentmods" height="20"></a>
Per session 30 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,174 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin unknown No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00030 $0.01174
Opus 5 $0.00015 $0.00587
Sonnet 5 $0.00006 $0.00235
Haiku 4.5 $0.00003 $0.00117

Measured today against content hash bca1ad016cba, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

reality-checker scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

| "API endpoint live" | `curl` output with full request + response, including auth |
packages/codex/skills/reality-checker/SKILL.md · 121 lines

How it starts

The opening of the file, as written. The whole thing — 121 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Reality Checker

The last line of defense against premature "production ready" claims. Defaults to NEEDS WORK unless overwhelming evidence proves otherwise. No fantasy 98/100 ratings. No "looks good to me" without screenshots, logs, or test runs.

Phase 0: Load Project Context

Read if exists:

  1. AGENTS.md or CLAUDE.md — project's quality bar, deployment requirements
  2. docs/architecture/*.md — what "done" means for this system
  3. The original task / spec — what was actually requested

Verification Discipline

  • No approval without fresh evidence. A claim of "done/fixed/passing" requires fresh command output (test/build/run) printed in this turn — not a description, not "should work".
  • Hedge words auto-reject. If the work is justified with "should", "probably", "seems to", "I believe", or "appears to" instead of evidence, mark it NOT verified.
  • Verification is a separate pass from the one that authored the change — re-derive the result, don't trust the author's summary.
  • Work is done when verification passes — not when it compiles. A missing "yes" means "no".

When to Use

  • Before merging a PR that claims to "complete" a feature
  • Before tagging a release
  • When another agent reports "this is done"
  • After UI/UX work to verify what was claimed actually exists
  • When a previous review gave a high score without evidence

Default Verdict

NEEDS WORK until disproven by evidence. In practice, most "ready" claims are 30-60% complete. Defaulting to ready is statistically wrong.

Evidence Requirements

Claim Required evidence
"Feature X works" Screenshot or recording end-to-end on actual app, not a localhost mock
"Tests pass" Test runner output + green count + coverage % + names of new tests
"API endpoint live" curl output with full request + response, including auth
"DB migration safe" EXPLAIN plan + rollback script + tested on real-size dataset
"No regressions" Diff of test results before/after OR exhaustive list of tested flows
"UI matches design" Side-by-side: spec image + actual screenshot at correct viewport
"Performance improved" Before/after benchmark with same input, run 3+ times
"Security reviewed" Specific threats considered + mitigations applied

Read the full file on GitHub · 121 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. today First seen · 121 lines · 30 tokens per session scan A bca1ad016cba

Subscribe to this mod's changes

reality-checker is a skill published in the GitHub repository RaNDoM6913/claude-code-superkit (2 stars, last pushed 1mo ago), licensed MIT. It adds 30 tokens to every session and 1,174 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.

Related

Other skills, from other repositories

apm-review-panel

Use this skill to run a multi-persona expert advisory review on a labelled pull request in microsoft/apm. The panel fans out to five mandatory specialists plus a test-coverage specialist (active on every PR that touches src/) plus three conditional specialists (auth, doc-writer, performance-expert), all running in…

microsoft/apm · 178 tokens

batch-bug-shepherd

Use this skill to drive a batch of suspected bugs in microsoft/apm from raw issue list to mergeable PR queue. Fan out one triage subagent per issue (LEGIT / UNCLEAR / FIXED-AT-HEAD), gate every legit bug against PRINCIPLES.md via an apm-ceo strategic-alignment pass, cross-reference legit issues against open PRs, then…

microsoft/apm · 227 tokens

apm-issue-autopilot

Use this skill to drive any open microsoft/apm issue (bug, feature, docs, refactor, perf) from raw intake to a mergeable PR with triage as the central, paramount gate. Run the apm-triage-panel rubric per issue first, then present ONE consolidated triage review for the whole batch and escalate to the maintainer BY…

microsoft/apm · 238 tokens

apm-spec-guardian

Use this skill to run a four-panel adversarial advisory review on any pull request that touches the OpenAPM specification artifact (docs/src/content/docs/specs/openapm-.md), its inline / sidecar JSON Schemas (docs/src/content/docs/specs/schemas/.schema.json), or the conformance fixture seed…

microsoft/apm · 215 tokens

apm-triage-panel

Use this skill to triage one microsoft/apm issue selected by the daily sweep, the status/needs-triage fast path, or manual dispatch. Emit one synthesized comment with a decision, label set, exact milestone, and suggested next action.

microsoft/apm · 59 tokens

cli-logging-ux

Use this skill when editing or creating CLI output, logging, warnings, error messages, progress indicators, or diagnostic summaries in the APM codebase. Activate whenever code touches console helpers (richsuccess, richwarning, richerror, richinfo, richecho), DiagnosticCollector, STATUSSYMBOLS, CommandLogger, or any…

microsoft/apm · 94 tokens