Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/nullhack/temple8/confirm-red-failurenpx skills add nullhack/temple8 --skill confirm-red-failuregit clone --depth 1 https://github.com/nullhack/temple8What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00030 | $0.00328 |
| Opus 5 | $0.00015 | $0.00164 |
| Sonnet 5 | $0.00006 | $0.00066 |
| Haiku 4.5 | $0.00003 | $0.00033 |
Grade A, and why
confirm-red-failure scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Confirm Red Failure
- Load [[software-craft/tdd]] — red/green discipline and the right-reason rule.
- Remove the pending marker from all the target contract's tests for this cycle. Removing the decorator may orphan
import pytestwhen the test used it for nothing else — run ruff and drop the unused import; that is a lint side-effect of un-marking, not a behaviour edit. Do not author tests; they already exist with full bodies. - Run the tests. Confirm the failure is ours: IF the source .py is absent and the deferred in-body import raises
ImportErrorTHEN new work; IF the source .py exists but is stale against the changed contract and an assertion fails THEN rework. Either is the right reason per [[software-craft/tdd]]. - IF the red fails for the wrong reason (a typo, a bad fixture, a contract the tests themselves violate) THEN reject it — do not patch the test.
- IF investigation reveals the contract itself is the problem — a referenced cassette is missing, an assertion no correct implementation could satisfy, the spec contradicts itself — THEN append the finding to
.cache/<session_id>/journal.md(contract, the gap, the evidence) and firereveals-gap. This routes back to plan, not green. - Read the target contract's .pyi first, not the whole tree.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 14 lines · 30 tokens per session scan A 04a0b2dace4e
confirm-red-failure is a skill published in the GitHub repository nullhack/temple8 (11 stars, last pushed 28d ago), licensed MIT. It adds 30 tokens to every session and 328 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
test-guide
Test-writing guide for Backend.AI — propose success/exception/edge scenarios first, refine them with the user, then implement while reporting per-scenario verification status. Covers fixtures, withtables, mock repositories, pants test, optional TDD cadence.
tdd
Drive a red → green → refactor cycle for the active task.
tdd
Guided TDD workflow — plan, tracer bullet, incremental RED-GREEN.
agent-integration
Run all three agent integration phases sequentially: research, write-tests, and implement using E2E-first TDD (unit tests written last). For individual phases, use /agent-integration:research, /agent-integration:write-tests, or /agent-integration:implement. Use when the user says "integrate agent", "add agent…
engram-testing-coverage
TDD and coverage standards for Engram. Trigger: When implementing behavior changes in any package.
code-assist
Guides implementation of code tasks using test-driven development in an Explore, Plan, Code, Commit workflow. Acts as a Technical Implementation Partner and TDD Coach — following existing patterns, avoiding over-engineering, and producing idiomatic, modern code.