Getting it into your agent
This one installs as part of its plugin. Adding the marketplace and installing the plugin brings it with everything else the plugin ships.
/plugin marketplace add gigsmart/haiku-method/plugin install haikuWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/gigsmart/haiku-method/haiku-pressure-testing)<a href="https://agentmods.dev/skills/gigsmart/haiku-method/haiku-pressure-testing"><img src="https://agentmods.dev/badge/skills/gigsmart/haiku-method/haiku-pressure-testing.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00018 | $0.00305 |
| Opus 5 | $0.00009 | $0.00152 |
| Sonnet 5 | $0.00004 | $0.00061 |
| Haiku 4.5 | $0.00002 | $0.00030 |
Grade A, and why
haiku-pressure-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Pressure Testing
Adversarially test hat definitions using Evaluation-Driven Development (RED-GREEN-REFACTOR).
Process
-
If no hat specified, list available hats from
$CLAUDE_PLUGIN_ROOT/studios/software/stages/*/hats/and ask the user to pick one. -
Load the hat definition file.
-
Design Pressure Scenario:
- Combine 3+ pressure types (time, sunk cost, authority, economic, exhaustion, social, pragmatic)
- Target the hat's most important constraints
- Present scenario to user for approval
-
RED Phase (baseline without anti-rationalization table):
- Run scenario with a subagent that has hat instructions minus anti-rationalization table
- Document verbatim: decisions made, rationalizations used, sections violated
-
GREEN Phase (full hat definition):
- Run same scenario with full hat definition
- Agent MUST cite specific hat sections, acknowledge temptations
- PASS if correct decision citing hat sections; FAIL if rationalized past them
-
REFACTOR Phase:
- If GREEN failed: capture rationalization, add to anti-rationalization table, re-run
- If GREEN passed: document that hat held under pressure
-
Commit artifacts to
.haiku/pressure-tests/
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 35 lines · 18 tokens per session scan A fe8ee09a348a
haiku-pressure-testing is a skill published in the GitHub repository gigsmart/haiku-method (24 stars, last pushed 3mo ago), licensed Apache-2.0. It adds 18 tokens to every session and 305 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
red-green-refactor
Guides the red-green-refactor TDD workflow: write a failing test first, implement the minimum code to make it pass, then refactor while keeping tests green. Use when a user asks to practice TDD, write tests first, follow red-green-refactor, do test-driven development, write failing tests before code, or phrases like…
superpowers-zh
Use when constraining AI coding with Chinese TDD methodology, systematic debugging, code review, and verification workflows. Superpowers-zh: Chinese adaptation of the Superpowers AI-assisted programming skills and methodologies.
simple-tdd
Test-driven development methodology — write the failing test first, then the minimum code to pass it.
tsq-bdd
A guide to Behavior-Driven Development (BDD), a way to describe software behavior in business language using Given-When-Then scenarios. These scenarios can serve as shared documentation and acceptance tests.
test-case-design
A method for turning product and technical requirements into detailed tests with clear checks and links back to the original requirements. TDD means writing tests before implementation, but this add-on describes test design rather than test execution.
test-execution-router
A process for preparing, routing, running, and recording tests in two stages: an initial health check and later formal verification.