Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/mehrad-dm/mastermind/qanpx skills add mehrad-dm/mastermind --skill qagit clone --depth 1 https://github.com/mehrad-dm/mastermindWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mehrad-dm/mastermind/qa)<a href="https://agentmods.dev/skills/mehrad-dm/mastermind/qa"><img src="https://agentmods.dev/badge/skills/mehrad-dm/mastermind/qa.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00066 | $0.01777 |
| Opus 5 | $0.00033 | $0.00889 |
| Sonnet 5 | $0.00013 | $0.00355 |
| Haiku 4.5 | $0.00007 | $0.00178 |
Grade A, and why
qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 105 lines — stays where its author put it; the contents beside it link to each section on GitHub.
QA: prove it works (verify by default, tests on request)
Proving a change works is mandatory: though the proof is rarely a test suite. Default to verify (drive the real thing, watch what it does). Tests / TDD are a project choice: offer them, don't impose them. After a build, the honest close is: "Built it: want me to QA it? (I can verify it end-to-end, and add tests / do it test-first if you want.)"
Mode 1: Verify (the default; always do this)
-
Write the checklist before you look. Turn the expected behavior into a short list of criteria that are each individually checkable: one line per criterion, phrased so the answer is met or not met, never "looks fine". Write it before exercising anything, because a list written while looking is a list that describes what you found. If you can't say what correct looks like, you can't verify it.
Prefer criteria a person could count or observe over ones needing an opinion: "a second submit while pending is rejected" beats "handles concurrency well." This list is what you report against in step 5, and what makes "it works" mean something.
-
Pick the lightest real check. Drive the actual thing over reasoning about it: run the app and click the flow, hit the endpoint, run the script/CLI, render the component. Reuse the project's run/dev command; reach for harness the project already has.
-
Happy path, then the edges that matter: empty, null, error, loading, zero/one/many, unauthorized, malformed, offline/slow (
~/.mastermind/engineering/core/rigor.md). Observe actual output and state. -
Check the invisible: typecheck, lint, build; console/network for errors; for UI, keyboard + focus, contrast, no layout shift/regression nearby.
-
Say what must NOT happen, not only what must. Half of a real check is a negative: no request was sent, the tokens were not cleared, the user was not bounced to login, no second write landed. A test that only asserts the visible message passes while the damage happens behind it.
-
Report with evidence: what you ran and what you observed (command output, response, screenshot). State confidence plainly. Couldn't run a check? Say so; never present unrun work as verified.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 105 lines · 66 tokens per session scan A 292d99dcf88e
qa is a skill published in the GitHub repository mehrad-dm/mastermind (24 stars, last pushed 4d ago), licensed MIT. It adds 66 tokens to every session and 1,777 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
docker-extend
Use when: User wants to extend Docker with custom tools, personalize the Docker environment, or set up user-specific Docker customization. Triggers: 'extend docker', 'docker-extend', 'add tools to docker', 'customize docker', 'add my tools to the container', 'personalize docker setup', 'docker user setup', 'install…
write-zot-themes
Help the user create, install, or package zot themes, including theme-only extensions.
st-full-workflow
Use when the user asks to run the complete end-to-end Strikethroo workflow for a work order in one shot in this repository — triggers include full workflow, end-to-end, plan and execute, do everything, run the whole strikethroo workflow. Do not use when the user wants only one stage (create a plan, generate tasks, or…
st-code-review
Use when the blueprint execution gate asks for an independent second-harness review of a Strikethroo plan's cumulative diff in this repository — triggers include code review gate, review the plan diff, second-model review, CODEREVIEW hook, review the cumulative diff. Do not use to review a single task, to review code…
st-refine-plan
Use when the user asks to review, refine, improve, interrogate, pressure-test, or update an existing Strikethroo plan by plan ID in this repository — triggers include refine plan, improve plan, review plan, red-team the plan, update plan. Do not use to create a new plan, to generate tasks, or for generic brainstorming…
tlamatini-daily-chat-test
Run the daily automated Tlamatini chat regression — drive a visible Chrome via Playwright, log into agentpage.html, ask up to 1000 curated safe questions one-by-one (Multi-Turn ON, ACPX/Ask-Execs/Exec-Report/Internet OFF), wait for and qualify each answer (heuristic + LLM judge on failures), then write a dated report…