qa

qa is a skill for Claude Code, Codex from mehrad-dm/mastermind. It costs 66 tokens per session (1,777 once invoked), scanned A, original, MIT.

A verification step used after a feature or fix, or before shipping, to check the real behavior against a written list of expected results. Test-driven development, or TDD, is offered as an optional way to write tests before implementation rather than required by default.

In plain words
What is it for?
Use it to drive a feature or fix end to end, confirm each acceptance criterion, write tests when requested, or work test-first when that suits the project.
Why use it?
It replaces the vague question “does it work?” with observable checks that can reveal failures in the finished change.

Skill for Claude CodeCodex

Part of the mastermind plugin — 23 skills, 4 agents, 1 hook shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/mehrad-dm/mastermind/qa
Any agent
npx skills add mehrad-dm/mastermind --skill qa
Clone the repo
git clone --depth 1 https://github.com/mehrad-dm/mastermind

Made for: Claude Code, Codex.

Or install mastermind, the plugin that ships this one along with the rest of its 23 skills, 4 agents, 1 hook.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for qa

README.md
[![agentmods](https://agentmods.dev/badge/skills/mehrad-dm/mastermind/qa.svg)](https://agentmods.dev/skills/mehrad-dm/mastermind/qa)
Your own site
<a href="https://agentmods.dev/skills/mehrad-dm/mastermind/qa"><img src="https://agentmods.dev/badge/skills/mehrad-dm/mastermind/qa.svg" alt="Measured on agentmods" height="20"></a>
Per session 66 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,777 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00066 $0.01777
Opus 5 $0.00033 $0.00889
Sonnet 5 $0.00013 $0.00355
Haiku 4.5 $0.00007 $0.00178

Measured 4d ago against content hash 292d99dcf88e, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/qa/SKILL.md · 105 lines

How it starts

The opening of the file, as written. The whole thing — 105 lines — stays where its author put it; the contents beside it link to each section on GitHub.

QA: prove it works (verify by default, tests on request)

Proving a change works is mandatory: though the proof is rarely a test suite. Default to verify (drive the real thing, watch what it does). Tests / TDD are a project choice: offer them, don't impose them. After a build, the honest close is: "Built it: want me to QA it? (I can verify it end-to-end, and add tests / do it test-first if you want.)"

Mode 1: Verify (the default; always do this)

  1. Write the checklist before you look. Turn the expected behavior into a short list of criteria that are each individually checkable: one line per criterion, phrased so the answer is met or not met, never "looks fine". Write it before exercising anything, because a list written while looking is a list that describes what you found. If you can't say what correct looks like, you can't verify it.

    Prefer criteria a person could count or observe over ones needing an opinion: "a second submit while pending is rejected" beats "handles concurrency well." This list is what you report against in step 5, and what makes "it works" mean something.

  2. Pick the lightest real check. Drive the actual thing over reasoning about it: run the app and click the flow, hit the endpoint, run the script/CLI, render the component. Reuse the project's run/dev command; reach for harness the project already has.

  3. Happy path, then the edges that matter: empty, null, error, loading, zero/one/many, unauthorized, malformed, offline/slow (~/.mastermind/engineering/core/rigor.md). Observe actual output and state.

  4. Check the invisible: typecheck, lint, build; console/network for errors; for UI, keyboard + focus, contrast, no layout shift/regression nearby.

  5. Say what must NOT happen, not only what must. Half of a real check is a negative: no request was sent, the tokens were not cleared, the user was not bounced to login, no second write landed. A test that only asserts the visible message passes while the damage happens behind it.

  6. Report with evidence: what you ran and what you observed (command output, response, screenshot). State confidence plainly. Couldn't run a check? Say so; never present unrun work as verified.

Read the full file on GitHub · 105 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 105 lines · 66 tokens per session scan A 292d99dcf88e

Subscribe to this mod's changes

qa is a skill published in the GitHub repository mehrad-dm/mastermind (24 stars, last pushed 4d ago), licensed MIT. It adds 66 tokens to every session and 1,777 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

docker-extend

Use when: User wants to extend Docker with custom tools, personalize the Docker environment, or set up user-specific Docker customization. Triggers: 'extend docker', 'docker-extend', 'add tools to docker', 'customize docker', 'add my tools to the container', 'personalize docker setup', 'docker user setup', 'install…

coleam00/Archon · 118 tokens

write-zot-themes

Help the user create, install, or package zot themes, including theme-only extensions.

patriceckhart/zot · 23 tokens

st-full-workflow

Use when the user asks to run the complete end-to-end Strikethroo workflow for a work order in one shot in this repository — triggers include full workflow, end-to-end, plan and execute, do everything, run the whole strikethroo workflow. Do not use when the user wants only one stage (create a plan, generate tasks, or…

e0ipso/strikethroo · 90 tokens

st-code-review

Use when the blueprint execution gate asks for an independent second-harness review of a Strikethroo plan's cumulative diff in this repository — triggers include code review gate, review the plan diff, second-model review, CODEREVIEW hook, review the cumulative diff. Do not use to review a single task, to review code…

e0ipso/strikethroo · 87 tokens

st-refine-plan

Use when the user asks to review, refine, improve, interrogate, pressure-test, or update an existing Strikethroo plan by plan ID in this repository — triggers include refine plan, improve plan, review plan, red-team the plan, update plan. Do not use to create a new plan, to generate tasks, or for generic brainstorming…

e0ipso/strikethroo · 81 tokens

tlamatini-daily-chat-test

Run the daily automated Tlamatini chat regression — drive a visible Chrome via Playwright, log into agentpage.html, ask up to 1000 curated safe questions one-by-one (Multi-Turn ON, ACPX/Ask-Execs/Exec-Report/Internet OFF), wait for and qualify each answer (heuristic + LLM judge on failures), then write a dated report…

XAIHT/Tlamatini · 125 tokens