Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add andrewstellman/quality-playbook --skill quality-playbook-harnessgit clone --depth 1 https://github.com/andrewstellman/quality-playbookWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/andrewstellman/quality-playbook/quality-playbook-harness)<a href="https://agentmods.dev/skills/andrewstellman/quality-playbook/quality-playbook-harness"><img src="https://agentmods.dev/badge/skills/andrewstellman/quality-playbook/quality-playbook-harness/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/andrewstellman/quality-playbook/quality-playbook-harness"><img src="https://agentmods.dev/badge/skills/andrewstellman/quality-playbook/quality-playbook-harness.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00157 | $0.02254 |
| Opus 5 | $0.00078 | $0.01127 |
| Sonnet 5 | $0.00031 | $0.00451 |
| Haiku 4.5 | $0.00016 | $0.00225 |
Grade A, and why
quality-playbook-harness scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 151 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Quality Playbook harness
You are the harness orchestrator. Your entire per-tick job is small and
fixed: run one Python script, dispatch the worker subagents it lists,
print the table it formats, and schedule the next tick. All the
state-machine logic lives in bin/qpb_harness_tick.py — you never reason
about run state yourself. (Details: references/STATE_MACHINE.md.)
The worker contract (the whole of it): a job is anything that appends
JSON lines to a file. status is the only field the harness
interprets; label (a short free string shown in the ACTIVITY column),
message, and the opaque data object are displayed but never read. The
contract honors Postel's law — conservative in what the harness emits,
liberal in what it accepts: a worker that never writes, dies, or writes
garbage degrades to a visible STALLED / failed / LAUNCH-FAIL row, never to
a wedged state machine; a malformed line is skipped with a warning, never
fatal. Heartbeats are schema_version: "2" (label/data); the reader
still accepts v1 (phase/step). You never read or transcribe a path:
every harness-known path — including {HARNESS_BIN} — is substituted
mechanically by the engine before dispatch (FR-21a). Pass each worker
prompt verbatim; it is already fully resolved.
Capability ladder — probe, announce, degrade (do this FIRST)
The harness degrades along two axes; the disk state machine is identical at every rung. At startup, PROBE your own tooling and ANNOUNCE the rungs you selected, in one line to the operator:
- Cadence (how the next tick happens): rung 1 = you have an in-session
scheduling primitive (
ScheduleWakeup); rung 2 = an OS scheduler; rung 3 = the foregroundharness_ticker.pyloop; rung 4 = manual ticks. - Dispatch (how workers start): rung 1 = in-session subagents
(
Task/Agent); rung 2 = detached host-CLI processes (dispatch_mode: "shell").
As a Claude Code session you run at cadence 1 + dispatch 1: you have
ScheduleWakeup and a subagent tool, and your session persists across the
workers' lifetime. Announce that: *"Harness: cadence rung 1 (ScheduleWakeup)
- dispatch rung 1 (subagent). Plan has N entries, pool P."* If the plan's
entries are
dispatch_mode: "shell", you cannot run them in-session — tell the operator to drive the run with the ticker (the printed command below) and stop.
What ships with it
7 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 151 lines · 157 tokens per session scan A 167efe466f24
quality-playbook-harness is a skill published in the GitHub repository andrewstellman/quality-playbook (84 stars, last pushed yesterday), licensed Apache-2.0. It adds 157 tokens to every session and 2,254 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
recipe-create-meet-space
Create a Google Meet meeting space and share the join link.
workthreads
SpecStory Workthreads - a weekly work-thread rollup across a team's repos from SpecStory coding histories (any agent - Claude Code, Codex, Cursor, Gemini, and more). It groups the window's sessions into threads of work per project and labels each new / open / recently closed, so a lead sees what shipped, what is still…
atmos-config
Atmos root configuration: atmos.yaml discovery, precedence, deep merging, basepath, imports, minimal bootstrap, and routing to narrower Atmos skills.
story-readiness
Validate that a story file is implementation-ready. Checks for embedded GDD requirements, ADR references, engine notes, clear acceptance criteria, and no open design questions. Produces READY / NEEDS WORK / BLOCKED verdict with specific gaps. Use when user says 'is this story ready', 'can I start on this story', 'is…
autotask-creator
Rules for automation CRUD from the group-chat commander. The commander does not call mutation tools and does not edit cloud/autotasks files directly. It emits one or more top-level ... containers in its final text; the bus parses and applies them after the turn.
projects
List all managed projects with status, branch, open PRs, and open issue counts — portfolio-level view.