Spec Kitty is an open-source command-line tool that turns product requirements into a repository-based workflow for AI-assisted software development. It stores specifications, plans, tasks, acceptance criteria, reviews, and merge decisions in Git while giving agents isolated git worktrees for parallel implementation. The catalogue add-ons support the project's workflows for coordinating agents and governing their work.
Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Priivacy-ai/spec-kitty --skill adversarial-squadgit clone --depth 1 https://github.com/Priivacy-ai/spec-kittyWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/priivacy-ai/spec-kitty/adversarial-squad)<a href="https://agentmods.dev/skills/priivacy-ai/spec-kitty/adversarial-squad"><img src="https://agentmods.dev/badge/skills/priivacy-ai/spec-kitty/adversarial-squad/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/priivacy-ai/spec-kitty/adversarial-squad"><img src="https://agentmods.dev/badge/skills/priivacy-ai/spec-kitty/adversarial-squad.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00160 | $0.01157 |
| Opus 5 | $0.00080 | $0.00579 |
| Sonnet 5 | $0.00032 | $0.00231 |
| Haiku 4.5 | $0.00016 | $0.00116 |
Grade A, and why
adversarial-squad scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 83 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Adversarial Squad Deployment (harness)
The operational HOW. The doctrinal WHEN/WHY is the procedure
adversarial-squad-deployment (packs/built-in/procedures/adversarial-squad-deployment.procedure.yaml),
which sits under the brownfield-onboarding paradigm. This skill changes no mission
type or guard; it is a technique the orchestrator opts into.
When to use
A squad is worth its tokens at a high-leverage point-cut where one reviewer's blind spot is expensive:
- after
/spec-kitty.specify→ pre-spec investigation (scope, prior art, live repros) - after
/spec-kitty.plan→ post-planning brownfield check (foldable issues, split-brain, deprecations) - after
/spec-kitty.tasks→ post-tasks anti-laziness pass (fakeable DoDs, decomposition realism) - before merge → architectural-gate / cross-base sweep
- ad-hoc decision → proponent + adversaries + synthesizer (e.g. delete-vs-migrate)
Do NOT use it as a rubber stamp, and do NOT wire it as a mandatory gate.
The recipe
- Frame one sharp question and pick the point-cut. A squad answers a question; it is not a vibe check.
- Select 3–4 distinct profiles by lens (bounded). Complementary, not redundant:
architect-alphonso— structure / seams / topologydebugger-debbie— live-evidence, coverage, "would this catch the regression?"reviewer-renata— anti-laziness, contract-vs-implementation, fakeable assertionsrandy-reducer— duplication / dead code (⚠ duct-tape bias — read critically)paula-patterns— decomposition, boundaries, second-opinion adjudicationplanner-priti— scope, sequencing, tracker hygienepython-pedro— implementer feasibilitydoctrine-daphne— doctrine integrity / DRG wiring Scale past 4 only for an explicit "audit / comprehensive" ask.
- Dispatch in parallel, profile-LOADED. Each delegate's prompt MUST begin with:
"FIRST run
spec-kitty agent profile show <id>andspec-kitty charter context --action <action> --json; apply the resolved initialization, boundaries, directives, and tactics, then state which you applied." Loading the profile — not naming a persona — is the point. Only a read-only harness that cannot invoke the CLI may readpacks/built-in/agent_profiles/<id>.agent.yaml; that degraded fallback can diverge because overlays,specializes_fromlineage, andenhances/overridessemantics are not applied. Keep delegates read-only unless the task is an isolated implementation in its own worktree. - Require structured, non-fakeable output. Each returns findings as
[SEVERITY] file:line — issue — recommendation, ending in a verdict, grounded in cited evidence, with honest concession of where its lens does not apply. A steelman that over-claims is weak; an adversary that concedes nothing is noise. - Match model tier to difficulty. Strong tier for analytical/adversarial lenses; lighter tier for mechanical/tracker delegates.
- Synthesize; second-opinion on divergence. Aggregate. Where delegates disagree on a consequential point, do NOT average — adjudicate from the source, or dispatch one focused second-opinion delegate. Be critical of any delegate with a known bias. If irreconcilable, escalate to the human with both positions.
- Record the convergent evidence and act. Capture confirmed findings (artifact, findings doc, or memory). The value is convergent evidence that survived independent scrutiny — not a single opinion.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 83 lines · 160 tokens per session scan A fb089048f502
adversarial-squad is a skill published in the GitHub repository Priivacy-ai/spec-kitty (1,603 stars, last pushed 2d ago), licensed MIT. It adds 160 tokens to every session and 1,157 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
clarity-over-cleverness
Apply clarity-over-cleverness rewrites — prefer code a junior engineer can read at a glance over compact-but-clever code. Use during /build's simplify step and during /code-simplify (alias "simplify the code"). Never weakens behavior; suite must remain green.
spring-code-review-rubric
Pre-commit code review rubric for Spring Boot 4 changes. Used by spring-code-reviewer to produce 08-code-review.md before any commit. Covers traceability, architecture, Spring idioms, error handling, data access, security, test quality, clarity, and migration.
6_gofer_validate
Validate implementation with 10-category engineering rubric (100 points).
enforcement-audit
Run a compliance audit against a technology instruction file, detecting discrepancies, planning workstreams, implementing fixes, and validating quality gates.
spec-to-code-compliance
Check code against the documentation that specifies it - which requirements hold, which the code contradicts, which are absent, and what the code does that no document mentions. Use when comparing an implementation against a whitepaper, protocol spec, or design document.
code-review
Run a structured Spec Kit review focused on code compliance, documentation quality, or test coverage, positioned within the spec-driven development pipeline.