Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/juanklagos/spec-driven-development-template/specgit clone --depth 1 https://github.com/juanklagos/spec-driven-development-templateWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00037 | $0.00355 |
| Opus 5 | $0.00018 | $0.00178 |
| Sonnet 5 | $0.00007 | $0.00071 |
| Haiku 4.5 | $0.00004 | $0.00036 |
Grade A, and why
spec scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Create or refine a spec bundle. Respond in the user's language (EN/ES).
- If
$ARGUMENTSnames a new feature: run./scripts/new-spec.sh "<feature>" "<owner>"(sidecar:./spec/scripts/new-spec.sh). If it names an existing spec (e.g.,002), open that bundle. - Work the bundle in order:
spec.md: user story, Given/When/Then scenarios, EARS acceptance criteria (WHEN [trigger], THE SYSTEM SHALL [observable behavior]), requirements, out of scope. Keep it 1-3 pages.plan.md: technical approach consistent with the spec — every requirement covered, nothing outside scope.tasks.md: checklist tasks, test tasks before implementation tasks (TDD).history.md: record every scope/requirement change with date.
- Quality bar (see
docs/en/12-tdd-and-bdd-how-to-write-specs.md): no vague words without measurable values; every criterion maps to at least one test task. - If the spec is ready, ask the user to approve it and record the approval (status, date, approver, evidence) in
spec.md. - Update
specs/INDEX.md(status/priority/date) and point to/sdd:gateas the next step.
Hard stop: no implementation here. One active spec at a time.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 19 lines · 37 tokens per session scan A ff9e60e27d09
spec is a command published in the GitHub repository juanklagos/spec-driven-development-template (14 stars, last pushed yesterday), licensed MIT. It adds 37 tokens to every session and 355 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
paul:help
Show available PAUL commands and usage guide.
safe-refactor
Safe refactoring with automated review, testing, and rollback capabilities.
test
Smart test runner with filtering, coverage, and health monitoring.
paul:verify
Guide manual user acceptance testing of recently built features.
drift
Post-implementation spec drift check — verify the implementation matches existing OpenSpec specifications.
expect
Diff-aware AI browser testing — reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP, ARIA-tree-first) with pass/fail reporting. Use when testing UI changes, verifying PRs before merge, or running regression checks on…