Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/pierry/harness-kit/designgit clone --depth 1 https://github.com/Pierry/harness-kitWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00035 | $0.00664 |
| Opus 5 | $0.00017 | $0.00332 |
| Sonnet 5 | $0.00007 | $0.00133 |
| Haiku 4.5 | $0.00003 | $0.00066 |
Grade A, and why
design scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Design a system. Follow .claude/agents/system-architect/guides/pipeline.md for retry and approval, and .claude/shared/pipeline-pattern.md for inputs (resolve-mark-proceed) and eval (adversarial).
Inputs are resolved, not asked (resolve-mark-proceed). Resolve scale target (users / QPS / data /
latency SLO), internal-vs-web-scale, and known constraints from any provided description plus
context-library/. Infer a sensible order-of-magnitude scale and mark it ASSUMPTION: {x}; mark
genuine unknowns NOT FOUND - NEEDS REVIEW: {detail} and keep going. The one permitted stop: if NO
system or problem statement exists at all, ask once for the one-liner. Everything else resolves from
context. Never stop to ask for a resolvable input.
Route to a topic skill if one fits. Check
.claude/agents/system-architect/skills/{topic}/SKILL.md (currently: search-engine, url-shortener,
rate-limiter). If the problem matches, use that skill's reference architecture. Otherwise use the generic skill
.claude/agents/system-architect/skills/design/SKILL.md.
Compute feature_id = {YYYY-MM-DD}-{slug}.
Read:
- .claude/agents/system-architect/guides/design-method.md
- .claude/agents/system-architect/guides/writing-style.md
- .claude/agents/system-architect/guides/templates/system-design.md
- .claude/agents/system-architect/guides/pipeline.md
- .claude/agents/system-architect/guides/examples/good-system-design-example.md
- the matched topic skill, if any
Save to .claude/runtime/outputs/architect/design/{feature_id}.md.
Sensors: .claude/agents/system-architect/sensors/design-structure.md, .../sensors/design-rigor.md. Evals: .claude/agents/system-architect/evals/design-quality.md.
Run the evals adversarially: dispatch a fresh evaluator via the Task tool (subagent_type: general-purpose) that did not author this design. Hand it only the artifact path and the one rubric path; it scores against the rubric and reports the weighted total plus the low-scoring dimensions. Below threshold (8.0) retries per pipeline.md, regenerating only the flagged dimensions.
After save, reply with this exact shape (name the actual skill/sensors/evals/guides that ran):
Design saved at {path}. Score: {N}/10.
skill: {search-engine | design (generic)}
sensors: design-structure ok, design-rigor ok
eval: design-quality {N}/10 (attempts: N)
guides: design-method.md, writing-style.md, templates/system-design.md
canon: {names applied}
next: /system-design:review
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 47 lines · 35 tokens per session scan A 3203891d402f
design is a command published in the GitHub repository Pierry/harness-kit (3 stars, last pushed 1mo ago), licensed MIT. It adds 35 tokens to every session and 664 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
setup-pm-skills
Onboard a new user — find out what they do, recommend the right bundles & top skills, and set up a project CONTEXT.md so every skill is tailored to them.
statusbar-style
Switch the status-bar style (classic / capsule / hairline).
fest-show
Show festival progression (in-progress tasks, roadmap, and dependency view).
superpowers-execute
Execute the current GSD phase plan with Superpowers instead of gsd-execute-phase.
config
Command "config" from sdebruyn/fabric-dw-mcp-cli, covering configuration & defaults, http retry budget, sql retry budget, mcp workspace allowlist {#mcp-workspace-allowlist} and mcp server log level.
deps-age
Analyze dependency freshness and maintenance activity.