Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add lukasrepublic/agentic-foundry --skill sd-plan-testsgit clone --depth 1 https://github.com/lukasrepublic/agentic-foundryWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/lukasrepublic/agentic-foundry/sd-plan-tests)<a href="https://agentmods.dev/skills/lukasrepublic/agentic-foundry/sd-plan-tests"><img src="https://agentmods.dev/badge/skills/lukasrepublic/agentic-foundry/sd-plan-tests.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00133 | $0.02401 |
| Opus 5 | $0.00067 | $0.01201 |
| Sonnet 5 | $0.00027 | $0.00480 |
| Haiku 4.5 | $0.00013 | $0.00240 |
Grade A, and why
sd-plan-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 155 lines — stays where its author put it; the contents beside it link to each section on GitHub.
sd-plan-tests — derive a per-AC test plan before implementation (SDLC step 5)
The software-delivery step sequence (a documented procedure this skill family forms — no workflow engine or state-machine file ships)'s step 5 is test-planning: the craft step that sits
between front-authorization (the ACs are frozen) and implementation (code is written).
Its job is to enumerate, before code exists, the behavioral cases each authorized AC implies —
so the downstream test suite the merge floor's CI runs has real behavioral coverage instead of
accidental coverage reverse-engineered from whatever the implementation happened to do.
This is a PROCEDURE craft skill: a deterministic enumeration procedure with a structured,
parseable output. It is ADVISORY — it advises the trusted operator; it is NOT a gate and
NOT a defense against the operator. It catches the omission mistake (an AC whose
edge/negative/baseline cases were never enumerated); it does not enforce that you follow the plan,
nor police a hostile operator. Current reality (named honestly, not silently papered over): the
completeness accounting below is presently an author self-check against the enumerated rules,
not a machine-verified predicate — the v0.25.0 test-suite realignment's doctor-thinning (2,900 → 255 lines) retired the
the drop-in per-check selftest + its foundry-doctor.py --sd-plan-tests-selftest
registration along with the whole drop-in-check registry, and (unlike several other checks) this one
was not yet ported to the tests/ pytest suite. Applying the accounting below is still the
right procedure; it is a self-check, not a gate, until a pytest port lands. The judgment
(are these the RIGHT cases? are the assertions meaningful?) is trusted craft this skill does not
claim to verify.
When to trigger
- After
/foundry:authorizehas frozen an atom'sacceptance-contract.yaml(spec_sha256 + contract_sha256 written) and before dispatching its implementation. - "plan the tests for
<atom>", "derive the test plan", or SDLC step 5 ofsoftware-delivery.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 155 lines · 133 tokens per session scan A 971691eedc1b
sd-plan-tests is a skill published in the GitHub repository lukasrepublic/agentic-foundry (1 stars, last pushed 2d ago), licensed MIT. It adds 133 tokens to every session and 2,401 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
journey-simulation
Use when caller wants to observe how a stranger encounters a flow, artifact, or sandbox — triggers like "simulate a user journey", "test our onboarding / checkout / signup", "will my ICP convert", "how does a cold reader experience this README", "first-time user test", "cognitive walkthrough", or any request to…
are-you-done
A completion checker for coding work that asks for evidence before allowing an assistant to say a task is finished.
yolo-verify
Use to check a feature's work against its successcriteria and record the result. Produces verification.md and, on pass, the YOLO-Verified trailer. Triggers on "verify this", "does it meet the criteria", or as the verify step of yolo-feature.
test-architect
Acts as a Principal Test Architect to plan, write, prune, run, and hand off automated tests for any codebase in any language or framework — never touching production code, never committing, never running against production. Use whenever the user wants to add, fix, expand, refactor, or review tests for a feature, bug…
scenarios-from-requirements
Write brutally thorough, fully-traceable test scenarios from a requirement — a Jira/Confluence (or Trello/Linear/Azure DevOps/GitHub Issues) source of truth — BEFORE any test code, then a self-contained HTML/CSV coverage report. This is the requirements-first counterpart to the Test Architect: it trusts the spec and…
g-review
Run the review gate on the current branch diff. Runs the test suite, captures the diff, and dispatches code-lead, which verifies done conditions and reviews the diff itself. Issues MERGE READY or HOLD.