Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/stefan-stepzero/shipkit/shipkit-stagenpx skills add stefan-stepzero/shipkit --skill shipkit-stagegit clone --depth 1 https://github.com/stefan-stepzero/shipkitWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00030 | $0.05303 |
| Opus 5 | $0.00015 | $0.02652 |
| Sonnet 5 | $0.00006 | $0.01061 |
| Haiku 4.5 | $0.00003 | $0.00530 |
Grade A, and why
shipkit-stage scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 465 lines — stays where its author put it; the contents beside it link to each section on GitHub.
shipkit-stage — Project Stage & Graduation Criteria
Purpose: Define the project stage (POC/Alpha/MVP/Scale), derive scope constraints and quality bars from that stage, set business-metric criteria (S-*), and define stage gates. Evaluate mode assesses gate readiness for human-approved graduation.
What it does: Reads the project vision and codebase signals, grounds the stage from cited signals, derives constraints and business metrics, defines gates with S-* criteria, and writes goals/strategic.json. Runs in fork context — when stage is grounded by cited signals, proceeds autonomously; when stage is genuinely ungrounded, emits NEEDS_ELICITATION:shipkit-stage and pauses rather than guessing silently. Evaluate mode is user-invoked and human-gated — it cross-references ALL goal files to build a graduation evidence table for human approval.
Protocol: This skill follows the canonical elicitation protocol defined in install/shared/references/elicitation-protocol.md (the mechanics — marker, state files, resume). The steps below are this skill's specific application of that protocol.
Calibration: Apply install/shared/references/ground-or-ask-calibration.md (the intelligence — propose vs ask). Ground first: propose every stage field you can tie to a cited signal (the opening prompt, why.json stage/currentState/approach/constraints keywords, codebase maturity signals — file counts, CI/CD config, specs, tests, deployment infrastructure), tagged with its source; flag low-leverage guesses as guessed. For shipkit-stage the high-leverage decisions are: (1) the stage itself (POC / Alpha / MVP / Scale) when no cited signal resolves it, and (2) graduation scope constraints when stage is genuinely ambiguous. These set hard scope boundaries — a wrong stage silently misaligns every constraint, quality bar, and success criterion derived from it. When stage is grounded by a clear cited signal, proceed autonomously (no marker). Only emit NEEDS_ELICITATION:shipkit-stage when stage is genuinely ungrounded after the full inference pass (no explicit stage field, no unambiguous keywords, no codebase maturity signal points clearly at one stage).
Output: One JSON file:
goals/strategic.json— Stage, constraints, stageImplications, business-metric criteria, gates
Product goals (user outcomes: P-) are handled by
/shipkit-product-goals. Engineering goals (technical performance: E-) are handled by/shipkit-engineering-goals. Both skills append their criteria IDs to the gates defined here.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 465 lines · 30 tokens per session scan A db5406ff82bc
shipkit-stage is a skill published in the GitHub repository stefan-stepzero/shipkit (1 stars, last pushed 1mo ago), licensed MIT. It adds 30 tokens to every session and 5,303 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
opencli-sitemap-author
Use when creating or maintaining OpenCLI site sitemaps: agent-facing navigation, page-state, action, workflow, API-reference, pitfall, and fallback knowledge for a website. Use after browser exploration discovers durable site context, when a sitemap is stale, or when promoting local site knowledge into the repo.
golden-rss
Use when testing the rss golden build.
omh-buzz
This is a Hermes-native buzz workflow skill.
omh-code-review
This is a Hermes-native code-review workflow skill.
redteam-api-detail-pack
Domain routing and boundary guidance for authorized API security testing, including BOLA/IDOR, authentication bypass, mass assignment, missing rate limits, and GraphQL issues. Use when a task belongs to the API testing domain and needs scope, evidence, pivot, or exit criteria.
redteam-postex-detail-pack
Domain routing and boundary guidance for authorized post-exploitation testing after initial access, including privilege escalation, persistence, lateral movement, data collection, and cleanup considerations. Use when a task belongs to the post-exploitation domain and needs scope, evidence, pivot, or exit criteria.