Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add juliusz-cwiakalski/agentic-delivery-os --skill run-plangit clone --depth 1 https://github.com/juliusz-cwiakalski/agentic-delivery-osWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/juliusz-cwiakalski/agentic-delivery-os/run-plan)<a href="https://agentmods.dev/skills/juliusz-cwiakalski/agentic-delivery-os/run-plan"><img src="https://agentmods.dev/badge/skills/juliusz-cwiakalski/agentic-delivery-os/run-plan/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/juliusz-cwiakalski/agentic-delivery-os/run-plan"><img src="https://agentmods.dev/badge/skills/juliusz-cwiakalski/agentic-delivery-os/run-plan.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00011 | $0.02334 |
| Opus 5 | $0.00005 | $0.01167 |
| Sonnet 5 | $0.00002 | $0.00467 |
| Haiku 4.5 | $0.00001 | $0.00233 |
Grade A, and why
run-plan scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 194 lines — stays where its author put it; the contents beside it link to each section on GitHub.
<discovery_rules>
Locate change folder: search doc/changes/**/*--<workItemRef>--*/
If not found, search for spec file: doc/changes/**/chg-<workItemRef>-spec.md
Plan file: chg-<workItemRef>-plan.md inside the change folder.
Spec file: chg-<workItemRef>-spec.md for change.type and slug validation.
Folder pattern: doc/changes/YYYY-MM/YYYY-MM-DD--<workItemRef>--<slug>/
Never reference doc/changes/current/ in edits or evidence.
</discovery_rules>
<branch_rules>
<change.type>/<workItemRef>/<slug>
1. If current branch matches pattern → keep. 2. Else attempt checkout existing branch. 3. Else create new branch. 4. Abort if dirty working tree contains unrelated staged changes.
</branch_rules>
<directive_parsing> Supported (case-insensitive):
- "execute next N phases" → run current incomplete phase + N−1 subsequent.
- "execute all remaining phases" → run all incomplete phases.
- "execute phase N" → run only phase N (must be incomplete).
- "ask for review" / "and then ask for review" → pause after requested phases (default).
- "no review" / "continue without review" → do not pause.
- "dry run" → simulate; no edits/commits.
- "commit per task" → commit after each task instead of per phase.
Defaults (no directive): phasesToRun=1; askForReview=true; commitMode=per-phase. </directive_parsing>
<plan_parsing_rules>
Phase header regex: /^### Phase (\d+):/
Tasks: checkboxes under "Tasks:" until blank line or "Acceptance Criteria:"
Unchecked: - [ ]; Completed: - [x]
Completion: replace token only, append note e.g. (done: added config & tests)
Acceptance criteria: lines start with "- Must:" or "- Should:"; append (PASSED: <summary>) or (FAILED: <summary>)
Execution Log: header "## Execution Log"; append if missing
</plan_parsing_rules>
<core_principles> Determinism: parsing & updates never reorder tasks or phases. Minimality: edit only necessary lines; avoid broad formatting changes. Traceability: every edit tied to task completion; commits atomic. Autonomy: proceed without prompting unless blocked. Idempotence: re-running resumes cleanly without duplicating evidence. </core_principles>
<phase_execution_rules> For each selected phase:
- Identify pending tasks (unchecked). If none but acceptance evidence missing, treat as completion-only phase.
- For each pending task:
a. Form internal contract (goal, inputs, outputs, success checks).
b. Discover relevant files in: src/, app/, packages/, modules/, lib/, services/, infra/, config/, scripts/, tests/, static/, doc/.
c. Implement minimal edits; avoid unrelated refactors.
d. Add/adjust tests when functional behavior changes.
e. Run quick validations (typecheck/build/test subset).
f. Mark task completed with concise evidence note.
g. If commitMode=per-task: stage only task changes + plan update, then
/commitusing actual values forworkItemRef, the task's specificoutcome, supportedwhyfrom the spec/plan, and observedverification. - After tasks complete:
a. Run full quality gates; capture PASS/FAIL summaries.
b. Append evidence to acceptance criteria lines (once only).
c. Append Execution Log entry.
d. If commitMode=per-phase:
/commitusing actual values forworkItemRef, the completed work's specificoutcome, supportedwhyfrom the spec/plan, and observedverification. - Stop after phasesToRun. If askForReview=true, pause with summary. </phase_execution_rules>
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 194 lines · 11 tokens per session scan A f051d4326eda
run-plan is a skill published in the GitHub repository juliusz-cwiakalski/agentic-delivery-os (38 stars, last pushed yesterday), licensed MIT. It adds 11 tokens to every session and 2,334 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
Verification & Quality Assurance
Comprehensive truth scoring, code quality verification, and automatic rollback system with 0.95 accuracy threshold for ensuring high-quality agent outputs and codebase reliability.
agent-github-modes
Agent skill for github-modes - invoke with $agent-github-modes.
workflow-automation
Workflow creation, execution, and template management. Automates complex multi-step processes with agent coordination. Use when: automating processes, creating reusable workflows, orchestrating multi-step tasks. Skip when: simple single-step tasks, ad-hoc operations.
deployment-pipeline-design
Design multi-stage CI/CD pipelines with approval gates, security checks, and deployment orchestration. Use this skill when designing zero-downtime deployment pipelines, implementing canary rollout strategies, setting up multi-environment promotion workflows, or debugging failed deployment gates in CI/CD.
gitlab-ci-patterns
Build GitLab CI/CD pipelines with multi-stage workflows, caching, and distributed runners for scalable automation. Use when implementing GitLab CI/CD, optimizing pipeline performance, or setting up automated testing and deployment.
bazel-build-optimization
Optimize Bazel builds for large-scale monorepos. Use when configuring Bazel, implementing remote execution, or optimizing build performance for enterprise codebases.