Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add haiggoh/run-to-completion --skill triage-for-autonomygit clone --depth 1 https://github.com/haiggoh/run-to-completionWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/haiggoh/run-to-completion/triage-for-autonomy)<a href="https://agentmods.dev/skills/haiggoh/run-to-completion/triage-for-autonomy"><img src="https://agentmods.dev/badge/skills/haiggoh/run-to-completion/triage-for-autonomy/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/haiggoh/run-to-completion/triage-for-autonomy"><img src="https://agentmods.dev/badge/skills/haiggoh/run-to-completion/triage-for-autonomy.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00136 | $0.02375 |
| Opus 5 | $0.00068 | $0.01188 |
| Sonnet 5 | $0.00027 | $0.00475 |
| Haiku 4.5 | $0.00014 | $0.00237 |
Grade A, and why
triage-for-autonomy scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 104 lines — stays where its author put it; the contents beside it link to each section on GitHub.
triage-for-autonomy — score the queue before you touch it
Selection is what makes unattended work safe. The failure mode of autonomous execution is not incompetence; it is taste-dependent guesses made while nobody is watching. This skill prevents that by forcing an explicit, recorded verdict on every item before execution begins — so that nothing gets attempted by default.
The three scoring axes
Apply these three filters to every item in your queue. An item must pass all three to be Tier 1.
- Not gated on a human: Does this require user input, a supervised interactive session, a download, or a fresh session state? If yes, it is gated.
- Objective decision or spelled-out plan: Is the outcome defined by objective criteria (a bug fix, a known-correct doc update, a written test procedure)? If it requires design taste, methodology choice, or creative judgment, it is not objective.
- Verifiable by observable outcome: Can you prove it is done by checking a diff, a test result, or a file state? If success is subjective or internal, it is not verifiable.
The three tiers
Bucket items based on the axes above.
- Tier 1 — do now: Fully autonomous, objective, bounded, and verifiable. Examples: documentation fixes with known-correct outcomes, clear bug fixes, test procedures already written out, read-only investigations, and objectively-closeable items.
- Tier 2 — autonomous but heavy: Delegatable but liable to sprawl. Examples: benchmarks, research, and measurement tasks. Do these only when Tier 1 is exhausted and budget remains.
- Gated — do not start: Needs user input, a supervised session, a download, is a design/taste decision, or needs fresh session state. Gated is a destination for this run, not a permanent verdict — see below.
Gate reason is a required output
For every gated item, record which category blocks it. This is not a note; it is a required field. The gate reason is what makes the closing wrap and any subsequent unblocking pass nearly free. Without it, you must re-analyze the item later to understand why it stopped.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago Changed · +21 lines 1eeb4f57d40c
- 10d ago First seen · 83 lines · 136 tokens per session scan A 90e61c42bd44
triage-for-autonomy is a skill published in the GitHub repository haiggoh/run-to-completion (1 stars, last pushed 3d ago), licensed MIT. It adds 136 tokens to every session and 2,375 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
design
Create a doc-as-code design package from a PRD or SPEC. Conditionally generates C4 diagrams (Context/Container/Component), sequence diagrams, ER diagram + Data Dictionary, OpenAPI 3.0, AsyncAPI 3.0, ADRs, domain glossary, state diagrams, and deployment view as Mermaid-rendered Markdown files. Use when PM mentions…
design-corpus
EXPERIMENTAL (Claude Code only). Apply a SPEC increment to a single LIVING architecture corpus under docs/architecture/ instead of a per-SPEC silo package — treats docs as event-sourcing (SPEC = commit, corpus = working tree), so C4 Context/Container, glossary and the data model stay system-wide and never drift across…
review-pr
A pull-request review skill for evaluating proposed code changes and posting a quality verdict. A pull request is a request to merge changes into a shared codebase.
tasks
A task-breakdown tool that turns a plan, specification, feature brief, bug report, technical-debt item, or chore into small TASK-NNN work items. Technical debt means postponed cleanup or design work in a codebase.
continue
Autonomous work — find and execute ready tasks.
reconcile-docs
Best-effort reconciliation of the living architecture corpus against the code: the model reads docs/architecture/ plus the code and lists divergences with coordinates. Advisory only — an opinion of the model, never a verified verdict, no gates and no stamps. Nothing is applied automatically: corpus edits are a…