run

A task-running workflow that records a goal, checks progress against defined requirements, and uses separate validation steps to gather evidence before closing the task.

In plain words
What is it for?
Use it for implementation work that needs an explicit plan, validation gates, fix tracking, and evidence that each requirement is complete.
Why use it?
It helps prevent tasks from being marked complete based only on an unverified claim that the work is finished.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/chafoo/anchored/run
Any agent
npx skills add chafoo/anchored --skill run
Clone the repo
git clone --depth 1 https://github.com/chafoo/anchored

Made for: Claude Code, Codex.

Per session 118 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,858 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00118 $0.01858
Opus 5 $0.00059 $0.00929
Sonnet 5 $0.00024 $0.00372
Haiku 4.5 $0.00012 $0.00186

Measured 2d ago against content hash a2c431c9a464, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

run scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugin/skills/run/SKILL.md · 143 lines

How it starts

The opening of the file, as written. The whole thing — 143 lines — stays where its author put it; the contents beside it link to each section on GitHub.

/a:run — one task, one run file, evidence-gated

You are the working session AND the orchestrator. The one thing you never do: author evidence. Only the spawned validator calls anchored evidence / anchored fail — that role separation is what makes a checked-off criterion mean proven, not claimed.

All run-file access goes through the anchored CLI over Bash (JSON envelopes; see references/api.md). Never Write/Edit a live run file directly — that would bypass the evidence invariant and the atomic writes.

Communication style

Plain, short, outcome-first. Say what was proven, what failed and why, what happens next. No ceremony, no re-narrating the loop.

Pre-flight

  1. anchored version — if command not found, the plugin isn't enabled or bin/anchored is missing (dev checkout: npm --prefix core run bundle:plugin). Never npm i -g.
  2. If the argument matches an existing slug (anchored status lists them), RESUME: read anchored status <slug> and continue at the open/failed criteria below.
  3. Read anchored.yml (if present) for the declared setups and their instructions. No file = defaults; that's fine. If the project has never used anchored, offer /a:setup once — don't push it.

Anchor

One breath, no planning ceremony. From the user's words + whatever plan already exists:

  • Plan: if a plan-mode plan was just accepted, or a spec/ticket/chat plan is in context, copy it VERBATIM into plan — it is the immutable record of what was asked. No plan source → omit plan, the goal carries the run.
  • Rigor: from the user's own words — "keep it simple" → light, default standard, "must be clean / important" → high, "release-critical" → max. The choice is visible in the file and correctable.
  • Criteria: testable derivations of the plan — phrased so a validator can later prove or refute each one. Tag each criterion with the setup that knows how to verify it (a /a:run frontend … argument is a tagging hint, not a field); no fitting setup → leave it untagged (defaults).
  • Judgment: phrase criteria so something CAN be run against them — an execution is the sharpest proof, whatever the subject (a test, a render, a request, a checksum). Where the subject genuinely has nothing to execute (the copy reads calm, the asset matches the brand sheet, the solution follows the pattern), mark judgment: true. It is a note to the reader, not an escape hatch: "looks right in the browser" is not judgment — drive the browser. A validator can never award itself that mark.
  • Gates: slice them yourself, sized to the rigor (light: one final gate · standard: by risk · high: fine-grained · max: one gate per criterion). A gate is setup-homogeneous — slice along setup boundaries too.

Read the full file on GitHub · 143 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 143 lines · 118 tokens per session scan A a2c431c9a464

Subscribe to this mod's changes

run is a skill published in the GitHub repository chafoo/anchored (3 stars, last pushed 1mo ago), licensed MIT. It adds 118 tokens to every session and 1,858 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

doubt-driven-review

In-flight adversarial check on a non-trivial decision BEFORE it stands — distinct from post-hoc review of a finished diff. Use on "stress-test this decision", "are we sure about this", "verify before commit", "poke holes in this", when working in unfamiliar code, or before an irreversible step (migration, prod deploy…

BlackBeltTechnology/pi-agent-dashboard · 90 tokens

release-cut

Cut a new pi-agent-dashboard release: promote ## [Unreleased] in CHANGELOG.md, bump every workspace package.json per SemVer, commit, tag v , and push — triggering the Release workflow that publishes every non-private workspace, builds the Electron artifacts, and creates a GitHub Release. Use on "cut a release"…

BlackBeltTechnology/pi-agent-dashboard · 93 tokens

ship-it

Worktree-side implementation orchestrator for an OpenSpec change. Idempotent: gates automated scenarios on filesystem reality, owns the red-test fix loop, runs the docker harness with always-teardown, then drives ship-change inline. Escape hatch writes SHIPITBLOCKED.md. Runnable headless. Triggers: "ship it", "build…

BlackBeltTechnology/pi-agent-dashboard · 92 tokens

faq-mine

Mine docs/faq.md from README.md, docs/.md, and the pi-hermes memory stores. Dispatches @fast subagents per source, dedupes against the existing FAQ, and merges entries in caveman style. Use when asked to "build / regenerate / extend the FAQ", "mine docs into FAQ", "mine hermes memory into FAQ", "surface runtime…

BlackBeltTechnology/pi-agent-dashboard · 94 tokens

session-to-guideline

Turn a pi session into a Markdown "how-we-did-it" collaboration guideline: reads the session's JSONL transcript and synthesizes a reusable playbook of which prompts worked, what had to be steered, and how to reproduce the result faster. Use when: "document this session", "write up how we did X with the AI", "make a…

BlackBeltTechnology/pi-agent-dashboard · 93 tokens

scenario-design

Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec. Derives edge-case, performance, frontend-quirk and error-handling scenarios with ISTQB techniques, routes each to a test level, and writes test-plan.md, emitting clarification questions on a spec gap. Use on "design test scenarios", "what…

BlackBeltTechnology/pi-agent-dashboard · 88 tokens