task-verify

A command for running the approved checks for a task in progress and saving the results as evidence. TDD means test-driven development, where automated tests are used to check that code behaves as expected.

In plain words
What is it for?
Use it to run a task's tests or other verification command, optionally with a type-check command and time limit, before review.
Why use it?
It replaces an informal rerun and visual inspection with a recorded result tied to the task's allowed files and checks.

Command

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/thixpin/pitway/task-verify
Clone the repo
git clone --depth 1 https://github.com/thixpin/pitway
Per session 15 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 623 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00015 $0.00623
Opus 5 $0.00008 $0.00311
Sonnet 5 $0.00003 $0.00125
Haiku 4.5 $0.00002 $0.00062

Measured 2d ago against content hash 6ed794181e40, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

task-verify scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

src/integrations/claude/commands/task-verify.md · 50 lines

What it actually says

task-verify

Role: Orchestrator

pitway task-verify <id> [--typecheck <command>] [--timeout <ms>] [--json]

Runs an in_progress task's own approved command/tdd verification command (plus optional --typecheck <command> and --timeout <ms> -- 1000..3600000; when omitted, the task's own verification.timeout_ms applies if declared, else 120000) and persists a verification record — a full task_verify_evidence journal entry covering the run's exit code, pass/fail counts, and a fingerprint of the task's declared write_scope/relevant_files. Each record is named by an evidence id (e.g. tve-a1b2c3). This is the formal replacement for an ad hoc independent rerun-and-eyeball of the verification command — see ../protocol-driver.md. It does not replace your own diff/write_scope review, which still happens first, every time.

Must run while the task is in_progress — before pitway task-update <id> review, not after (review's own tasks.yaml rewrite never invalidates a recorded evidence record). If completion later refuses because no recorded record resolves cleanly, pitway task-update <id> in_progress is a legal recovery transition from review — run task-verify again from there to produce a fresh record.

pitway task-update <id> completed then resolves one record — implicitly, the newest one whose own run passed (searched newest-to-oldest, so a later failing re-run never masks an earlier passing record), or explicitly via --evidence <id> (strict: no such search, a failing or stale record named explicitly still refuses) — and validates it (task identity, run success, attempt match, command match, write_scope match, fingerprint match) before completing. On success, its captured evidence unconditionally becomes the persisted result evidence — the final, capped string that lands in task.result.evidence — replacing whatever --result's file carried for that field.

Once any evidence record exists for a task, plain --result/--message completion can no longer bypass it. But a task is never permanently stuck: if no recorded record resolves cleanly at completion, recover with pitway task-update <id> in_progress, then run task-verify again to produce a fresh, valid record.

See ../protocol-driver.md and ../dispatch.md. Run pitway task-verify --help for flags.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 50 lines · 0 tokens per session scan A 6ed794181e40

Subscribe to this mod's changes

task-verify is a command published in the GitHub repository thixpin/pitway (19 stars, last pushed 3d ago), licensed MIT. It adds 15 tokens to every session and 623 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.