verified-delivery

verified-delivery is a skill for Claude Code from Caxson/proof-of-work. It costs 74 tokens per session (581 once invoked), scanned A, original, Apache-2.0.

A workflow for proving that software changes work by checking each step against real services, data, screens, or devices.

In plain words
What is it for?
Use it when building or fixing web interfaces, backends, databases, APIs, or mobile features that need evidence such as screenshots, real API calls, database checks, or device tests.
Why use it?
It prevents a task from being marked complete based only on code changes or mocked results.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin. Also seen: mentions subagents.

Part of the proof-of-work plugin — 2 skills, 3 commands, 1 hook shipped together

Good fit Use it when building or fixing web interfaces, backends, databases, APIs, or mobile features that need evidence such as screenshots, real API calls, database checks, or device tests.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/caxson/proof-of-work/verified-delivery
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Caxson/proof-of-work --skill verified-delivery
Clone the repo
git clone --depth 1 https://github.com/Caxson/proof-of-work

Made for: Claude Code.

Or install proof-of-work, the plugin that ships this one along with the rest of its 2 skills, 3 commands, 1 hook.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for verified-delivery

README.md
[![agentmods](https://agentmods.dev/badge/skills/caxson/proof-of-work/verified-delivery/github.svg)](https://agentmods.dev/skills/caxson/proof-of-work/verified-delivery)
Your own site
<a href="https://agentmods.dev/skills/caxson/proof-of-work/verified-delivery"><img src="https://agentmods.dev/badge/skills/caxson/proof-of-work/verified-delivery/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for verified-delivery

Your own site · 80×15
<a href="https://agentmods.dev/skills/caxson/proof-of-work/verified-delivery"><img src="https://agentmods.dev/badge/skills/caxson/proof-of-work/verified-delivery.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 74 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 581 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00074 $0.00581
Opus 5 $0.00037 $0.00291
Sonnet 5 $0.00015 $0.00116
Haiku 4.5 $0.00007 $0.00058

Measured 12d ago against content hash 562575bc338f, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

verified-delivery scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/verified-delivery/SKILL.md · 51 lines

How it starts

The opening of the file, as written. The whole thing — 51 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Verified Delivery

Turn the request into a verification-backed chain. Do not claim success without proof.

Procedure

  1. Locate & read. Scope the relevant code and logs — use search to bound the area; do NOT read an entire large repo. Understand the real interaction logic. If a design or assumption looks wrong, STOP and confirm with the user before building.

  2. Decompose, and attach a verification to each step. For every step, decide up front HOW it will be proven, by surface:

    • Web / UI → drive it with Playwright and screenshot the result.
    • Data / backend → write a script that calls the REAL service, then query the DB/SQL to confirm the write actually landed.
    • Blocked by hardware/device → close the loop via API first, and tell the user this is a stand-in for the real device test.
    • Mobile → drive the device and screenshot.
    • Always use REAL data. Never mock data unless the user explicitly asks for mock data.
  3. Confirm the plan (for non-trivial work). Send the user the step plan plus how each step will be verified, before writing code. Obvious or small tasks may proceed directly with a one-line note.

  4. Build and verify step by step. Each step must be proven by a REAL action on REAL data — not a simulated action and not an asserted success.

  5. On repeated failure (three or more). Stop. Give a full cause analysis that FIRST rules out your own bug. Do NOT fake a pass and do NOT degrade or work around the problem. Re-plan from a new angle.

  6. End-to-end proof. When everything is done, run one end-to-end verification. Save the evidence chain + screenshots to a TEMP directory (never the user's home root). In chat, give only a compact table: step → how verified → result → evidence path. The task is "done" ONLY after the user confirms.

Leverage

  • Dispatch an INDEPENDENT subagent for adversarial verification — let it try to DISPROVE that a step works, rather than confirm it.
  • For multi-step or parallel verification, orchestrate it as a workflow.

Read the full file on GitHub · 51 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 51 lines · 74 tokens per session scan A 562575bc338f

Subscribe to this mod's changes

verified-delivery is a skill published in the GitHub repository Caxson/proof-of-work (2 stars, last pushed 3mo ago), licensed Apache-2.0. It adds 74 tokens to every session and 581 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

research-engineer

An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.

davila7/claude-code-templates · 43 tokens

tika-eval-compare

Compare extracts from two Tika builds over a corpus to detect regressions in content, encoding, exceptions, and embedded-document handling. Use for "compare before/after extracts", "eval this change against the corpus".

apache/tika · 50 tokens

neuron-evaluation-engineer

Create and run AI evaluations with datasets, assertions, and output drivers in Neuron AI. Use this skill whenever the user mentions evaluation, testing AI systems, creating evaluators, dataset-driven testing, assertion-based validation, or wants to measure AI system performance. Also trigger for tasks involving…

neuron-core/neuron-ai · 77 tokens

jetson-validate-image

Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.

NVIDIA/skills · 50 tokens

atmos-validation

Validate Atmos projects, components, arbitrary JSON Schema inputs, EditorConfig, and GitHub Actions; use affected-file selection and native CI annotations.

cloudposse/atmos · 31 tokens

skill-benchmark

Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.

HoangNguyen0403/agent-skills-standard · 16 tokens