role-verify

role-verify is a skill for Claude Code, Codex from ryan-scheinberg/harness. It costs 41 tokens per session (913 once invoked), scanned A, original, MIT.

A session role that independently checks whether a claimed coding artifact is actually finished. It examines the artifact and its tests, tries realistic edge cases, and reports a verdict without changing the work.

In plain words
What is it for?
Use it before declaring an artifact complete, opening a pull request, deploying, or publishing it.
Why use it?
It reduces the risk of reporting work as done when it is incomplete or only works on the normal path.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/ryan-scheinberg/harness/role-verify
Any agent
npx skills add ryan-scheinberg/harness --skill role-verify
Clone the repo
git clone --depth 1 https://github.com/ryan-scheinberg/harness

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for role-verify

README.md
[![agentmods](https://agentmods.dev/badge/skills/ryan-scheinberg/harness/role-verify.svg)](https://agentmods.dev/skills/ryan-scheinberg/harness/role-verify)
Your own site
<a href="https://agentmods.dev/skills/ryan-scheinberg/harness/role-verify"><img src="https://agentmods.dev/badge/skills/ryan-scheinberg/harness/role-verify.svg" alt="Measured on agentmods" height="20"></a>
Per session 41 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 913 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00041 $0.00913
Opus 5 $0.00020 $0.00456
Sonnet 5 $0.00008 $0.00183
Haiku 4.5 $0.00004 $0.00091

Measured 5d ago against content hash 8b41f828b333, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

role-verify scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

codex/roles/role-verify/SKILL.md · 57 lines

How it starts

The opening of the file, as written. The whole thing — 57 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are Verify. Your single job is to answer honestly: Is this actually done?

The parent hands you a task and a pointer to the claimed artifact. Exercise that artifact with whatever tools and skills fit the domain, probe edge cases the parent might have skipped, and return a tight, honest verdict

How you work

  • You report, you do not fix. Never edit, propose fixes, or speculate about causes. The parent decides what to do
  • You lean on skills. Before inventing a check, mention the relevant skill (e.g. claude-api for Anthropic SDK code)
  • You probe edge cases. If you can think of a realistic input or scenario that would break the artifact — empty, null, boundary, concurrent, malformed, default vs exception paths, blast radius — try it
  • You do not fake confidence. If "done" cannot be verified with available tools (requires live human judgment, production traffic, a real customer), say so explicitly

Not QA

role-qa gates a whole batch: it stands the assembled system up in a dev environment and breaks it at the seams, and its verdict gates the deploy. You check one artifact, fast — verify in the code and its tests, no environment stood up, back in minutes. One artifact, one verdict; the assembled running system is QA's pass

Domain playbook

  • Code: run tests, typecheck, lint; read the diff; confirm it addresses the stated brief; try edge inputs; for bugfixes, reproduce the original scenario against the fix
  • Infra (Terraform/OpenTofu, Akamai, K8s, Fargate, BigQuery): tofu validate / tofu plan; use akamai / kubectl / gcloud / bq to diff declared vs actual; hit the resulting endpoint or rule; check default and exception paths; check blast radius
  • Slice / brief completion: re-read PROJECT_BRIEF.md / SLICES.md; enumerate each acceptance criterion; confirm honestly satisfied (not just "tests pass")
  • Marketing: invoke the project's marketing skill; compare the draft; flag generic AI-sounding lines, tone drift, missing hooks, misaligned claims
  • Docs / skills: re-read in full context; confirm the change closes the stated gap without breaking flow or leaving stale references

Read the full file on GitHub · 57 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 57 lines · 41 tokens per session scan A 8b41f828b333

Subscribe to this mod's changes

role-verify is a skill published in the GitHub repository ryan-scheinberg/harness (2 stars, last pushed 1mo ago), licensed MIT. It adds 41 tokens to every session and 913 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

gleanin

Rules distilled from past session friction (errors, edit loops, corrections), analyzed by a stronger model and applied to guide future work. Regenerated from glean/artifacts/ plus an index of trigger-loaded protocols from GLEANPROTOCOLDIRS by sync.sh — load every session via CLAUDE.md.

pcx-wave/glean · 62 tokens

exploratory-data-analysis

Perform bounded, local exploratory analysis of explicitly supported scientific files. Use for redacted CSV/TSV/JSON profiles; optional NumPy, HDF5, FASTA/FASTQ, and basic image metadata inspection; missingness/leakage audits; outlier and transformation sensitivity; and rigorous EDA report scaffolds. Other domain…

K-Dense-AI/scientific-agent-skills · 83 tokens

golden-rss

Use when testing the rss golden build.

yusufkaraaslan/Skill_Seekers · 12 tokens

data-charts-tako

Search and visualize the world's data - get charts, insights, and embeddable knowledge cards for finance, economics, demographics, sports, and more.

gooseworks-ai/goose-skills · 35 tokens

google-ads-audit

Google Ads account audit and business context setup. Run this first — it gathers business information, analyzes account health, and saves context that all other ads skills reuse. Trigger on "audit my ads", "ads audit", "set up my ads", "onboard", "account overview", "how's my account", "ads health check", "what should…

nowork-studio/notfair-plugin · 114 tokens

webhook-management

Configure and validate CCAM webhook targets across supported chat, incident, automation, and generic providers. Use when listing provider requirements, creating or updating a target, scoping it to alert rules, sending a test notification, reviewing delivery history, or deleting a target.

hoangsonww/Claude-Code-Agent-Monitor · 56 tokens