verify-behavior

verify-behavior is a skill for Claude Code, Codex from nicknisi/dotfiles. It costs 68 tokens per session (1,824 once invoked), scanned A, original, MIT.

A procedure for checking whether a product behaves correctly by using its real user interface. It covers reproducing reported bugs and verifying implemented changes with recorded interface state.

In plain words
What is it for?
Use it to reproduce interface bugs, confirm feature fixes, and review visible workflows such as browser interactions.
Why use it?
It replaces assumptions based only on source code with evidence from the running application.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/nicknisi/dotfiles/verify-behavior
Any agent
npx skills add nicknisi/dotfiles --skill verify-behavior
Clone the repo
git clone --depth 1 https://github.com/nicknisi/dotfiles

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for verify-behavior

README.md
[![agentmods](https://agentmods.dev/badge/skills/nicknisi/dotfiles/verify-behavior.svg)](https://agentmods.dev/skills/nicknisi/dotfiles/verify-behavior)
Your own site
<a href="https://agentmods.dev/skills/nicknisi/dotfiles/verify-behavior"><img src="https://agentmods.dev/badge/skills/nicknisi/dotfiles/verify-behavior.svg" alt="Measured on agentmods" height="20"></a>
Per session 68 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,824 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00068 $0.01824
Opus 5 $0.00034 $0.00912
Sonnet 5 $0.00014 $0.00365
Haiku 4.5 $0.00007 $0.00182

Measured 4d ago against content hash ad5bb96cc460, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

verify-behavior scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

home/.pi/agent/skills/verify-behavior/SKILL.md · 122 lines

How it starts

The opening of the file, as written. The whole thing — 122 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Verify behavior

Prove or disprove visible product behavior for bugs and greenfield features using pi-computer-use. Triage, implementation, and review stages invoke this instead of guessing from code alone.

Modes

  • reproduce — does the reported bug still happen on baseline? (usually default branch; triage)
  • verify — does the implemented change match expected behavior? (features and fixes; implementation/review)

Infer if unnamed: issue-only → reproduce; implementation/PR branch → verify.

Platform contract (do not invent alternatives)

pi-computer-use exposes a direct, state-scoped tool surface, not a delegated capability and not raw screen capture. The normal loop:

Step Tools
Find find_roots returns ranked @r roots (desktop windows and CDP browser pages share one forest). launch_browser for a managed CDP page.
Observe observe_ui captures a root and returns a folded outline, @e refs, and a stateId. Every later @e use requires its owning stateId.
Query search_ui, expand_ui, inspect_ui, and read_text query the cached state without re-capturing. Refine broad searches instead of paging matches.
Act act_ui performs checked, transactional steps and returns the successor stateId. Attach expect when the action has an observable completion signal.
Wait wait_for for asynchronous UI changes; navigate_browser / evaluate_browser only on CDP page states.

Consume the successor stateId from act_ui directly; observe again only after an uncertain external mutation or state eviction.

Read the full file on GitHub · 122 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 122 lines · 68 tokens per session scan A ad5bb96cc460

Subscribe to this mod's changes

verify-behavior is a skill published in the GitHub repository nicknisi/dotfiles (2,987 stars, last pushed yesterday), licensed MIT. It adds 68 tokens to every session and 1,824 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

coding

Principles for writing and designing code, covering API and abstraction design, naming, comment discipline, and standards of evidence for claims about code behavior. Use whenever the task is to write or modify code, design an API or software architecture, or review code.

ryota2357/dotfiles · 53 tokens

assess-quality

Foundational quality framework: the five questions (readable, easy to start, expands without bloat, consistent, intentional) every other dev skill is judged against, plus the dual-audience and workshop principles. Use when onboarding to a project, defining a quality bar, setting an assessment checklist, or arbitrating…

urmzd/dotfiles · 120 tokens

create-oss-skill

Create well-formed Agent Skills following the agentskills.io specification. Scaffold directories, write SKILL.md files, bundle scripts, and structure instructions for progressive disclosure. Use when creating a new skill, reviewing skill structure, optimizing a skill description, or setting up evals for skill quality.

urmzd/dotfiles · 63 tokens

extend-oss-skills-to-claude

Extend standard agentskills.io skills with Claude Code-specific features. Invocation control, subagent execution, dynamic context injection, string substitutions, model/effort overrides, and deployment scoping. Use when adapting a portable skill for Claude Code, adding Claude-specific frontmatter, setting up subagent…

urmzd/dotfiles · 74 tokens

orchestrate-agents

Orchestrate multiple agent CLIs (Claude, Codex, Antigravity) via tmux with a shared fleet store, dispatching one guardian subagent per pane. Survey-first: inspects and adopts existing tmux sessions, windows, and agent panes before creating anything new. Use when running a multi-agent session, dispatching parallel…

urmzd/dotfiles · 83 tokens

scaffold-project

Generates cross-language standard files (README, AGENTS.md, LICENSE, CONTRIBUTING.md, SECURITY.md, sr.yaml, .envrc, llms.txt), documentation conventions, and project structure, then dispatches to language-specific scaffolds. Use first for cross-language standard files and structure, THEN load the matching scaffold…

urmzd/dotfiles · 137 tokens