vulnerability-validation

A security investigation workflow for finding and validating previously unknown vulnerabilities in a code repository. It examines whether a suspected flaw is reachable, exploitable, and relevant to the defined threat model.

In plain words
What is it for?
Use it to investigate new vulnerability leads or discover flaws directly in code. It supports attack-surface analysis, proof-of-concept testing, reachability checks, and replayable evidence.
Why use it?
Security findings can be false positives or lack enough evidence to act on. This guides careful validation, reproducible proof, risk assessment, and minimal fixes.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/omkhar/vulnerability-validation-skill/vulnerability-validation
Any agent
npx skills add omkhar/vulnerability-validation-skill --skill vulnerability-validation
Clone the repo
git clone --depth 1 https://github.com/omkhar/vulnerability-validation-skill

Made for: Claude Code, Codex.

Per session 149 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 7,950 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00149 $0.07950
Opus 5 $0.00075 $0.03975
Sonnet 5 $0.00030 $0.01590
Haiku 4.5 $0.00015 $0.00795

Measured 2d ago against content hash 039d31bb8cae, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

vulnerability-validation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

The scan reads SKILL.md. This mod also ships 4 executable files (scripts/discovery_engine.py, scripts/quarantine_extract.py, scripts/surface_tool_inventory.py, …), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agent/skills/vulnerability-validation/SKILL.md · 450 lines

How it starts

The opening of the file, as written. The whole thing — 450 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Vulnerability Validation

Use this workflow to discover new vulnerabilities (0-days) in a target repository and validate them. By default it runs native discovery: with no findings list it analyzes the code to find vulnerabilities itself; a scanner finding, researcher report, or PoC is only an optional seed. Reduce every candidate to a defensible disposition, fix true issues with minimal idiomatic patches, and leave evidence another engineer can replay.

Quick Start

For target reads, PHASE-0 starts with an external operator-controlled verifier authenticating the whole package outside the target and neutralizing project-local discovery before startup; the loaded package cannot attest itself. Without both proofs, stop at documentation/package compatibility with no target read. After proof and operator authorization, record run-state.bootstrap.json, then search AI-use/security policy, threat model/boundaries, disclosure process, and upstream sources. Collect target repository/branch, latest upstream HEAD, baseline SHA, dirty state, finding source type, frozen snapshot or seed sources/counts, artifact root, local-only validation environment, dependency installation/provisioning method, upstream URL, repo commands, and output mode.

Read references/portable-invariants.md and create run-state.json before triage. For the default native discovery run, load references/discovery-intake.md; optionally run scripts/discovery_engine.py for a first pass,. Load references/artifact-contract.md when needed: references/artifact-index.md, references/controlled-fields.md, references/closure-bundle.md, references/drift-gate.md, references/evidence-bundle.md. Read references/review-lenses.md before claiming review consensus and references/policy-friction.md when the LLM refuses or hedges.

Read the full file on GitHub · 450 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 450 lines · 149 tokens per session scan A 039d31bb8cae

Subscribe to this mod's changes

vulnerability-validation is a skill published in the GitHub repository omkhar/vulnerability-validation-skill (5 stars, last pushed 9d ago), licensed Apache-2.0. It adds 149 tokens to every session and 7,950 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

microsoft/vscode · 51 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens