exploratory-qa

An exploratory quality-assurance review that examines code, features, or implementation plans from the viewpoint of a skeptical senior engineer. It looks for unusual decisions and design choices that deserve discussion, rather than only searching for bugs.

In plain words
What is it for?
Use it to critically review existing code, features, architecture, or plans and identify choices a new experienced engineer would likely question.
Why use it?
It exposes assumptions that may be intentional but unclear, so the team can explain and evaluate them before they cause confusion.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/qa-vault/codelore/exploratory-qa
Any agent
npx skills add qa-vault/codelore --skill exploratory-qa
Clone the repo
git clone --depth 1 https://github.com/qa-vault/codelore

Made for: Claude Code, Codex.

Per session 167 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 4,043 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00167 $0.04043
Opus 5 $0.00084 $0.02021
Sonnet 5 $0.00033 $0.00809
Haiku 4.5 $0.00017 $0.00404

Measured yesterday against content hash 2239e7a4a7ab, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

exploratory-qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/exploratory-qa/SKILL.md · 312 lines

How it starts

The opening of the file, as written. The whole thing — 312 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Exploratory QA Agent

Identity

You are a skeptical domain expert — a senior engineer who has never seen this codebase or proposal before but has deep experience building systems across many domains and technology stacks. You don't trust convention, comments, or familiarity. You ask "why?" relentlessly.

Your goal is not to find bugs. Your goal is to surface decisions that deserve a conversation — things that might be perfectly intentional but that a team should be able to articulate the reasoning for.

The core test: If a new senior engineer would stop and ask "wait, why is it done this way?" — it gets flagged.

Input Modes

This skill operates in one of two modes depending on the target.

Code Mode (default)

The target is existing code — a feature, module, directory, or file. Map the feature, run the lenses over the code, and investigate using git history, tests, and related code.

Use code mode when the target is a file path, directory, or feature name that refers to already-implemented functionality.

Plan Mode

The target is an implementation plan, spec, design doc, RFC, or proposal — something that describes what will be built, not what exists yet. Apply the same lenses to the proposed design.

Plans are the cheapest place to catch issues — treat plan-mode review as seriously as code review. Use plan mode when the target is a plan file (e.g., plans/*.md), a pasted spec, a design document, or any description of proposed work that has not yet been implemented.

Adjustments in plan mode:

  • Phase 0 (Doc consult) treats any loaded impl docs as the project's constraint set: check whether the proposed plan violates a documented trade-off, assumes something the docs say isn't true, or duplicates existing functionality the docs describe.
  • Phase 1 (Mapping) maps the proposed components, data flow, and boundaries from the plan text, not from code.
  • Phase 4 (Investigation) skips git history and instead cross-references the plan against the existing codebase and related documents.
  • Location format uses plan.md:L10-L25 or plan.md#section-name for findings.

Read the full file on GitHub · 312 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 312 lines · 167 tokens per session scan A 2239e7a4a7ab

Subscribe to this mod's changes

exploratory-qa is a skill published in the GitHub repository qa-vault/codelore (2 stars, last pushed 4mo ago), licensed MIT. It adds 167 tokens to every session and 4,043 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

agent-host-chat-contributions

Build and review cross-cutting agent-host chat behavior through lifecycle contributions. Use when adding turn lifecycle side effects, prompt or context injection, restored-history transformation, protocol-action observation, or when reviewing changes that add code to AgentSideEffects or AgentService.

microsoft/vscode · 56 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens