test-hypothesis

A targeted codebase analysis workflow for checking whether a specific suspicion about a project is supported by evidence.

In plain words
What is it for?
Use it to test concrete hypotheses about code behavior, dependencies, or likely causes of a problem.
Why use it?
It focuses the investigation on confirming or denying one claim instead of performing a broad, unfocused review.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/quangphu1912/codebase-analyzer/test-hypothesis
Any agent
npx skills add quangphu1912/codebase-analyzer --skill test-hypothesis
Clone the repo
git clone --depth 1 https://github.com/quangphu1912/codebase-analyzer

Made for: Claude Code, Codex.

Per session 32 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,017 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00032 $0.01017
Opus 5 $0.00016 $0.00508
Sonnet 5 $0.00006 $0.00203
Haiku 4.5 $0.00003 $0.00102

Measured yesterday against content hash 383ac68afab9, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

test-hypothesis scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/test-hypothesis/SKILL.md · 86 lines

How it starts

The opening of the file, as written. The whole thing — 86 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Announce at start: "Using codebase-analyzer to test hypothesis: [hypothesis]."

Overview

User states a hypothesis. Plugin runs targeted analysis focused on confirming or denying it. This is hypothesis-driven analysis, not exploratory scanning.

Prerequisite Bypass

This skill carries its own prerequisite resolution. It can invoke ANY skill (Track A or Track B) regardless of normal phase prerequisites. Trade-off: invoking Track B without Phase 1 produces shallower analysis, but still produces a valid verdict. Check .state and note which prerequisites were unavailable so the user can interpret confidence levels correctly.

When a prerequisite skill was skipped or unavailable, explicitly note this in the verdict:

  • Which prerequisite was missing
  • How it affects the depth of evidence gathered
  • Whether running the prerequisite first would likely change the verdict

Process

  1. Parse hypothesis: Extract the specific claim and what evidence would confirm/deny it
  2. Select relevant skills: Which Track A/B skills address this hypothesis?
  3. Run targeted analysis: Only invoke skills relevant to the hypothesis
  4. Gather evidence: Collect file:line references that support or contradict
  5. Render verdict: CONFIRMED / DENIED / PARTIALLY CONFIRMED / INCONCLUSIVE
  6. Present evidence: For each piece of evidence, explain how it relates to the hypothesis

Example Hypotheses

Hypothesis Skills to Run Evidence to Find
"This app sends data to third parties" tech-stack, deps, trace-data-flows Outbound HTTP calls to non-first-party domains, data flow paths to external sinks
"There's a hidden admin panel" architecture, api-surface, extract-tool-graph, map-feature-gates Routes gated by role, undocumented endpoints, tool nodes not reachable from main UI
"This code was decompiled, not original" provenance, build-pipeline Sourcemap artifacts, decompiled patterns
"Feature X is coming but not yet released" dead-code, extract-tool-graph, map-feature-gates Dead code behind feature flags, unreleased API endpoints, gated tool nodes
"This system can do more than it exposes" extract-tool-graph, map-feature-gates, prompt-influence Tools defined but gated, capabilities hidden behind config
"Two modules depend on the same hidden contract" detect-hidden-contracts Implicit interfaces, shared assumptions between modules
"What was this system originally designed to do?" reconstruct-system-intent, provenance Design traces, architectural intent recovered from structure

Read the full file on GitHub · 86 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 86 lines · 32 tokens per session scan A 383ac68afab9

Subscribe to this mod's changes

test-hypothesis is a skill published in the GitHub repository quangphu1912/codebase-analyzer (2 stars, last pushed 4mo ago), licensed MIT. It adds 32 tokens to every session and 1,017 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

agent-host-chat-contributions

Build and review cross-cutting agent-host chat behavior through lifecycle contributions. Use when adding turn lifecycle side effects, prompt or context injection, restored-history transformation, protocol-action observation, or when reviewing changes that add code to AgentSideEffects or AgentService.

microsoft/vscode · 56 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens