test-classification

test-classification is a skill for Claude Code from QBall-Inc/the-bulwark. It costs 14 tokens per session (2,427 once invoked), scanned A, original, MIT.

A prompt template for the first stage of a test-suite audit. It helps a fast, lightweight agent classify test files and identify which ones need deeper review.

In plain words
What is it for?
Use it inside the Test Audit process to categorize test files and flag candidates for analysis. It is an internal pipeline component, not a general-purpose user command.
Why use it?
Analyzing every test file deeply from the start can consume unnecessary time. Initial classification narrows the audit and feeds later mock detection and report-writing stages.

Skill for Claude Code

Written for Claude Code: allowed-tools in frontmatter. Also seen: mentions subagents.

Part of the the-bulwark plugin — 30 skills, 17 agents, 6 hooks shipped together

Good fit Use it inside the Test Audit process to categorize test files and flag candidates for analysis. It is an internal pipeline component, not a general-purpose user command.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/qball-inc/the-bulwark/test-classification
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add QBall-Inc/the-bulwark --skill test-classification
Clone the repo
git clone --depth 1 https://github.com/QBall-Inc/the-bulwark

Made for: Claude Code.

Or install the-bulwark, the plugin that ships this one along with the rest of its 30 skills, 17 agents, 6 hooks.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for test-classification

README.md
[![agentmods](https://agentmods.dev/badge/skills/qball-inc/the-bulwark/test-classification.svg)](https://agentmods.dev/skills/qball-inc/the-bulwark/test-classification)
Your own site
<a href="https://agentmods.dev/skills/qball-inc/the-bulwark/test-classification"><img src="https://agentmods.dev/badge/skills/qball-inc/the-bulwark/test-classification.svg" alt="Measured on agentmods" height="20"></a>
Per session 14 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,427 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00014 $0.02427
Opus 5 $0.00007 $0.01213
Sonnet 5 $0.00003 $0.00485
Haiku 4.5 $0.00001 $0.00243

Measured 7d ago against content hash 49436b034a3e, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-07, from the pricing page.

Security

Grade A, and why

test-classification scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Runs shell commandslowCapability

Expected in a hook, worth knowing in a rule or an instructions file.

| Imports system modules (`child_process`, `fs`, `http`) | Note for risk assessment |
Origin

Copies of this mod

1 near-identical copy found in the catalogue:

skills/test-classification/SKILL.md · 332 lines

How it starts

The opening of the file, as written. The whole thing — 332 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Test Classification

Prompt template for surface-level test classification and triage. Designed for a Haiku sub-agent to quickly categorize test files and flag those needing deep analysis.


When to Use This Skill

This is an internal skill loaded by the orchestrator during Test Audit pipeline.

Context Action
/test-audit invoked Orchestrator loads this skill for Stage 1
Test Audit pipeline triggered by hook Orchestrator loads this skill for Stage 1
Need to classify test files Load directly as prompt template for Haiku

DO NOT use for:

  • Direct user invocation (not user-invocable)
  • Mock detection (use mock-detection skill)
  • Full audit synthesis (use test-audit skill)

Role in Test Audit Pipeline

This skill provides the first stage prompt template:

test-audit (P0.8) orchestrates:
  Stage 1: test-classification (Haiku) → classification YAML
  Stage 2: mock-detection (Sonnet) → violations YAML
  Stage 3: synthesis (Sonnet) → audit report

The orchestrator loads this skill and constructs a 4-part prompt for a general-purpose Haiku sub-agent.


4-Part Prompt Template

GOAL

Classify all test files in {target} by type and flag files needing deep analysis for mock appropriateness.

CONSTRAINTS

  • Do NOT modify any files
  • Classify every test file found
  • Use filename-first classification (content validates)
  • Flag mock+integration mismatches for deep analysis
  • Use AST verification_lines as ground truth when provided in context (do NOT re-count). Only fall back to manual counting if AST data is unavailable.
  • Complete within 30 tool calls

CONTEXT

Target directory: {target}

Test file patterns: *.test.*, *.spec.*, test_*, *.integration.*, *.e2e.*

Classification rules: See "Classification Logic" section below

Deep analysis triggers: See "needs_deep_analysis Triggers" section below

Line counting rules: See "Verification Line Counting" section below

Read the full file on GitHub · 332 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago First seen · 332 lines · 14 tokens per session scan A 49436b034a3e

Subscribe to this mod's changes

test-classification is a skill published in the GitHub repository QBall-Inc/the-bulwark (8 stars, last pushed yesterday), licensed MIT. It adds 14 tokens to every session and 2,427 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 1 finding (runs shell commands). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

journey-simulation

Use when caller wants to observe how a stranger encounters a flow, artifact, or sandbox — triggers like "simulate a user journey", "test our onboarding / checkout / signup", "will my ICP convert", "how does a cold reader experience this README", "first-time user test", "cognitive walkthrough", or any request to…

RockyHong/super-bootstrap · 79 tokens

black-box-test

Use as an INDEPENDENT tester agent to test another agent's intesting ticket from the OUTSIDE — never your own, and never from the implementation diff. You test from the operational test contract + acceptance criteria only (never HOW it was built), writing automated tests that invoke the changed surfaces and recording…

tmj-90/gaffer · 122 tokens

java-conventions

Use when a ticket adds or changes Java code and it must follow the repo's Java conventions — modern Java (records, sealed types, pattern matching, switch expressions), Optional discipline, immutability, Spring Boot constructor injection, and JUnit 5 + Mockito tests. Invoke for "add this in Java", "fix the Java build"…

tmj-90/gaffer · 89 tokens

python-conventions

Use when a ticket adds or changes Python code and it must follow the repo's Python conventions — PEP 8, full type hints, dataclasses, pythonic idioms, explicit error handling, and pytest with coverage. Invoke for "add this in Python", "fix the type/lint errors", "add the FastAPI/Django endpoint", or as the language…

tmj-90/gaffer · 84 tokens

add-integration-test

Use when a ticket asks for integration or end-to-end coverage across components — an API route hitting a database, a service-to-service call, a multi-step flow — rather than a single unit. Invoke for "test the endpoint end to end", "cover the checkout flow", or "verify the migration + query together".

tmj-90/gaffer · 69 tokens

add-unit-test

Use when a ticket asks for new unit tests, or when an acceptance criterion requires test coverage for a function/module/component and none exists. Invoke for "add tests for X", "cover the Y edge case", or raising coverage on a specific unit.

tmj-90/gaffer · 54 tokens