tailtest-hunt

An adversarial test-planning command for a specified source file. It examines the code and creates a separate test file with scenarios designed to expose failures at boundaries, during errors, under unusual inputs, or under heavy use.

In plain words
What is it for?
Use it to generate 8–12 targeted stress scenarios for one file, such as invalid formats, type mix-ups, timing problems, partial failures, resource exhaustion, or off-by-one errors. It writes the scenarios using the project’s expected hunt-test naming convention.
Why use it?
It encourages deliberate attempts to break code instead of checking only expected behavior. Keeping these tests separate lets you review them before adding them to the main test suite.

Command

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/avansaber/tailtest/tailtest-hunt
Clone the repo
git clone --depth 1 https://github.com/avansaber/tailtest
Per session 0 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 533 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.00533
Opus 5 $0.00000 $0.00267
Sonnet 5 $0.00000 $0.00107
Haiku 4.5 $0.00000 $0.00053

Measured 2d ago against content hash 6d028481e674, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

tailtest-hunt scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

commands/tailtest-hunt.md · 31 lines

What it actually says

Run an adversarial pass on $ARGUMENTS -- explicitly try to break the source code.

Read the source file at $ARGUMENTS. Generate 8-12 adversarial test scenarios drawn from the R15 categories in CLAUDE.md (boundary inputs, format / injection, type confusion, concurrent state, time / locale edges, error handling under partial failures, resource exhaustion, off-by-one logic). Pick categories that genuinely apply to this file; skip any that do not (and note the skip).

This command bypasses the project's configured depth and forces an adversarial-biased pass on the named file regardless of depth setting in .tailtest/config.json.

Where to write the test file: write to a SEPARATE hunt test file, not the regular test file for this source. Naming convention:

Source file Hunt test file
services/billing.py tests/test_billing_hunt.py
app/Http/Controllers/OrderController.php tests/Feature/OrderControllerHuntTest.php
internal/handler.go internal/handler_hunt_test.go

The hunt file is intentionally separate so it does not contaminate the main test suite. The user decides after review whether to keep, merge into the main test file, or discard.

Step-by-step behavior:

  1. Read the source file at $ARGUMENTS
  2. Output a SCENARIO PLAN with 8-12 adversarial scenarios, each labeled [adversarial: <category>]. State which categories were skipped and why.
  3. Write the test file at the hunt path (see table above)
  4. Run the hunt test file using the configured runner (pytest -q tests/test_<basename>_hunt.py etc.)
  5. For any failure, apply R12 classification (real_bug / environment / test_bug). Report each failing scenario with category and classification: [adversarial: type-confusion] real_bug -- function returns None on int input where str expected.
  6. If all pass: tailtest hunt: {N} adversarial scenarios on {file}, all passed.

Do not auto-fix. Always ask before fixing any real_bug found by hunt.

No update-existing-tests behavior. Hunt always writes to the separate hunt test file. If the hunt file already exists, replace its contents (the user is asking for a fresh hunt).

Treat the file as new-file regardless of git status -- hunt explicitly requests generation even on legacy files or files tailtest would normally skip.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 31 lines · 0 tokens per session scan A 6d028481e674

Subscribe to this mod's changes

tailtest-hunt is a command published in the GitHub repository avansaber/tailtest (11 stars, last pushed 2mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 533 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.