Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/avansaber/tailtest/tailtest-huntgit clone --depth 1 https://github.com/avansaber/tailtestWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00533 |
| Opus 5 | $0.00000 | $0.00267 |
| Sonnet 5 | $0.00000 | $0.00107 |
| Haiku 4.5 | $0.00000 | $0.00053 |
Grade A, and why
tailtest-hunt scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Run an adversarial pass on $ARGUMENTS -- explicitly try to break the source code.
Read the source file at $ARGUMENTS. Generate 8-12 adversarial test scenarios drawn from the R15 categories in CLAUDE.md (boundary inputs, format / injection, type confusion, concurrent state, time / locale edges, error handling under partial failures, resource exhaustion, off-by-one logic). Pick categories that genuinely apply to this file; skip any that do not (and note the skip).
This command bypasses the project's configured depth and forces an adversarial-biased pass on the named file regardless of depth setting in .tailtest/config.json.
Where to write the test file: write to a SEPARATE hunt test file, not the regular test file for this source. Naming convention:
| Source file | Hunt test file |
|---|---|
services/billing.py |
tests/test_billing_hunt.py |
app/Http/Controllers/OrderController.php |
tests/Feature/OrderControllerHuntTest.php |
internal/handler.go |
internal/handler_hunt_test.go |
The hunt file is intentionally separate so it does not contaminate the main test suite. The user decides after review whether to keep, merge into the main test file, or discard.
Step-by-step behavior:
- Read the source file at
$ARGUMENTS - Output a SCENARIO PLAN with 8-12 adversarial scenarios, each labeled
[adversarial: <category>]. State which categories were skipped and why. - Write the test file at the hunt path (see table above)
- Run the hunt test file using the configured runner (
pytest -q tests/test_<basename>_hunt.pyetc.) - For any failure, apply R12 classification (real_bug / environment / test_bug). Report each failing scenario with category and classification:
[adversarial: type-confusion] real_bug -- function returns None on int input where str expected. - If all pass:
tailtest hunt: {N} adversarial scenarios on {file}, all passed.
Do not auto-fix. Always ask before fixing any real_bug found by hunt.
No update-existing-tests behavior. Hunt always writes to the separate hunt test file. If the hunt file already exists, replace its contents (the user is asking for a fresh hunt).
Treat the file as new-file regardless of git status -- hunt explicitly requests generation even on legacy files or files tailtest would normally skip.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 31 lines · 0 tokens per session scan A 6d028481e674
tailtest-hunt is a command published in the GitHub repository avansaber/tailtest (11 stars, last pushed 2mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 533 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
review-codex
Adversarial security review via OpenAI model — second opinion from a different AI.
sessions
Generate implementation session prompts from a PRD or research doc.
deploy
Command "deploy" from PraveenJayaprakash-JP/repomemory, covering deployment workflow, 1. pre-deploy checks, 2. build verification, 3. environment validation and 4. deployment steps (platform detection).
diagnose
Command "diagnose" from Vimalk0703/shipworthy, covering process, example output, project diagnosis report, critical (fix immediately) and high (fix soon).
plan-review
Deep plan review before implementing any feature, refactor, or significant code change. Enters plan mode - no code changes, just evaluation.
monitor
Monitor repository health, track issues, and identify maintenance tasks.