Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/awarexone/agentic-bug-hunter/llm-redteamgit clone --depth 1 https://github.com/Awarexone/Agentic-Bug-HunterWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/awarexone/agentic-bug-hunter/llm-redteam)<a href="https://agentmods.dev/commands/awarexone/agentic-bug-hunter/llm-redteam"><img src="https://agentmods.dev/badge/commands/awarexone/agentic-bug-hunter/llm-redteam.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00630 |
| Opus 5 | $0.00000 | $0.00315 |
| Sonnet 5 | $0.00000 | $0.00126 |
| Haiku 4.5 | $0.00000 | $0.00063 |
Grade A, and why
llm-redteam scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 60 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/llm-redteam
Automated LLM/agentic red-teaming. Instead of hand-firing one prompt-injection at a time, this runs a categorized corpus against a chat endpoint and uses a canary token to reliably detect which payloads succeeded.
Usage
# simple endpoint that takes {"message": "..."}
/llm-redteam --url https://t/api/chat --field message
# only one category
/llm-redteam --url https://t/api/chat --field message --category jailbreak
# OpenAI-style body + nested response path
/llm-redteam --url https://t/api/chat \
--template '{"messages":[{"role":"user","content":"{{PAYLOAD}}"}]}' \
--response-path choices.0.message.content
Run directly:
tools/llm_redteam.py --url https://t/api/chat --field message --json
tools/llm_redteam.py --list-categories
Categories (OWASP LLM Top 10 / ASI)
| Category | Tests | OWASP |
|---|---|---|
prompt-injection |
direct instruction override | LLM01 |
jailbreak |
DAN / developer-mode persona escape | LLM01 |
system-prompt-leak |
extract the hidden system prompt | LLM07 |
data-exfil |
markdown-image beacon to attacker host | LLM02/LLM06 |
indirect-injection |
payload framed as a retrieved document | LLM01 (indirect) |
guardrail-bypass |
base64 / split-instruction filter evasion | LLM01 |
How detection works
Most payloads instruct the model to emit a unique token (RT_PWNED_xxxx). If
that token appears in the response, the injection landed — far more reliable
than keyword matching. System-prompt-leak uses a multi-signal heuristic;
data-exfil confirms when the canary URL is reflected in the output.
--header "Authorization: Bearer ..." (repeatable) for authed chatbots.
Turn a hit into a report
A bare prompt-injection is Informational until chained. Escalate:
injection → chatbot IDOR (read another user's data), data exfil (the
markdown-beacon hit proves a working channel), or RCE if the agent has a
code/tool execution capability. See skills/web2-vuln-classes §11 and
skills/bug-bounty Agentic AI (ASI01–ASI10).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 60 lines · 0 tokens per session scan A 5092b3da9e6e
llm-redteam is a command published in the GitHub repository Awarexone/Agentic-Bug-Hunter (4,689 stars, last pushed 2d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 630 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
recon
Run full recon pipeline on a target — subdomain enum (Chaos API + subfinder), live host discovery (dnsx + httpx), URL crawl (katana + waybackurls + gau), gf pattern classification, nuclei scan. Outputs to recon/ / directory. Usage: /recon target.com.
hunt
Active vulnerability hunting. Two-track dispatcher — asks Red Team vs WAPT, hands off to hunt-dispatch skill and sibling commands. Usage: /hunt target.com | /hunt .target.com | /hunt targets.txt [--vuln-class X] [--source-code P] [--chrome].
token-scan
Meme coin and token security scan — checks for rug pull vectors (hidden mint, honeypot, fee manipulation, LP lock bypass, authority retention, bonding curve exploits, fake renounce, sandwich amplification). Manual 8-class grep audit (with an optional automated scanner if present). Usage: /token-scan [--chain solana].
autopilot
Run autonomous hunt loop on a target — scope check → recon → rank surface → hunt → validate → report with configurable checkpoints. Usage: /autopilot target.com [--paranoid|--normal|--yolo].
triage
Quick 7-Question Gate triage on a finding before writing a report. Kills N/A submissions before they happen. Faster than /validate — for quick go/no-go decisions. Usage: /triage.
validate
Validate a finding — runs 7-Question Gate + 4-gate checklist. Kills weak findings before report writing. Prevents N/A submissions that hurt validity ratio. Usage: /validate.