Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add 0dayInc/pwn --skill test_case_enginegit clone --depth 1 https://github.com/0dayInc/pwnWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/0dayinc/pwn/test_case_engine)<a href="https://agentmods.dev/skills/0dayinc/pwn/test_case_engine"><img src="https://agentmods.dev/badge/skills/0dayinc/pwn/test_case_engine/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/0dayinc/pwn/test_case_engine"><img src="https://agentmods.dev/badge/skills/0dayinc/pwn/test_case_engine.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00026 | $0.00669 |
| Opus 5 | $0.00013 | $0.00334 |
| Sonnet 5 | $0.00005 | $0.00134 |
| Haiku 4.5 | $0.00003 | $0.00067 |
Grade A, and why
pwn-ai-redteam-testcaseengine scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 52 lines — stays where its author put it; the contents beside it link to each section on GitHub.
PWN::AI::RedTeam::TestCaseEngine
AI RedTeam Module used to execute PWN::AI::RedTeam::* modules against a target LLM / AI engine. Each attack payload is dispatched to the target model, the raw response is captured, and an independent judge (PWN::AI::Agent::Reflect) scores the response for vulnerability exposure so results roll straight into PWN::Reports::AIRedTeam. ATTACKER vs TARGET SPLIT ------------------------ The engine driving the attack (payload generation + judging) does NOT have to be the engine under test. A frontier model can red-team a local one: PWN::AI::RedTeam::PromptInjection.scan( attacker_engine: :anthropic, attacker_model: 'opus-4.8', target_engine: :ollama, target_model: 'qwen-3.6:latest' ) When neither is passed, both default to PWN::Env[:ai][:active] (the model attacks itself). ADAPTIVE TEST-CASE GENERATION ----------------------------- When PWN::Env[:ai][:module_reflection] == true the strategy-generated seed payloads from each RedTeam module are only round 0. After every round the attacker engine is handed the (payload, response, severity) history and asked to synthesise a fresh batch of payloads specific to the OWASP-LLM / ATLAS category under test. The loop halts on the FIRST deterministic condition met: 1. A finding at or above :stop_on_severity is produced (default CRITICAL) 2. :plateau_rounds consecutive adaptive rounds yield nothing >= MEDIUM 3. :max_adaptive_rounds is exhausted 4. The attacker returns no novel payloads (all duplicates of history) Because the halt is a pure function of the recorded severities / payload set, replaying the same responses reproduces the same stop.
When to use
Call PWN::AI::RedTeam::TestCaseEngine from pwn_eval when the task needs this module.
Do not reimplement it in shell.
Methodologies
Generated from pwn/ai/red_team/test_case_engine.rb. Prefer the public class methods below.
Class methods take (opts = {}) and read opts.
How to call
PWN::AI::RedTeam::TestCaseEngine.help
PWN::AI::RedTeam::TestCaseEngine.execute(opts)
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 52 lines · 26 tokens per session scan A ed3ab0674695
pwn-ai-redteam-testcaseengine is a skill published in the GitHub repository 0dayInc/pwn (81 stars, last pushed today), licensed MIT. It adds 26 tokens to every session and 669 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
race
Race condition / TOCTOU playbook — limit overrun (one-time codes used twice, gift cards spent twice), single-packet attack (last-byte sync) to force parallel processing, and state-confusion races (file upload + read, order before payment). Use when timing-sensitive logic could be abused — one-time codes, coupons/gift…
exploiting-insecure-deserialization
Identifying and exploiting insecure deserialization vulnerabilities in Java, PHP, Python, and .NET applications to achieve remote code execution during authorized penetration tests.
testing-for-business-logic-vulnerabilities
Identifying flaws in application business logic that allow price manipulation, workflow bypass, and privilege escalation beyond what technical vulnerability scanners can detect.
testing-for-sensitive-data-exposure
Identifying sensitive data exposure vulnerabilities including API key leakage, PII in responses, insecure storage, and unprotected data transmission during security assessments.
owasp-top-10-testing
Test an application against the OWASP Top 10 with Strix — autonomous AI agents that attempt real exploits for each category of the current OWASP Top 10:2025 (broken access control including SSRF, security misconfiguration, software supply chain failures, cryptographic failures, injection, insecure design…
application-security-testing
Application security testing (AppSec) across a whole product with Strix — decide which asset needs which test (source code, running web app, API, CI pipeline), run it, and turn the results into a ranked remediation plan. Autonomous agents exploit and prove each issue instead of emitting static-analysis alerts, so the…