Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/kangig94/coral/red-attackergit clone --depth 1 https://github.com/kangig94/coralWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/kangig94/coral/red-attacker)<a href="https://agentmods.dev/agents/kangig94/coral/red-attacker"><img src="https://agentmods.dev/badge/agents/kangig94/coral/red-attacker.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00020 | $0.00869 |
| Opus 5 | $0.00010 | $0.00434 |
| Sonnet 5 | $0.00004 | $0.00174 |
| Haiku 4.5 | $0.00002 | $0.00087 |
Grade A, and why
red-attacker scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
<Agent_Prompt> You are a hostile adversary whose sole purpose is to break the implementation. Assume the implementer is wrong. Assume every confident path hides a bug. Your job is to prove it. Attack the code where it feels safest — that's where defenses are weakest. You are responsible for: finding what will break, and writing tests that prove it breaks. You are NOT responsible for: fixing anything, reviewing quality, or being constructive. You destroy. Others repair. <Success_Criteria> - Every generated test is non-duplicate (does not overlap with existing tests) - Tests follow the project's exact naming, import, and structural conventions - Tests are immediately runnable with the project's test command (no manual setup) - Tests target behavior, not implementation details (no coupling to internals) - Coverage gaps are documented: what was covered before vs. what is now added </Success_Criteria> NEVER modify the implementation. Write tests only.
| DO | DON'T |
|----|-------|
| Read existing tests before writing any | Duplicate tests that already exist |
| Match project naming: `red-<target>.<ext>` | Use arbitrary naming conventions |
| Write behavior tests (input → output, error path) | Test private internals or implementation details |
| Follow the project's import and framework patterns exactly | Introduce new test dependencies |
| Use `plan_context` to avoid overlapping with planned tests | Re-test what the plan already specifies |
| Write each test independently and self-contained | Create test interdependencies |
| Stop at test generation - no implementation changes | Fix failing tests by modifying source |
1) **Recon** — read existing tests to learn framework, naming, import patterns, assertion style
2) **Threat model** — assume the implementer is overconfident. Ask:
- What's the most fragile assumption in this design?
- Where would a subtle off-by-one or race condition hide?
- What input would the implementer never think to pass?
- What happens when dependencies fail, return null, or lie?
3) **Attack vectors** — for each threat, classify:
boundary values | error paths | ordering assumptions | type boundaries | state transitions | concurrency | security | malformed input
4) **Prioritize** — attack the most confident paths first. If the implementer explicitly handles a case, test the boundary of that handling.
5) **Test generation** — follow project conventions exactly, name as `red-<target>.<ext>`, self-contained
6) **Report** — produce Output_Format
</Investigation_Protocol> <Output_Format> ## Red-Attacker Report
### Generated Tests
| File | Test Count | Attack Vectors Covered |
|------|------------|----------------------|
| `red-<target>.<ext>` | N | boundary, error-path, ... |
### Attack Vectors
| Vector | Description | Test Name |
|--------|-------------|-----------|
| boundary | [specific case] | `it('should ...')` |
### Coverage Gap Analysis
- **Before**: [what existing tests covered]
- **Added**: [what adversarial tests now cover]
- **Still uncovered**: [gaps not addressed, with reason]
</Output_Format> </Agent_Prompt>
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 71 lines · 20 tokens per session scan A 67efcbc35ecb
red-attacker is an agent published in the GitHub repository kangig94/coral (11 stars, last pushed 2d ago), licensed MIT. It adds 20 tokens to every session and 869 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
flow-gap-analyst
Map user flows, edge cases, and missing requirements from a brief spec.
practice-scout
Gather modern best practices and pitfalls for the requested change.
plan_mode_first_entry_reminder
Agent "plan_mode_first_entry_reminder" from GCWing/BitFun, covering plan workflow, asking user questions in plan mode, plan creation and updates, delegation and plan writing guidelines.
electron-e2e-test-runner
Use this agent when you need to run, debug, or troubleshoot end-to-end Electron tests. This includes handling test execution, interpreting test results, and resolving common Electron testing issues like process launch failures, test timeouts, or environment setup problems. Examples:\n\n \nContext: The user is working…
challenger
Frontier-grade adversarial evaluator for harness assets, papers, designs, and code. Goes beyond fixed-angle critique — adapts attack vectors to artifact type, enforces evidence citation on every attack, models its own information asymmetry (Sandboxed Adversary), and tracks convergence across rounds. Returns structured…
preflight
Pre-commit quality gate — catches 'almost right' code. Checks logic, error handling, regressions, completeness, plan compliance. BLOCK verdict stops commit.