ai red teaming agents

20 tagged ai red teaming, measured the same way as everything else here.

Browse within: ai-security 20jailbreak 18llm-evaluation 18

false-green-hunter

01

guardana/guardana

Agent Claude Code

Read-only adversarial reviewer for Guardana. Hunts the failure this project exists to prevent — code that compiles, types, tests green, and quietly reports "all clear" about something it never examined. Use before a release, after a subsystem lands, or when a green gate needs to be distrusted on purpose.

89 2d ago A 69 tokens original Apache-2.0

gate-runner

02

guardana/guardana

Agent Claude Code

Runs Guardana's full gate — lint, format, strict types, import contract, tests with coverage floors, dogfood, generated docs and the three isolated example suites — and reports what actually passed. Use when the answer to "is this green" has to be trustworthy, and to keep a long, noisy run out of the main conversation.

89 2d ago C 73 tokens original Apache-2.0

nuguard-test-writer

03

NuGuardAI/nuguard

Agent Claude Code

Use this agent when you need to write real-world integration or unit tests for NuGuard's key capabilities (SBOM generation, analysis, policy, redteam, CLI, configuration, etc.). Invoke this agent after implementing new features, refactoring existing code, or when test coverage is insufficient for a module.\n\n…

36 2d ago A 407 tokens

e2e-runner

04

NuGuardAI/nuguard

Agent

End-to-end testing specialist using Playwright. Use PROACTIVELY for generating, maintaining, and running E2E tests. Manages test journeys, quarantines flaky tests, uploads artifacts (screenshots, videos, traces), and ensures critical user flows work.

36 2d ago A 59 tokens

security-reviewer

05

NuGuardAI/nuguard

Agent

Security vulnerability detection and remediation specialist. Use PROACTIVELY after writing code that handles user input, authentication, API endpoints, or sensitive data. Flags secrets, SSRF, injection, unsafe crypto, and OWASP Top 10 vulnerabilities.

36 2d ago A 52 tokens