llm evaluation agents

19 tagged llm evaluation, measured the same way as everything else here.

Browse within: ai-red-teaming 18ai-security 18jailbreak 18

nuguard-test-writer

01

NuGuardAI/nuguard

Agent Claude Code

Use this agent when you need to write real-world integration or unit tests for NuGuard's key capabilities (SBOM generation, analysis, policy, redteam, CLI, configuration, etc.). Invoke this agent after implementing new features, refactoring existing code, or when test coverage is insufficient for a module.\n\n…

36 yesterday A 407 tokens

e2e-runner

02

NuGuardAI/nuguard

Agent

End-to-end testing specialist using Playwright. Use PROACTIVELY for generating, maintaining, and running E2E tests. Manages test journeys, quarantines flaky tests, uploads artifacts (screenshots, videos, traces), and ensures critical user flows work.

36 yesterday A 59 tokens

security-reviewer

03

NuGuardAI/nuguard

Agent

Security vulnerability detection and remediation specialist. Use PROACTIVELY after writing code that handles user input, authentication, API endpoints, or sensitive data. Flags secrets, SSRF, injection, unsafe crypto, and OWASP Top 10 vulnerabilities.

36 yesterday A 52 tokens

homemade-software-inc/completion-kit

Agent Claude Code

Assesses a pull request against CompletionKit's merge bar — is it worth merging at all, is it secure, is the code excellent, is it as simple as possible, and does it pass the project's hard gates (CI, 100% coverage, the inline test-schema gotcha, conventions). Use when triaging or reviewing an incoming PR, especially…

3 20d ago A 90 tokens