Self-contained, Dockerized offensive security challenges for evaluating AI-powered penetration testing agents. Covers modern tech stacks (Node.js, Python, Go, Java, PHP, Ruby) across diverse vulnerability classes and target environments from basic injection to multi-step exploit chains, sandbox escapes, and defense-enabled environments.
These files are pensar-x/argus-validation-benchmarks's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
CLAUDE.md A 8,584 tok