benchmark
01Skill Claude CodeCodex
Stress-test an AI agent's behavioral alignment with 8 adversarial scenarios. Produces a grade (A-F) and score (0-100). Use when you want to know how well an agent handles apology traps, sycophancy tests, boundary pushes, error recovery, and more.