Skill Claude CodeCodex
Performs comprehensive code reviews with security, quality, and best practice checks.
21 tagged agent benchmark, measured the same way as everything else here.
Browse within: agent-evaluation 18ai-evaluation 12ai-coding-agents 11CrewAI 7Evaluation 7langgraph 7
Skill Claude CodeCodex
Performs comprehensive code reviews with security, quality, and best practice checks.
Skill Claude CodeCodex
Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.
Skill Claude CodeCodex
Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.
Skill Claude CodeCodex
Give your AI agent a professional health checkup (AgentVitals). Use when the user asks the agent to run a checkup / test itself / benchmark itself ("run a checkup", "check your vitals", "test yourself", "how stable are you", "/checkup"), or an advanced personality checkup (backbone, proactivity, creativity). 给 AI…
Skill Claude CodeCodex
Run coding agent benchmarks and verify results with nasde. Use this skill when the user wants to: Run a benchmark (all tasks, single task, specific variant) Re-run assessment evaluation on existing trial results Check or verify results in Opik (traces, feedback scores, experiments) Troubleshoot a failed benchmark run…
Skill Claude CodeCodex
Design, refactor, analyze, and review code by applying the principles and patterns of tactical domain-driven design. Triggers on: domain modeling, aggregate design, 'entity', 'value object', 'repository', 'bounded context', 'domain event', 'domain service', code touching domain/ directories, rich domain model…
Skill Claude CodeCodex
Skill "refactor" from NoesisVision/nasde-toolkit, covering refactor, when to use, refactoring principles, the golden rules and when not to refactor.
Skill Claude CodeCodex
You are an experienced SRE + backend engineer reviewing a system-test bundle produced by silicon-system-test. The bundle captures everything from one run: server logs, per-agent logs, replay files, plus the orchestrator's own log and manifest.
Skill Claude CodeCodex
Validate scenario configuration files for correctness, consistency, balance, localization, and playability. Run with no arguments to check all scenarios, or pass a scenario name to check one.