agent benchmark skills

21 tagged agent benchmark, measured the same way as everything else here.

Browse within: agent-evaluation 18ai-evaluation 12ai-coding-agents 11CrewAI 7Evaluation 7langgraph 7

code-reviewer

01

hidai25/eval-view

Skill Claude CodeCodex

Performs comprehensive code reviews with security, quality, and best practice checks.

132 9d ago A 18 tokens original Apache-2.0

generate-tests

02

hidai25/eval-view

Skill Claude CodeCodex

Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

132 9d ago A 32 tokens original Apache-2.0

run-eval

03

hidai25/eval-view

Skill Claude CodeCodex

Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

132 9d ago A 30 tokens original Apache-2.0

checkup

04

agentvitals/checkup

Skill Claude CodeCodex

Give your AI agent a professional health checkup (AgentVitals). Use when the user asks the agent to run a checkup / test itself / benchmark itself ("run a checkup", "check your vitals", "test yourself", "how stable are you", "/checkup"), or an advanced personality checkup (backbone, proactivity, creativity). 给 AI…

95 15d ago C 129 tokens AGPL-3.0

NoesisVision/nasde-toolkit

Skill Claude CodeCodex

Run coding agent benchmarks and verify results with nasde. Use this skill when the user wants to: Run a benchmark (all tasks, single task, specific variant) Re-run assessment evaluation on existing trial results Check or verify results in Opik (traces, feedback scores, experiments) Troubleshoot a failed benchmark run…

12 8d ago A 139 tokens original MIT

tactical-ddd

06

NoesisVision/nasde-toolkit

Skill Claude CodeCodex

Design, refactor, analyze, and review code by applying the principles and patterns of tactical domain-driven design. Triggers on: domain modeling, aggregate design, 'entity', 'value object', 'repository', 'bounded context', 'domain event', 'domain service', code touching domain/ directories, rich domain model…

12 8d ago A 70 tokens original MIT

refactor

07

NoesisVision/nasde-toolkit

Skill Claude CodeCodex

Skill "refactor" from NoesisVision/nasde-toolkit, covering refactor, when to use, refactoring principles, the golden rules and when not to refactor.

12 8d ago A 0 tokens original MIT

review-system-test

08

haoyifan/Silicon-Pantheon

Skill Claude CodeCodex

You are an experienced SRE + backend engineer reviewing a system-test bundle produced by silicon-system-test. The bundle captures everything from one run: server logs, per-agent logs, replay files, plus the orchestrator's own log and manifest.

6 3mo ago A 0 tokens original Apache-2.0

scenario-check

09

haoyifan/Silicon-Pantheon

Skill Claude CodeCodex

Validate scenario configuration files for correctness, consistency, balance, localization, and playability. Run with no arguments to check all scenarios, or pass a scenario name to check one.

6 3mo ago A 38 tokens original Apache-2.0