agent evaluation skills

196 tagged agent evaluation, measured the same way as everything else here.

Browse within: benchmark 65llm-agents 61ci 59python-cli 59release-gate 59coding-agents 26codex-cli 22agent-benchmark 18Evaluation 13ab-testing 13agent-governance 13agent-orchestration 13ai-agent-harness 13ai-product-management 13

ifixai

01

ifixai-ai/iFixAi

Skill Claude CodeCodex

Guide the user through an independent iFixAi audit of their own agent, checking whether it does the job it is supposed to do given their business rules and org structure. Prefer pointing it at the user's REAL deployed agent over its HTTP endpoint (its actual tools, retrieval, and governance) with --provider http…

12k 3d ago A 178 tokens original Apache-2.0

api-caller

02

NVIDIA/SkillEvaluator

Skill Claude CodeCodex

Call any REST API dynamically. Make GET, POST, PUT, DELETE requests to any endpoint with custom headers and JSON body.

357 2d ago A 29 tokens original Apache-2.0

calculator

03

NVIDIA/SkillEvaluator

Skill Claude CodeCodex

Evaluate mathematical expressions and unit conversions. Handles arithmetic, percentages, exponents, and common unit conversions (temperature, distance, weight). No external dependencies.

357 2d ago A 32 tokens original Apache-2.0

NVIDIA/SkillEvaluator

Skill Claude CodeCodex

Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

357 2d ago A 35 tokens original Apache-2.0

code-reviewer

05

hidai25/eval-view

Skill Claude CodeCodex

Performs comprehensive code reviews with security, quality, and best practice checks.

132 8d ago A 18 tokens original Apache-2.0

generate-tests

06

hidai25/eval-view

Skill Claude CodeCodex

Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

132 8d ago A 32 tokens original Apache-2.0

run-eval

07

hidai25/eval-view

Skill Claude CodeCodex

Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

132 8d ago A 30 tokens original Apache-2.0

dialogue-graph

08

Raidriar7170/hermes-skilleval

Skill Claude CodeCodex

A library for building, validating, visualizing, and serializing dialogue graphs. Use this when parsing scripts or creating branching narrative structures.

127 1mo ago A 32 tokens original MIT

docx

09

Raidriar7170/hermes-skilleval

Skill Claude CodeCodex

Word document manipulation with python-docx - handling split placeholders, headers/footers, nested tables.

127 1mo ago A 22 tokens original MIT

powerlifting

10

Raidriar7170/hermes-skilleval

Skill Claude CodeCodex

Calculating powerlifting scores to determine the performance of lifters across different weight classes.

127 1mo ago A 21 tokens original MIT

analyze

11

UiPath/coder_eval

Skill Claude CodeCodex

Analyze a finished coder-eval run and write analysis.md — cluster failures into systemic patterns and recommend fixes. Use when the user wants to know why a run failed, what regressed or got worse since a previous run, what to fix, or what a run says about their tasks.

119 3d ago A 57 tokens original Apache-2.0

check-skill

12

UiPath/coder_eval

Skill Claude CodeCodex

Generate and run a coder-eval activation suite for a Claude Code skill. Use when the user asks whether a skill triggers, wants to test skill activation, or worries a skill has silently stopped firing.

119 3d ago A 40 tokens original Apache-2.0

ci

13

UiPath/coder_eval

Skill Claude CodeCodex

Generate a GitHub Actions workflow that runs a coder-eval suite as a CI gate or on a schedule, using the published composite action — with the agent runtime, credentials, JUnit output and a score floor wired correctly.

119 3d ago A 45 tokens original Apache-2.0

samarailly51-pixel/claimpilot-harness

Skill Claude CodeCodex

Review and structure auto insurance bodily injury claims, including intake triage, evidence completeness, injury causation, treatment chronology, medical necessity, wage-loss support, negotiation risks, and human escalation. Use when analyzing bodily injury claim files, designing claims-agent workflows, drafting…

117 8d ago A 76 tokens original MIT

execution

15

Towow-ai/Flowness

Skill Claude CodeCodex

M-1.4 execution skill — 跑 single task 产 patch + 提交 envelope。.

102 24d ago A 22 tokens original Apache-2.0

fix-self-check

16

Towow-ai/Flowness

Skill Claude CodeCodex

An independent, read-only check for a code-fix report. It verifies that each claimed check was run against the real contract, rather than trusting a summary or allowing the fixer to alter its own evidence.

102 24d ago A 74 tokens original Apache-2.0

review

17

Towow-ai/Flowness

Skill Claude CodeCodex

M-1.5 review skill — 在 patch 跟 contract 之间找 finding,produce Finding 一等对象。.

102 24d ago A 26 tokens original Apache-2.0

frontend-design

18

thiientv/godmode

Skill Claude CodeCodex

Designs, builds, or refactors user-facing web or application interfaces with deliberate visual direction, typography, color and spacing tokens, content hierarchy, responsive behavior, accessible interaction states, and real rendered-surface verification. Use for pages, components, dashboards, landing pages, design…

93 6d ago A 83 tokens original MIT

pr-code-reviewer

19

thiientv/godmode

Skill Claude CodeCodex

Reviews a GitHub pull request or focused branch diff for correctness, regressions, security, compatibility, test gaps, and maintainability. Use when an implementation needs an independent source-aware review before merge. Produces prioritized findings with evidence and actionable fixes. Not for architecture-only…

93 6d ago A 70 tokens original MIT

using-godmode

20

thiientv/godmode

Skill Claude CodeCodex

Selects and composes the Godmode skills for a coding task, explains the activation boundary, keeps workflow skills from being skipped when they materially reduce risk, and identifies the task's current engineering lifecycle state. Use when starting a task with Godmode installed, deciding which skill to invoke…

93 6d ago A 91 tokens original MIT

code-review

21

ARTPARK-SAHAI-ORG/calibrate

Skill Claude CodeCodex

Project code-review style for this repo. Review the current branch's changes in two passes (correctness + justify-every-line), tracing any bug claim to ground before reporting.

20 2d ago A 38 tokens CC-BY-SA-4.0

multi-agent-systems-failure-taxonomy/ATLAS

Skill Claude CodeCodex

Use when working on an agent task where AdaMAST failure-mode checkpoints, final submission gates, trace capture, taxonomy generation/refinement, or AdaMAST CLI setup should guide Codex. This skill helps Codex apply AdaMAST during software or research tasks, diagnose its own trajectory against the active taxonomy, and…

19 1mo ago A 83 tokens original Apache-2.0

multi-agent-systems-failure-taxonomy/ATLAS

Skill Claude CodeCodex

Use when working on an agent task where AdaMAST failure-mode checkpoints, final submission gates, trace capture, taxonomy generation/refinement, or AdaMAST CLI setup should guide Codex. This skill helps Codex apply AdaMAST during software or research tasks, diagnose its own trajectory against the active taxonomy, and…

19 1mo ago A 83 tokens original Apache-2.0