ai-evaluation skills

68 tagged ai-evaluation, measured the same way as everything else here.

Browse within: ai-governance 24ai-verification 23ai-workflows 23agent-benchmark 12agent-evaluation 12ai-coding-agents 11code-quality 9agentskills 8

ifixai

01

ifixai-ai/iFixAi

Skill Claude CodeCodex

Guide the user through an independent iFixAi audit of their own agent, checking whether it does the job it is supposed to do given their business rules and org structure. Prefer pointing it at the user's REAL deployed agent over its HTTP endpoint (its actual tools, retrieval, and governance) with --provider http…

12k +889 4d ago A 178 tokens original Apache-2.0

checkup

02

agentvitals/checkup

Skill Claude CodeCodex

Give your AI agent a professional health checkup (AgentVitals). Use when the user asks the agent to run a checkup / test itself / benchmark itself ("run a checkup", "check your vitals", "test yourself", "how stable are you", "/checkup"), or an advanced personality checkup (backbone, proactivity, creativity). 给 AI…

95 15d ago C 129 tokens AGPL-3.0

eval-engineer

03

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when a user is unsure which Eval Engineer command to run for AI agents/RAG apps, needs onboarding/status for a .galileo workspace, or asks where to start.

41 21d ago A 39 tokens original MIT

eval-fetch

04

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when a user asks to fetch Galileo evidence, provides Galileo URLs or IDs, says "fetch this Galileo link", or needs traces, sessions, experiments, or log streams saved locally.

41 21d ago A 40 tokens original MIT

eval-setup

05

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when the user asks to set up Eval Engineer, check .galileo readiness, create workspace scaffolding, or configure editable files, verification commands, app type, or evidence paths.

41 21d ago A 41 tokens original MIT

NoesisVision/nasde-toolkit

Skill Claude CodeCodex

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR…

12 9d ago A 171 tokens original MIT

NoesisVision/nasde-toolkit

Skill Claude CodeCodex

Run coding agent benchmarks and verify results with nasde. Use this skill when the user wants to: Run a benchmark (all tasks, single task, specific variant) Re-run assessment evaluation on existing trial results Check or verify results in Opik (traces, feedback scores, experiments) Troubleshoot a failed benchmark run…

12 9d ago A 139 tokens original MIT

tactical-ddd

08

NoesisVision/nasde-toolkit

Skill Claude CodeCodex

Design, refactor, analyze, and review code by applying the principles and patterns of tactical domain-driven design. Triggers on: domain modeling, aggregate design, 'entity', 'value object', 'repository', 'bounded context', 'domain event', 'domain service', code touching domain/ directories, rich domain model…

12 9d ago A 70 tokens original MIT

create-evaluation

09

Goodeye-Labs/truesight-mcp-skills

Skill Claude CodeCodex

Scope what quality should be measured, convert it into one or more actionable binary evaluations, deploy those evaluations through Truesight MCP, and generate a companion skill that applies them correctly. Use when a user wants to create new evals, quality checks, guardrails, or pass/fail criteria for AI outputs.

7 5mo ago A 66 tokens original MIT

error-analysis

10

Goodeye-Labs/truesight-mcp-skills

Skill Claude CodeCodex

Systematically identify and categorize failure modes in evaluated traces using Truesight datasets and error-analysis tools. Use when quality issues are unclear, after major pipeline changes, or when incidents indicate drift.

7 5mo ago A 41 tokens original MIT

Goodeye-Labs/truesight-mcp-skills

Skill Claude CodeCodex

Generate synthetic test data for LLM evaluations using dimension-based tuple expansion. Use when the user needs synthetic traces, test cases, eval datasets, or when create-evaluation needs synthetic fallback data.

7 5mo ago A 43 tokens original MIT

Ker102/Harneloop

Skill Claude CodeCodex

Operates Harneloop harness units through adaptive intake, environment mapping, artifact-aware attempts, explicit evaluation decisions, candidates, and evidence-gated promotion. Use when creating, continuing, evaluating, or recovering context for a Harneloop harness unit.

4 1mo ago A 62 tokens original Apache-2.0

rag-evaluation

13

karthikrshet/aiskills

Skill Claude CodeCodex

Use this skill to evaluate the quality of a RAG pipeline on faithfulness, answer relevancy, context precision, context recall, and hallucination rate. Activates after a RAG system is implemented or when retrieval quality is in question. Produces a structured evaluation report with measurable results.

3 12d ago A 63 tokens

deep-aggregate

14

Deep-Process/deep-process

Skill Claude CodeCodex

Use when user has multiple analysis outputs (risk, feasibility, architecture, verification) and needs a combined decision. Triggers: "combine these analyses", "decision brief", "GO/NO-GO", "aggregate results", "what's the overall verdict".

2 5mo ago A 56 tokens

deep-venture

15

Deep-Process/deep-process

Skill Claude CodeCodex

Use when user wants to find or create a business opportunity — whether discovering market gaps, inventing novel concepts, or both. Combines signal harvesting, gap detection, novel synthesis, and revenue architecture into one pipeline. Triggers: "find a niche", "what should I build", "business opportunity", "invent…

2 5mo ago A 82 tokens

deep-verify

16

Deep-Process/deep-process

Skill Claude CodeCodex

Use when user asks to verify, fact-check, validate, or check correctness of code, documents, specs, claims, or LLM-generated artifacts. Triggers: "verify this", "is this correct", "check against spec", "find contradictions", "review this document". Do NOT use for code review (style/quality) — this is for…

2 5mo ago A 80 tokens

accessibility-audit

17

vikast908/agent-repo-card

Skill Claude CodeCodex

Use when the user wants a WCAG 2.2 accessibility review of a UI — semantics, keyboard operability, focus management, color contrast, ARIA, forms/labels, reduced-motion, and screen-reader support, including streaming AI output via live regions. Triggers on "accessibility audit", "is this WCAG compliant", "a11y review"…

1 2mo ago A 90 tokens original MIT

token-efficiency

18

vikast908/agent-repo-card

Skill Claude CodeCodex

Use when the user wants to reduce LLM token usage, context-window pressure, or API cost in an AI/agent codebase without hurting quality — reviewing prompt construction, context assembly, chat-history retention, tool definitions, retrieval, caching, batching, and output verbosity. Triggers on "reduce token usage", "cut…

1 2mo ago A 90 tokens original MIT

ux-audit

19

vikast908/agent-repo-card

Skill Claude CodeCodex

Use when the user wants a UX / UI / interaction-design review or redesign of an app, dashboard, editor, canvas, AI/agentic product, web app, or mobile app — including microinteractions, motion, loading, error recovery, empty states, accessibility, perceived performance, and AI trust/observability. Triggers on "audit…

1 2mo ago A 99 tokens original MIT

adherencia-reglas

20

jleonceo/skill-adherencia-reglas

Skill Claude CodeCodex

Mide qué fracción de las reglas propias del proyecto se cumple de verdad, contándolo sobre el historial de sesiones que Claude Code ya guarda en disco. Se activa al preguntar si una norma, protocolo o convención se está aplicando, al revisar si un CLAUDE.md sirve de algo, al decidir si una regla debe pasar a ser hook…

1 1mo ago A 127 tokens original MIT

pocket-i-lab

21

yukakust/joinmultiplayer.ai

Skill Claude CodeCodex

Connect the current Codex task to a consented, redacted public experiment journal on joinmultiplayer.ai. Use when the user asks to start, continue, inspect, finish, or publicly document a Pocket i / joinmultiplayer experiment without leaving Codex. Do not use for ordinary private coding tasks or publish anything…

0 2d ago A 77 tokens original MIT

naturepedia

22

RobbieRazor/robbies-razor-benchmarks

Skill Claude CodeCodex

Use this skill whenever a user or autonomous agent needs scientifically grounded, machine-readable knowledge about natural systems, ecological relationships, geometry in nature, established mathematical references including Hopf Fibration, topology, fiber bundles, state-space geometry, weather, water systems, ocean…

0 2d ago A 100 tokens

clean-example

23

frankxai/starlight-evals

Skill Claude CodeCodex

A minimal, deliberately clean SKILL.md fixture used to CI-test scripts/lint-skills.mjs against a known-good file. It has valid frontmatter, no dollar-digit sequences, no secret-shaped strings, and no personal paths.

0 5d ago A 50 tokens