ai evals skills

51 tagged ai evals, measured the same way as everything else here.

Browse within: ai-tool 41

reviewing-prs

01

future-agi/future-agi

Skill Claude CodeCodex

Use when asked to review a pull request, branch, or diff — before approving, as a self-review before opening a PR, to judge whether a PR is ready — 'is PR 123 mergeable?', 'anything blocking here?' — or to answer 'does this change need an E2E flow'. Applies the FutureAGI coding standards and, in repos with an e2e/…

1.9k +29 today A 129 tokens original Apache-2.0

writing-e2e-flows

02

future-agi/future-agi

Skill Claude CodeCodex

Use when a feature or fix in future-agi needs an end-to-end Playwright flow under e2e/ — a new flow for user-visible behaviour, an update to a flow whose pinned endpoint, route, table or selector changed, or when a review flagged missing E2E coverage. Also use to check whether an existing flow already pins an…

1.9k +29 today A 174 tokens original Apache-2.0

code-review

03

modiqo/skillspec

Skill Claude CodeCodex

Multi-agent code review with deep analysis. Orchestrates codebase research, optional web research, parallel Rust engineers, codex second opinion, and general-purpose reviewers into a synthesized report. Use when the user asks to review code, review a PR, review changes, audit code quality, or says "review", "/review"…

737 +2 24d ago A 116 tokens original Apache-2.0

durable-executor

04

modiqo/skillspec

Skill Claude CodeCodex

Universal first-hop for any tool-backed request that may execute, inspect, mutate, automate, browse, call, fetch, generate, install, test, run, or create artifacts through CLI, shell, API, adapter, browser, provider, local process, or external tool, with trace, alignment, evidence capture, and future recall. Use for…

737 +2 24d ago A 175 tokens original Apache-2.0

modiqo/skillspec

Skill Claude CodeCodex

Generic example for creating or updating a Codex skill with appropriate resources, metadata, validation, and forward-testing. Use for generic skill creator, create a skill, update a skill, skill creator, skill authoring, SKILL.md, agents/openai.yaml, initskill.py, quickvalidate.py, forward-test a skill, progressive…

737 +2 24d ago A 150 tokens original Apache-2.0

eval-driven-dev

06

yiouli/pixie-qa

Skill Claude CodeCodex

Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals…

7 4mo ago A 89 tokens original MIT

fde-consultant

07

selectess/fde-consultants-protocoles

Skill Claude CodeCodex

Forward Deployed Engineering co-pilot for coding agents, personal agents, software engineering, AI/agent systems, SaaS architecture, and business AI upgrades. Use when scoping, building, prototyping, or shipping AI products, SaaS features, or business transformations. Produces scoping reports, prototype specs…

5 2mo ago A 120 tokens original Apache-2.0

ui-2026-cinematic

08

selectess/fde-consultants-protocoles

Skill Claude CodeCodex

Cinematic 3D / scroll-driven / motion design skill for landing pages in 2026. Use when building marketing pages, SaaS sites, or product showcases that need scroll storytelling, 3D elements, micro-interactions, kinetic typography, or glassmorphism. Encodes modern best practices from Figma, Anthropic, Apple, Vercel.…

5 2mo ago A 99 tokens original Apache-2.0

calibration-guard

10

ByteStack-Labs/claude-plugins

Skill Claude CodeCodex

Detects and quantifies confidently-wrong behavior: where a model, classifier, agent, or LLM judge is highly confident and incorrect, and where the coupling between confidence and correctness breaks down under distribution shift. Measures calibration on the in-distribution set and again on the production distribution…

2 2mo ago A 208 tokens original MIT

production-autopsy

11

ByteStack-Labs/claude-plugins

Skill Claude CodeCodex

Start here. Audits a deployed ML or LLM or agent system that scores well on evaluation but fails, regresses, or behaves unexpectedly in production. Runs a reproducible root-cause "autopsy": frames the eval-to-deployment gap, reproduces the production failure, quantifies it by slice, tests confidence calibration under…

2 2mo ago A 258 tokens original MIT

tool-eval

12

ByteStack-Labs/claude-plugins

Skill Claude CodeCodex

Verifies the tools an agent or orchestrator depends on, by re-deriving a tool evaluation's real accuracy instead of trusting a single pass/fail score. Separates a formatting miss (a correct value scored wrong) from a real failure (a wrong value scored right), recomputes the expected answer from the raw inputs rather…

2 2mo ago A 253 tokens original MIT