Evaluation skills

332 tagged Evaluation, measured the same way as everything else here.

Browse within: benchmark 61harness 55leaderboard 39benchmarking 35datasets 34gpt 33cyanheads 32evals 32grader 32inspect-ai 32lm-eval-harness 32browser-automation 28HumanEval 25ai-coding 24

benchmark-radar

25

ktwu01/benchmark-radar

Skill Claude CodeCodex

Find, inspect, and check AI benchmark records with the Benchmark Radar CLI. Use when a request needs benchmark discovery, details, recent Radar evidence, or local data health; do not assume why the user needs the results.

not rated 145 changed today A 48 tokens original MIT

ai-engineer

26

buiphucminhtam/forgewright

Skill Claude CodeCodex

Builds production AI/ML systems — model training, fine-tuning, MLOps pipelines, model serving, evaluation frameworks, RAG optimization, and agent orchestration at scale. Use when the user asks to build, train, or deploy ML models, set up MLOps pipelines, optimize RAG systems, create inference endpoints, or design…

not rated 49 yesterday A 78 tokens

benchjack

27

benchjack/benchjack

Skill Claude CodeCodex

Audit an AI agent benchmark for hackability. Detects evaluation vulnerabilities like missing isolation, leaked answers, eval() on untrusted input, prompt injection in LLM judges, weak scoring, logic gaps, and trust of untrusted output. Use when analyzing whether a benchmark can be gamed or exploited.

not rated 45 +1 3mo ago A 63 tokens original Apache-2.0

lai-gen-evals

28

eclipse-langium/langium-ai

Skill Claude CodeCodex

Expand and refine the evaluation suite for a Langium DSL project. Generates comprehensive eval files that cover syntactic correctness, semantic validity, user intent matching, edge cases, and language understanding.

not rated 30 7d ago A 42 tokens original MIT

lai

29

eclipse-langium/langium-ai

Skill Claude CodeCodex

Guide for using the langium-ai (LAI) CLI to generate language descriptors, synthesize system prompts, run evaluations, and iteratively refine AI-powered tooling in Langium projects. Use when working with lai commands, descriptors, or evaluation files.

not rated 30 7d ago A 52 tokens original MIT

langium

30

eclipse-langium/langium-ai

Skill Claude CodeCodex

A comprehensive skill to understanding how Langium-based projects work — from grammar definition through code generation, runtime parsing, linking, validation, and LSP integration.

not rated 30 7d ago A 33 tokens original MIT

protein-design-qc

31

zongtingwei/Bioclaw_Skills_Hub

Skill Claude CodeCodex

Protein design quality control, filtering thresholds, and ranking guidance. Use this skill when: (1) Evaluating design quality for binding, expression, or structure, (2) Setting filtering thresholds for pLDDT, ipTM, PAE, (3) Checking sequence liabilities (cysteines, deamidation, polybasic clusters), (4) Creating…

not rated 26 4mo ago A 143 tokens original MIT

run-aeon-benchmark

32

AEON-7/Aeon-Bench-Pod

Skill Claude CodeCodex

Use when asked to run, benchmark, evaluate, or score an LLM with AEON Bench. You run the AEON Bench Pod on the user's machine, point it at a model, run the benchmark, and submit the signed result to the public leaderboard at aeon-bench.com. All work happens on the pod. The mothership only shows the board and accepts…

not rated 25 +2 18d ago A 83 tokens original MIT

Accelerated-Innovation/governed-ai-delivery

Skill Claude CodeCodex

Plan implementation work as a sequence of the smallest independently demonstrable increments. Use when given a feature request, a body of work, or any implementation task that must be broken down or sequenced BEFORE coding — it classifies the request, identifies the next demonstrable behaviors, and recommends…

not rated 20 yesterday A 92 tokens

superior-unit-tests

34

Accelerated-Innovation/governed-ai-delivery

Skill Claude CodeCodex

Write or improve unit tests following a strict determinism, F.I.R.S.T., and fast-feedback-budget discipline. Use when asked to write tests, improve tests, review test quality, classify tests, or assess a test suite against a 30-second fast-feedback budget. Invokes a structured output protocol that identifies behaviors…

not rated 20 yesterday A 88 tokens

zom-representation

35

Accelerated-Innovation/governed-ai-delivery

Skill Claude CodeCodex

Zero-One-Many Representation Evolution — use this skill whenever working with existing code that needs to change, grow, or be refactored. Triggers include: "refactor this", "add a new X to this", "this is getting messy", "how should I represent this", "there's a lot of duplication here", or any time numbered…

not rated 20 yesterday A 137 tokens

test-skill

36

joeynyc/skillscore

Skill Claude CodeCodex

Fetches current weather data for any city using OpenWeatherMap API.

not rated 17 4mo ago A 0 tokens original MIT

File-Backup-Tool

37

joeynyc/skillscore

Skill Claude CodeCodex

Creates backups of important files to a remote location.

not rated 17 4mo ago A 16 tokens original MIT

api-integration

38

joeynyc/skillscore

Skill Claude CodeCodex

Integrates with various REST APIs using different URL patterns and endpoints when data fetching or authentication is needed.

not rated 17 4mo ago A 24 tokens original MIT

eval-surfer

39

di37/EvalSurfer

Skill Claude CodeCodex

Drive AI application evaluations using the EvalSurfer skill-first workflow. Use when creating AI eval rubrics, reviewing RAG outputs, checking agent tool use, assessing safety, or calculating operational metrics like latency, TTFT, inter-token latency, throughput (tokens per second), P99 tail latency, cost, cost per…

not rated 11 1mo ago A 81 tokens original MIT

skill-evaluator

40

lanyasheng/skill-evaluator

Skill Claude CodeCodex

A tool for evaluating and improving other coding-agent skills through benchmarks, adversarial tests, and repeated improvement cycles.

not rated 11 5mo ago A 31 tokens

michael-polanyi

41

August1314/Michael-Polanyi

Skill Claude CodeCodex

A writing guide for giving practical advice when the situation is unclear, trade-offs matter, or information is incomplete.

not rated 9 4mo ago A 87 tokens original MIT

ad-copy-google-ads

42

edholofy/dojo.md

Skill Claude CodeCodex

Writes Google Ads copy including RSA headlines, descriptions, extensions, DKI, and CTAs optimized for Quality Score and search intent. Use when creating, auditing, or rewriting Google Search ad campaigns from keyword research through final copy with extensions.

not rated 9 4mo ago A 53 tokens original MIT

ad-copy-google-ads

43

edholofy/dojo.md

Skill Claude CodeCodex

Writes Google Ads copy including RSA headlines, descriptions, extensions, DKI, CTAs, and complete campaign builds. Use when creating or optimizing Google Search ads, diagnosing Quality Score issues, or aligning ad copy with search intent.

not rated 9 4mo ago A 51 tokens original MIT

ad-copy-google-ads

44

edholofy/dojo.md

Skill Claude CodeCodex

Writes Google Ads copy including RSA headlines, descriptions, extensions, DKI, CTAs, and intent-matched campaigns. Use when creating, reviewing, or fixing Google Search ad copy, building complete ad groups, or optimizing Quality Score through copywriting.

not rated 9 4mo ago A 56 tokens original MIT

pitlane-ai/pitlane

Skill Claude CodeCodex

Design and create pitlane eval benchmarks that measure whether an AI coding skill or MCP server actually improves assistant performance. Use when the user wants to test a skill, evaluate an MCP server, create a pitlane eval YAML, benchmark an AI assistant, or compare baseline vs challenger configurations. Covers eval…

not rated 7 2mo ago A 76 tokens

skill-tester

47

topprismdata/skill-tester

Skill Claude CodeCodex

Tests and evaluates any Claude Code skill for structural validity, quality, and trigger accuracy. Implements the cc-plugin-eval 4-stage pipeline (Analysis → Generation → Execution → Evaluation) and the 4D scoring rubric (Documentation/Code/Completeness/Usability 25% each). Use before packaging or deploying any skill.

not rated 4 6d ago A 70 tokens

mcplab-assistant

48

inspectr-hq/mcplab

Skill Claude CodeCodex

Operator guide for MCPLab config authoring and execution workflows. Use when users need help writing or debugging MCPLab eval YAML, including response assertions, MCP tool constraints, tool-input assertions, and Judge/agent checks with optional prompt, tool-sequence, and tool-input context; running scenarios (prefer…

not rated 4 4d ago A 118 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: