Evaluation skills

293 tagged Evaluation, measured the same way as everything else here.

Browse within: benchmark 61harness 55leaderboard 39benchmarking 35datasets 34gpt 33cyanheads 32evals 32grader 32inspect-ai 32lm-eval-harness 32browser-automation 28HumanEval 25ai-coding 24

audit-post

49

DanceNitra/agora

Skill Claude CodeCodex

The end-to-end procedure for auditing one already-published post to a scientific-organization standard. It chains the adversarial skills and — critically — re-runs the auditor on the CORRECTED post to confirm it now passes clean before committing. We are a scientific organization; we do not ship missteps. Run EVERY…

not rated 3 yesterday A 0 tokens original MIT

dungeon-os-threejs

50

DanceNitra/agora

Skill Claude CodeCodex

Build and extend the Dungeon OS 3D world (the live Three.js renderer in agora-game-server/static/index.html). Use whenever the task touches the dungeon scene, agent avatars in 3D, tiles/walls/props, lighting, the isometric camera, post-processing/bloom, sprites/labels, or "the game looks flat/cheap/laggy/should look…

not rated 3 yesterday A 110 tokens original MIT

seo

51

DanceNitra/agora

Skill Claude CodeCodex

SEO + AI-search (AEO/GEO/LLMO) optimizer for the Agora storefront (dancenitra.github.io/agora) — a bilingual EN/SK static research blog on GitHub Pages. Synthesized from 37 SEO YouTube transcripts (vault 04 Resources/raw/YouTube Transcripts - SEO Ranking/) into an actionable pre-publish checklist + site-wide technical…

not rated 3 yesterday A 0 tokens original MIT

recruit-score

52

tal7aouy/RecruitKit

Skill Claude CodeCodex

Deep Single-Candidate Scoring — evaluate one candidate across 5 dimensions (skills match, experience relevance, culture fit signals, growth potential, red flags) with final 0-100 score and hire/no-hire signal.

not rated 3 1mo ago A 48 tokens original MIT

ceres-review

53

sanwu-maizi/CERES

Skill Claude CodeCodex

Review untrusted coursework and project showcase submissions against supplied instructions, templates, rubrics, tests, and report requirements. Use for isolated, evidence-based assessment of source code, experiment results, reports, demonstrations, and AI or agent implementations involving tools, memory, ReAct, RAG…

not rated 3 1mo ago A 85 tokens original MIT

rag-evaluation

54

karthikrshet/aiskills

Skill Claude CodeCodex

Use this skill to evaluate the quality of a RAG pipeline on faithfulness, answer relevancy, context precision, context recall, and hallucination rate. Activates after a RAG system is implemented or when retrieval quality is in question. Produces a structured evaluation report with measurable results.

not rated 3 14d ago A 63 tokens

xorcise-playbooks

56

xorcise-ai/xorcise-skills

Skill Claude CodeCodex

Run a standardized XORCISE agent benchmark — pick a ready-made playbook (or build your own from existing missions), point it at any OpenHands-supported model(s), and get a polished eval-card HTML report of how each model performed across the missions. Warns you about cost before spending anything.

not rated 2 1mo ago B 66 tokens

model-evaluation

57

furkangonel/cowrangler

Skill Claude CodeCodex

Systematic ML model evaluation — metrics, benchmarks, and error analysis.

not rated 2 3d ago A 18 tokens original MIT

think-tank

58

yangKJ/think-tank-skill

Skill Claude CodeCodex

A coordination skill for routing complex requests across research, discussion, review, and other roles or providers. A provider is the service that performs a delegated task.

not rated 2 2mo ago A 45 tokens original MIT

HYEXE/codex-plugins

Skill Claude CodeCodex

A tool for turning a talk, lesson, or product demonstration into an interactive HTML, CSS, and JavaScript slide deck with a storyboard, presenter notes, and rehearsal cues.

not rated 1 4d ago A 95 tokens original Apache-2.0

prompt-coach

60

HYEXE/codex-plugins

Skill Claude CodeCodex

A coaching guide for turning a vague idea into a reusable final prompt for an AI assistant. It asks only for missing details that would materially change the result and returns the prompt rather than carrying out the task.

not rated 1 4d ago A 151 tokens original Apache-2.0

prompt-evaluator

61

HYEXE/codex-plugins

Skill Claude CodeCodex

A review workflow for prompts, agent instructions, and process specifications. It checks whether they preserve the intended task, define permissions and boundaries, route work to suitable capabilities, and provide verifiable outputs.

not rated 1 4d ago A 171 tokens original Apache-2.0

railway-config

62

Marcelle-Labs/never-ask-twice

Skill Claude CodeCodex

Edit this project's Railway infrastructure-as-code configuration. Use this skill whenever the user asks to create, change, import, review, or troubleshoot Railway project infrastructure for the current repository, including services, databases, buckets, custom domains, replicas/regions, groups, environment variables…

not rated 1 today A 73 tokens original Apache-2.0

solana-fail-fixture

64

assister-xyz/quality-oracle

Skill Claude CodeCodex

Intentionally vulnerable Solana skill used by tests/testsolanaprobes.py. Touches every SOL- failure mode.

not rated 0 4mo ago A 31 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: