evaluation framework skills

53 tagged evaluation framework, measured the same way as everything else here.

Browse within: llms 15vscode 15agent-evaluation 6coding-agents 6

deepeval-otel

01

confident-ai/deepeval

Skill Claude CodeCodex

Export raw OpenTelemetry traces from an AI application to Confident AI's Observatory. TRIGGER when the user wants to send OpenTelemetry or OTLP traces/spans from an LLM app, agent, RAG pipeline, or chatbot to Confident AI; configure the Confident AI OTLP endpoint; set confident.span. or confident.trace. attributes…

18k 2d ago A 226 tokens original Apache-2.0

deepeval-tracing

02

confident-ai/deepeval

Skill Claude CodeCodex

Instrument an AI application with DeepEval's native tracing so its behavior is visible in Confident AI. TRIGGER when the user wants to add DeepEval tracing or @observe to an LLM app, agent, RAG pipeline, or chatbot; wire a framework, model-provider, or vector-database integration (LangGraph, LangChain, OpenAI Agents…

18k 2d ago A 208 tokens original Apache-2.0

deepeval

03

confident-ai/deepeval

Skill Claude CodeCodex

DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app; add evals; generate datasets or goldens; use deepeval generate; use deepeval test run; send results to Confident AI…

18k 2d ago A 229 tokens original Apache-2.0

analyze

04

UiPath/coder_eval

Skill Claude CodeCodex

Analyze a finished coder-eval run and write analysis.md — cluster failures into systemic patterns and recommend fixes. Use when the user wants to know why a run failed, what regressed or got worse since a previous run, what to fix, or what a run says about their tasks.

119 3d ago A 57 tokens original Apache-2.0

check-skill

05

UiPath/coder_eval

Skill Claude CodeCodex

Generate and run a coder-eval activation suite for a Claude Code skill. Use when the user asks whether a skill triggers, wants to test skill activation, or worries a skill has silently stopped firing.

119 3d ago A 40 tokens original Apache-2.0

ci

06

UiPath/coder_eval

Skill Claude CodeCodex

Generate a GitHub Actions workflow that runs a coder-eval suite as a CI gate or on a schedule, using the published composite action — with the agent runtime, credentials, JUnit output and a score floor wired correctly.

119 3d ago A 45 tokens original Apache-2.0

creating-reports

07

8ddieHu0314/Skill-Lab

Skill Claude CodeCodex

Creates detailed reports from data when the user asks for report generation or data summaries.

55 4mo ago A 20 tokens original Apache-2.0

clean

08

8ddieHu0314/Skill-Lab

Skill Claude CodeCodex

Use when you need to convert a CSV file to JSON format.

55 4mo ago A 15 tokens original Apache-2.0

testing-features

09

8ddieHu0314/Skill-Lab

Skill Claude CodeCodex

Tests various features in applications to ensure quality. Use when testing is needed.

55 4mo ago A 19 tokens original Apache-2.0

evalbench-review

10

GoogleCloudPlatform/evalbench

Skill Claude CodeCodex

Review a change in the EvalBench repo for (a) does it actually work — verified by running the tests and style checks, (b) does it follow EvalBench architecture — base-class contracts, config-key registration, PYTHONPATH-relative imports, sandbox isolation, concurrency safety, docs, (c) does it still build and deploy …

55 3d ago A 170 tokens original Apache-2.0

agent-plugin-review

11

EntityProcess/agentv

Skill Claude CodeCodex

Use when reviewing an AI plugin pull request, auditing plugin quality before release, or when asked to "review a plugin PR", "review skills in this PR", "check plugin quality", or "review workflow architecture". Covers skill quality, structural linting, and workflow architecture review.

15 1mo ago A 60 tokens original MIT

agentv-bench

12

EntityProcess/agentv

Skill Claude CodeCodex

Run AgentV evaluations and optimize agents through eval-driven iteration. Triggers: run evals, benchmark agents, optimize prompts/skills against evals, compare agent outputs across providers, analyze eval results, offline evaluation of recorded sessions, run autoresearch, optimize unattended, run overnight…

15 1mo ago A 101 tokens original MIT

agentv-eval-writer

13

EntityProcess/agentv

Skill Claude CodeCodex

Write, edit, review, and validate AgentV EVAL.yaml / .eval.yaml evaluation files. Use when asked to create new eval files, update or fix existing ones, add or remove test cases, configure graders (llm-rubric, script), review whether an eval is correct or complete, convert between EVAL.yaml and evals.json using agentv…

15 1mo ago A 129 tokens original MIT

auto-itera

14

clfhaha1234/auto-itera

Skill Claude CodeCodex

Use when the user wants to autonomously search for the best AI/engineering approach across competing candidates (prompts, models, retrieval strategies, architectures, algorithms) — give it a goal + candidate arms + success threshold, it runs the experiment to a defensible ship-or-kill verdict. Autonomously handles…

6 3mo ago A 114 tokens original MIT

gauntlet-loop

15

israeldegasperi/gauntlet-loop

Skill Claude CodeCodex

Run Matt Shumer's Gauntlet Loop methodology to raise the quality of an artifact through cycles of building, independent critique, comparison against real references, and improvement. Use when the user asks for the Gauntlet Loop, a builder-critic process, adversarial critique, iterative refinement, or the creation of…

1 5d ago A 95 tokens