Calibre-Labs/reforge-ai-evals

Market Map agent eval suite for the Reforge AI Evaluation course

68Stars on the repository
5Mods indexed here, across every type
4d agoLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied

eval-code

01

Calibre-Labs/reforge-ai-evals

Skill Claude CodeCodex

Write deterministic, code-based evaluator functions for an AI product. Use this skill whenever you need to write evaluators that check structural properties of AI outputs (format, count, presence, schema compliance) without calling another LLM. Also use when auditing existing code-based evaluators for brittleness…

not rated 68 +15 4d ago A 68 tokens

eval-llm-judge

02

Calibre-Labs/reforge-ai-evals

Skill Claude CodeCodex

Write LLM-as-judge evaluator prompts for AI products. Use this skill whenever you need to evaluate semantic properties of AI outputs that cannot be checked programmatically — output quality, tone, factual accuracy, edge case handling, reference alignment. Also use when auditing existing judge prompts for bias…

not rated 68 +15 4d ago A 92 tokens

llm-align

03

Calibre-Labs/reforge-ai-evals

Skill Claude CodeCodex

Analyze alignment between LLM judge scores and human labels in an eval dataset. Use this skill whenever someone wants to evaluate how well an LLM judge agrees with human reviewers, calculate TPR/TNR, investigate disagreements, or improve a scorer prompt. Trigger on phrases like: "calculate TPR/TNR", "judge alignment"…

not rated 68 +15 4d ago A 125 tokens

ticket-to-eval

04

Calibre-Labs/reforge-ai-evals

Skill Claude CodeCodex

Strip PII from a customer support ticket or eval trace and convert it into eval dataset rows. Produces two outputs: a regression dataset row (close to the original input, tagged with failuremode) and a generalized dataset row (abstracted for broader coverage). Works with local CSV files or any eval platform. Use this…

not rated 68 +15 4d ago A 127 tokens

uig

05

Calibre-Labs/reforge-ai-evals

Skill Claude CodeCodex

Build a User Input Grid (UIG) for an AI product or feature, evaluate existing eval datasets against it, and propose new inputs to fill coverage gaps. Use this skill whenever someone asks to design an eval framework, audit a test set, build a "user input grid", "synthetic query matrix" or "SQM", improve dataset…

not rated 68 +15 4d ago A 146 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: