Write deterministic, code-based evaluator functions for an AI product. Use this skill whenever you need to write evaluators that check structural properties of AI outputs (format, count, presence, schema compliance) without calling another LLM. Also use when auditing existing code-based evaluators for brittleness…
Write LLM-as-judge evaluator prompts for AI products. Use this skill whenever you need to evaluate semantic properties of AI outputs that cannot be checked programmatically — output quality, tone, factual accuracy, edge case handling, reference alignment. Also use when auditing existing judge prompts for bias…
Analyze alignment between LLM judge scores and human labels in an eval dataset. Use this skill whenever someone wants to evaluate how well an LLM judge agrees with human reviewers, calculate TPR/TNR, investigate disagreements, or improve a scorer prompt. Trigger on phrases like: "calculate TPR/TNR", "judge alignment"…
Strip PII from a customer support ticket or eval trace and convert it into eval dataset rows. Produces two outputs: a regression dataset row (close to the original input, tagged with failuremode) and a generalized dataset row (abstracted for broader coverage). Works with local CSV files or any eval platform. Use this…
Build a User Input Grid (UIG) for an AI product or feature, evaluate existing eval datasets against it, and propose new inputs to fill coverage gaps. Use this skill whenever someone asks to design an eval framework, audit a test set, build a "user input grid", "synthetic query matrix" or "SQM", improve dataset…
★not rated 68▲
+15 4d agoA146 tokens
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: