microsoft/eval-guide

A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook.

127Stars on the repository
11Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

eval-faq

01

microsoft/eval-guide

Skill Claude CodeCodex

Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage & Improvement Playbook, Eval Guidance Kit) supplemented by select industry sources.

127 2mo ago A 48 tokens original MIT

eval-generator

02

microsoft/eval-guide

Skill Claude CodeCodex

Generate standalone — turns the populated Eval Suite Planning workbook (output of /eval-suite-planner) into concrete capability eval sets and trust & safety eval sets. Delivers playbook Steps 2 & 3 and designs the Step 8 regression partition. Outputs 2-column Copilot Studio -for-import.csv files (Question + Expected…

127 2mo ago A 115 tokens original MIT

eval-guide

03

microsoft/eval-guide

Skill Claude CodeCodex

Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately. No built agent required — an idea or description is enough. Promotes eval-first development: write evals before building. Use when…

127 2mo ago B 98 tokens original MIT

microsoft/eval-guide

Skill Claude CodeCodex

Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics. Returns a gate-based SHIP / ITERATE / BLOCK verdict with root cause classification, remediation, and pattern analysis.

127 2mo ago A 65 tokens original MIT

eval-suite-planner

05

microsoft/eval-guide

Skill Claude CodeCodex

Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. Grounded in Practical Guidance on Agent Evaluation v5: Step 1 planning, Steps 2-3 eval-set decomposition, Step 4 gates/improvement targets, Step 5 human inputs, Step 6 grader-validation…

127 2mo ago A 122 tokens original MIT

microsoft/eval-guide

Skill Claude CodeCodex

Use this skill when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps, or analyze patterns to improve their agent. Always use this skill when the user mentions: "eval failed", "why did this fail"…

127 2mo ago A 121 tokens original MIT