microsoft/eval-guide

A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook.

About the project

eval-guide is an AI agent evaluation toolkit for Copilot Studio that helps users plan evaluations, create test cases, interpret results, and diagnose failures. It is used with Claude Code or GitHub Copilot to assess single-response and multi-turn agent behavior using Microsoft’s evaluation guidance.

These files are microsoft/eval-guide's own configuration. They tell GitHub Copilot, Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

128Stars on the repository
3Files it configures its agents with
3,504Tokens loaded in every session
4Agents configured

Instructions