A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook.
eval-guide is an AI agent evaluation toolkit for Copilot Studio that helps users plan evaluations, create test cases, interpret results, and diagnose failures. It is used with Claude Code or GitHub Copilot to assess single-response and multi-turn agent behavior using Microsoft’s evaluation guidance.
These files are microsoft/eval-guide's own configuration. They tell GitHub Copilot, Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
.github/copilot-instructions.md A 422 tok AGENTS.md A 1,507 tok CLAUDE.md A 1,575 tok