eval-guide AGENTS.md

Instructions for an AI-agent evaluation toolkit based on Microsoft's guidance for testing agents. An evaluation checks whether an agent gives suitable answers across planned test cases.

In plain words
What is it for?
Use it to plan evaluations, create test cases, choose grading methods, and review agent performance.
Why use it?
It helps organise testing when you need to decide what to measure, how to test it, and how to interpret the results.

Instructions file for CodexOpenCode

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add instructions/microsoft/eval-guide/agents-md
Clone the repo
git clone --depth 1 https://github.com/microsoft/eval-guide

Made for: Codex, OpenCode.

Per session 1,507 This file is loaded in full into every session.
When invoked 1,507 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.01507 $0.01507
Opus 5 $0.00754 $0.00754
Sonnet 5 $0.00301 $0.00301
Haiku 4.5 $0.00151 $0.00151

Measured yesterday against content hash 9904e2d1b347, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

eval-guide AGENTS.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

AGENTS.md · 79 lines

How it starts

The opening of the file, as written. The whole thing — 79 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Eval Guide — AI Agent Evaluation Toolkit

This repository contains an AI agent evaluation toolkit for Copilot Studio, grounded in Microsoft's Practical Guidance on Agent Evaluation — a 10-step playbook. The canonical methodology spine is skills/eval-guide/playbook.md.

What This Toolkit Does

Helps users go from "I don't know where to start with eval" to "I have a plan, test cases, and know how to interpret results" — in one session. No running agent required for planning and test generation.

Available Prompt Files

This toolkit provides 6 prompt files in .github/prompts/. When the user's request matches one of these, attach or reference the appropriate prompt file:

Prompt File When to Use
eval-guide.prompt.md Full eval lifecycle — discover, plan, generate, run, interpret. Start here when the user mentions agent evaluation, eval planning, "what should we test", or "how do we know if the agent is good".
eval-suite-planner.prompt.md Populated Eval Suite Template workbook plus an interactive HTML review page with eval sets, methods, gates, human inputs, and grader-validation notes. Use when the user has an agent description and needs a plan before generating test cases.
eval-generator.prompt.md Generate test cases (CSV for single-response, blueprints for multi-turn). Use after planning, or standalone with an agent description.
eval-result-interpreter.prompt.md SHIP / ITERATE / BLOCK verdict from eval results. Use when the user has CSV results or pass/fail data to interpret.
eval-triage-and-improvement.prompt.md Interactive diagnosis and remediation for failing evals. Use when the user needs help debugging specific failures.
eval-faq.prompt.md Methodology questions answered from Microsoft's eval ecosystem. Use for "how do I...", "what is...", "when should I..." eval questions.

Routing Guide

User says... Use this prompt
"We're planning to build an agent for..." eval-guide
"Help us think through what good looks like" eval-guide
"Here's our agent, plan the eval" eval-suite-planner
"I have a plan, generate test cases" eval-generator
"My evals came back, what do they mean?" eval-result-interpreter
"Some tests are failing and I don't know why" eval-triage-and-improvement
"How is evaluating X different from Y?" eval-faq

Read the full file on GitHub · 79 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 79 lines · 1,507 tokens per session scan A 9904e2d1b347

Subscribe to this mod's changes

eval-guide AGENTS.md is an instructions file published in the GitHub repository microsoft/eval-guide (127 stars, last pushed 2mo ago), licensed MIT. It adds 1,507 tokens to every session, about $0.0075 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other instructions, from other repositories

codex AGENTS.md

AGENTS.md instructions for openai/codex, covering rust/codex-rs, the codex-core crate, code review rules, crate api surface and model visible context.

openai/codex · 5,182 tokens

buildNext

Working notes and architecture documentation for the new esbuild-based build system in build/next. Use when making changes to the new build pipeline (transpile/bundle commands, NLS plugin, source-map handling, resource copying, or self-hosting watch tasks).

microsoft/vscode · 6,785 tokens

next.js AGENTS.md

Instructions for vercel/next.js, covering next.js development guide, codebase structure, monorepo overview, core package: packages/next and other important packages.

vercel/next.js · 7,296 tokens

vscode oss-third-party-notices.instructions.md

Instructions for microsoft/vscode, covering vs code oss third-party-notices pipeline, architecture, pipeline flow in ci, applying the notice (cutover) and fallback chain (never fail the build).

microsoft/vscode · 5,001 tokens

spec-kit AGENTS.md

Instructions for github/spec-kit, covering agents.md, about spec kit and specify, quickstart — add a new integration in 5 steps, integration architecture and integrationmanifest — file tracking.

github/spec-kit · 7,040 tokens

langchain AGENTS.md

Instructions for langchain-ai/langchain, covering global development guidelines for the langchain monorepo, corridor security analysis, project architecture and context, monorepo structure and development tools & commands.

langchain-ai/langchain · 4,345 tokens