Instructions file CodexOpenCode
AGENTS.md instructions for mastra-ai/mastra: Unless asked, don't inspect reference or modify examples. Use the most-specific AGENTS.md; for package work, read packages/ /AGENTS.md first.
63 tagged evals, measured the same way as everything else here.
Browse within: agentic 17evaluations 15framework 15net 15Evaluation 6
Instructions file CodexOpenCode
AGENTS.md instructions for mastra-ai/mastra: Unless asked, don't inspect reference or modify examples. Use the most-specific AGENTS.md; for package work, read packages/ /AGENTS.md first.
Instructions file
Claude Code instructions for mastra-ai/mastra, a project described as: Mastra is the modern TypeScript framework for AI-powered applications and agents.
Instructions file CodexOpenCode
Instructions for harbor-framework/harbor, covering claude.md - harbor framework, contributing, project overview, quick start commands and install.
Instructions file
Instructions for harbor-framework/harbor, a project described as: Framework for evaluating and improving agents.
Instructions file
Instructions for ShenSeanChen/waku-agent, covering waku-agent — working conventions, architecture map (file ↔ diagram box), rules and commands.
Instructions file
Claude Code instructions for Jwuthri/Tracely-ai, covering claude.md, commands, architecture, hard rules and gotchas.
Instructions file CodexOpenCode
Instructions for benchflow-ai/benchflow, covering benchflow, setup + test, conventions and skill catalog (.agents/skills, mirrored at .claude/skills).
Instructions file
Instructions for benchflow-ai/benchflow, a project described as: Research infra for creating RL environments, post-training, and evals.
Instructions file CodexOpenCode
Repository instructions for working on NiceEval, including discovery, bug fixes, tests, whole-repository rules, and downstream projects. An AGENTS.md file gives coding agents project-specific rules.
Instructions file
Claude Code instructions for NiceEval/NiceEval, a project described as: build eval for your agent in 10 mins.
Instructions file GitHub Copilot
Instructions for AgentEvalHQ/AgentEval, covering agenteval - ai coding agent instructions, architecture overview, environment setup, optional: secondary models for comparison and build & test commands.
Instructions file GitHub Copilot
Instructions for detecting and adapting to Microsoft Agent Framework (MAF) breaking changes.
Instructions file GitHub Copilot
Guidelines for working with AgentEval Memory module — benchmarks, reporting, HTML reports, and LongMemEval integration.
Instructions file CodexOpenCode
Instructions for minghinmatthewlam/openbench, covering openbench — agent context, local execution context, what openbench is, execution ownership and product goals (the two things we are building toward).
Instructions file CodexOpenCode
Instructions for callstackincubator/evals, covering agents guide, what this repo is, mental model of execution, think-before-coding rules (required) and execution principles (default policy).
prime-radiant-inc/superpowers-evals
Instructions file CodexOpenCode
AGENTS.md instructions for prime-radiant-inc/superpowers-evals, a project described as: Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and…
prime-radiant-inc/superpowers-evals
Instructions file
Claude Code instructions for prime-radiant-inc/superpowers-evals, covering superpowers evals, canonical actors, commands, architecture and scenario conventions.
pensar-x/argus-validation-benchmarks
Instructions file
Instructions for pensar-x/argus-validation-benchmarks, covering project overview, what you're building, the goal, success criteria and what is apex?.
Instructions file CodexOpenCode
Instructions for netlify/axis, covering agents.md, project overview, terminology, architecture and adapter pattern.
Instructions file
Instructions for PrefectHQ/prefect-mcp-server, covering rules for contributors, investigating eval failures in ci, get check run id for the evaluation results, get annotations with failure details and fastmcp client.
Instructions file CodexOpenCode
Instructions for edonadei/caliper, covering caliper — agent instructions, updating docs after api changes, formatting and decision docs.
Instructions file
Instructions for edonadei/caliper, a project described as: Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.
Instructions file GitHub Copilot
Instructions for geval-labs/geval, covering geval - ai coding instructions, project overview, repository structure, build and test and cli and exit codes.
Instructions file CodexOpenCode
Instructions for luanmorenommaciel/brief-spec: Before any task, read OPERATING.md. Do not invent a second loop. Do not skip the HMAC seal. Do not put family hops in docs/.