Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add instructions/agentevalhq/agenteval/testinggit clone --depth 1 https://github.com/AgentEvalHQ/AgentEvalWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00484 | $0.00484 |
| Opus 5 | $0.00242 | $0.00242 |
| Sonnet 5 | $0.00097 | $0.00097 |
| Haiku 4.5 | $0.00048 | $0.00048 |
Grade A, and why
AgentEval testing.instructions.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
AgentEval Test Guidelines
Test Naming Convention
Use: MethodName_StateUnderTest_ExpectedBehavior
[Fact]
public async Task HaveCalledTool_WhenToolWasCalled_ShouldPass()
Test Structure
- Tests mirror
src/folder structure intests/AgentEval.Tests/ - Use xUnit with
[Fact]and[Theory]attributes - Each public class/method should have corresponding tests
Using FakeChatClient
For metrics that call LLMs, use FakeChatClient to avoid API calls:
var fakeClient = new FakeChatClient("""{"score": 95, "explanation": "Good"}""");
var metric = new FaithfulnessMetric(fakeClient);
var result = await metric.EvaluateAsync(context);
Creating ToolUsageReport for Tests
var report = new ToolUsageReport(new List<ToolCallRecord>
{
new() { Name = "SearchTool", CallId = "call-1", Result = "found" },
new() { Name = "ProcessTool", CallId = "call-2", Result = "done" }
});
Creating PerformanceMetrics for Tests
var metrics = new PerformanceMetrics
{
TotalDuration = TimeSpan.FromSeconds(2.5),
TimeToFirstToken = TimeSpan.FromMilliseconds(250),
PromptTokens = 100,
CompletionTokens = 50,
ModelUsed = "gpt-4o"
};
Assertion Tests Pattern
When testing assertions, expect ToolAssertionException or PerformanceAssertionException:
[Fact]
public void HaveCalledTool_WhenToolNotCalled_ShouldThrow()
{
var report = new ToolUsageReport([]);
var ex = Assert.Throws<ToolAssertionException>(() =>
report.Should().HaveCalledTool("MissingTool"));
Assert.Contains("MissingTool", ex.Message);
Assert.NotNull(ex.Expected);
Assert.NotNull(ex.Actual);
}
Multi-Target Framework Testing
Tests run on net8.0, net9.0, net10.0 - ensure compatibility. Use #if NET8_0 sparingly.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 68 lines · 484 tokens per session scan A 6e1700aabfc2
AgentEval testing.instructions.md is an instructions file published in the GitHub repository AgentEvalHQ/AgentEval (138 stars, last pushed yesterday), licensed MIT. It adds 484 tokens to every session, about $0.0024 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other instructions, from other repositories
zeroclaw CLAUDE.md
Instructions for zeroclaw-labs/zeroclaw, covering claude.md — zeroclaw (claude code), claude code settings, hooks and slash commands.
openagent CLAUDE.md
Claude Code instructions for the-open-agent/openagent, covering claude.md, commands, architecture, backend (go / beego) and frontend (react).
Tracely-ai CLAUDE.md
Claude Code instructions for Jwuthri/Tracely-ai, covering claude.md, commands, architecture, hard rules and gotchas.
nuwax AGENTS.md
Instructions for nuwax-ai/nuwax, covering ai agent system documentation, 系统概述, ai agent 架构, 核心组件 and ai 功能特性.
zhin zhin-plugin.instructions.md
Instructions for zhinjs/zhin, covering zhin plugin runtime authoring, package contract, convention directories, imports and native typescript and command routes.
OpenPersona AGENTS.md
Instructions for acnlabs/OpenPersona, covering agents.md, project overview, setup, project structure and architecture rules.