AgentEval memory-benchmarks.instructions.md

A set of instructions for testing how well an AI agent remembers information across conversations. It includes built-in scenarios, the LongMemEval benchmark, saved baselines, and HTML reports.

In plain words
What is it for?
Use it to run memory benchmarks, compare results with baselines, test different history formats, and generate reports with charts and timelines.
Why use it?
It provides repeatable ways to measure memory instead of judging it from isolated examples.

Instructions file for GitHub Copilot

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add instructions/agentevalhq/agenteval/memory-benchmarks
Clone the repo
git clone --depth 1 https://github.com/AgentEvalHQ/AgentEval

Made for: GitHub Copilot.

Per session 4,074 This file is loaded in full into every session.
When invoked 4,074 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.04074 $0.04074
Opus 5 $0.02037 $0.02037
Sonnet 5 $0.00815 $0.00815
Haiku 4.5 $0.00407 $0.00407

Measured yesterday against content hash c109247199bf, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

AgentEval memory-benchmarks.instructions.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.github/instructions/memory-benchmarks.instructions.md · 362 lines

How it starts

The opening of the file, as written. The whole thing — 362 lines — stays where its author put it; the contents beside it link to each section on GitHub.

AgentEval Memory & Benchmarks — Agent Instructions

Module Overview

The AgentEval.Memory module provides comprehensive memory evaluation for AI agents:

  • Native benchmarks (12 scenario types across 5 presets)
  • External benchmarks (LongMemEval — ICLR 2025, 500 questions, 6 types)
  • Baseline persistence with JSON file storage
  • Interactive HTML reports with pentagon/radar charts, timeline, and comparison
  • History injection modes (Auto, Structured, TextBlob)

Architecture

src/AgentEval.Memory/
├── Abstractions/       Interfaces for memory operations
├── Assertions/         Memory-specific fluent assertion API
├── Data/               Embedded datasets, LongMemEval data files
├── DataLoading/        Scenario & corpus loaders, exporters
├── Engine/             Core: MemoryTestRunner, MemoryJudge
├── Evaluators/         MemoryBenchmarkRunner, ReachBack, Reducer, CrossSession evaluators
├── Extensions/         DI registration, helper methods
├── External/           External benchmarks (LongMemEval), interfaces
├── Metrics/            Memory-specific evaluation metrics
├── Models/             Core data: Benchmark, Result, Baseline, Config
├── Report/             HTML report template, pentagon mapper
├── Reporting/          Baseline store, comparer, output formatting
├── Scenarios/          Scenario providers: Memory, Chatty, Temporal, CrossSession
└── Temporal/           Temporal memory runner & scenarios

CRITICAL: Never Interrupt Running Benchmarks

NEVER kill, cancel, or interrupt a benchmark that is in progress. LLM benchmarks make hundreds of API calls and can take 10-100+ minutes to complete. Killing a benchmark wastes all progress, API costs, and time.

  • If you need to add console output, logging, or cosmetic changes — wait for the benchmark to finish first, then make changes and re-run.
  • If a benchmark appears to produce no output, it may still be running. Check the terminal for signs of life (CPU usage, network activity). The runner logs progress via Console.WriteLine after each question.
  • If you want to add progress reporting to a runner that lacks it, wait for the current run to complete.
  • The only valid reason to interrupt is an explicit user request or an unrecoverable crash.

Read the full file on GitHub · 362 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 362 lines · 4,074 tokens per session scan A c109247199bf

Subscribe to this mod's changes

AgentEval memory-benchmarks.instructions.md is an instructions file published in the GitHub repository AgentEvalHQ/AgentEval (138 stars, last pushed yesterday), licensed MIT. It adds 4,074 tokens to every session, about $0.0204 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.