replayd: Instructions file for Codex

AGENTS.md

replayd AGENTS.md is an instructions file for Codex, OpenCode from TaimoorKhan10/replayd. It costs 4,501 tokens per session, scanned A, original, MIT.

A set of instructions for replayd, a tool that records failed AI-agent runs and replays them as regression tests.

In plain words
What is it for?
Use it to capture inputs, outputs, and tool calls, mark a run as failed, save it as a test, and replay tests against later agent versions.
Why use it?
It helps catch an agent repeating an old mistake after a prompt, model, or code change. Failed behavior can block a release when it returns.

Instructions file for CodexOpenCode

Written for Codex and OpenCode: the file is AGENTS.md.

This is TaimoorKhan10/replayd's own configuration. It tells Codex and OpenCode how to work on replayd itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything replayd configures →

Reuse

Borrowing it

Nothing to install: this file belongs to TaimoorKhan10/replayd. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/TaimoorKhan10/replayd/main/AGENTS.md
Clone the repo
git clone --depth 1 https://github.com/TaimoorKhan10/replayd

Made for: Codex, OpenCode.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for replayd AGENTS.md

README.md
[![agentmods](https://agentmods.dev/badge/instructions/taimoorkhan10/replayd/agents-md.svg)](https://agentmods.dev/instructions/taimoorkhan10/replayd/agents-md)
Your own site
<a href="https://agentmods.dev/instructions/taimoorkhan10/replayd/agents-md"><img src="https://agentmods.dev/badge/instructions/taimoorkhan10/replayd/agents-md.svg" alt="Measured on agentmods" height="20"></a>
Per session 4,501 This file is loaded in full into every session.
When invoked 4,501 The same file — it is already loaded in full.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.04501 $0.04501
Opus 5 $0.02250 $0.02250
Sonnet 5 $0.00900 $0.00900
Haiku 4.5 $0.00450 $0.00450

Measured 8d ago against content hash c47a7e617fa5, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-07, from the pricing page.

Security

Grade A, and why

replayd AGENTS.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

AGENTS.md · 551 lines

How it starts

The opening of the file, as written. The whole thing — 551 lines — stays where its author put it; the contents beside it link to each section on GitHub.

replayd — Three Agent Examples

Turn failed AI agent runs into replayable regression tests.

This document covers the three example agents in the examples/ directory and explains how each one demonstrates the replayd release-control loop.


Table of Contents

  1. What replayd does
  2. The four-step loop
  3. Agent integration contract
  4. Example 1 — Multi-step Planning Agent
  5. Example 2 — RAG Policy Agent
  6. Example 3 — Production Incident Response Agent
  7. Running everything
  8. Test results
  9. Grading reference

What replayd does

AI agents regress silently. A team fixes a bug, ships a new prompt or model, and the same bad behavior quietly returns. Traditional software has regression tests and CI/CD to catch this. AI agents have had nothing equivalent.

replayd is the fix. It captures a failed agent run in full — input, output, every tool call — marks it as a known failure, saves it as a regression test, and then replays that test against every future version of the agent before you ship.

capture → mark_failed → save_test → replay_all

If the same bad behavior returns, the release is blocked. That is the entire idea.


The four-step loop

from replayd import Replayd

rp = Replayd()

# 1. Capture a run
with rp.capture(input=user_input, model="gpt-4o") as run:
    run.output = your_agent(user_input, run)

# 2. Mark it as failed
rp.mark_failed(run.id, reason="agent skipped constraint check")

# 3. Save as a regression test
rp.save_test(
    run.id,
    forbidden_actions=["finalize_plan"],
    expected_action="check_constraints",
)

# 4. Replay before every future deployment
results = rp.replay_all(agent=your_agent)
for r in results:
    print(r.verdict, r.reason)

Read the full file on GitHub · 551 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 8d ago First seen · 551 lines · 4,501 tokens per session scan A c47a7e617fa5

Subscribe to this mod's changes

replayd AGENTS.md is an instructions file published in the GitHub repository TaimoorKhan10/replayd (18 stars, last pushed 3mo ago), licensed MIT. It adds 4,501 tokens to every session, about $0.0225 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other instructions, from other repositories

codex AGENTS.md

AGENTS.md instructions for openai/codex, covering rust/codex-rs, the codex-core crate, code review rules, crate api surface and model visible context.

openai/codex · 5,182 tokens

vscode buildNext.instructions.md

Working notes and architecture documentation for the new esbuild-based build system in build/next. Use when making changes to the new build pipeline (transpile/bundle commands, NLS plugin, source-map handling, resource copying, or self-hosting watch tasks).

microsoft/vscode · 6,785 tokens

next.js AGENTS.md

AGENTS.md instructions for vercel/next.js, covering next.js development guide, codebase structure, monorepo overview, core package: packages/next and other important packages.

vercel/next.js · 7,296 tokens

langchain AGENTS.md

AGENTS.md instructions for langchain-ai/langchain, covering global development guidelines for the langchain monorepo, corridor security analysis, project architecture and context, monorepo structure and development tools & commands.

langchain-ai/langchain · 4,469 tokens

vscode oss-third-party-notices.instructions.md

Instructions for microsoft/vscode, covering vs code oss third-party-notices pipeline, architecture, pipeline flow in ci, applying the notice (cutover) and fallback chain (never fail the build).

microsoft/vscode · 5,001 tokens

spec-kit AGENTS.md

AGENTS.md instructions for github/spec-kit, covering agents.md, about spec kit and specify, quickstart — add a new integration in 5 steps, integration architecture and integrationmanifest — file tracking.

github/spec-kit · 7,104 tokens