retort CLAUDE.md

retort CLAUDE.md is an instructions file for coding agents from adrianco/retort. It costs 2,211 tokens per session, scanned A, original, Apache-2.0.

Repository instructions for Retort, a project that measures how reliably different coding setups complete tasks. A pass proportion is the share of runs that fully meet a specification.

In plain words
What is it for?
Use them when configuring experiments, checking tuning parameters with smoke tests, organizing experiment code, and publishing results or blog posts.
Why use it?
They help prevent misleading experiment results caused by settings that were recorded incorrectly or never verified.

Instructions file

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add instructions/adrianco/retort/claude-md
Clone the repo
git clone --depth 1 https://github.com/adrianco/retort

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for retort CLAUDE.md

README.md
[![agentmods](https://agentmods.dev/badge/instructions/adrianco/retort/claude-md.svg)](https://agentmods.dev/instructions/adrianco/retort/claude-md)
Your own site
<a href="https://agentmods.dev/instructions/adrianco/retort/claude-md"><img src="https://agentmods.dev/badge/instructions/adrianco/retort/claude-md.svg" alt="Measured on agentmods" height="20"></a>
Per session 2,211 This file is loaded in full into every session.
When invoked 2,211 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.02211 $0.02211
Opus 5 $0.01105 $0.01105
Sonnet 5 $0.00442 $0.00442
Haiku 4.5 $0.00221 $0.00221

Measured 4d ago against content hash 8f843f8dab69, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

retort CLAUDE.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

CLAUDE.md · 145 lines

How it starts

The opening of the file, as written. The whole thing — 145 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Working in this repo

Retort measures whole coding stacks (language × model × quantization × serving layer × agent × context engine × sampling × prompt), scoring each on pass-proportion — the fraction of runs that fully implement the spec. Guidance for running experiments here, and for any Claude session helping with them.

Principle: verify tuning parameters before a full experiment

Before starting any full experiment, RECORD every tuning parameter and VERIFY each one actually takes effect with a smoke test. A parameter set-but-not-verified is worse than none — it produces confident, wrong results.

Nearly every wrong conclusion this project has published came from a tuning parameter that was set-but-not-verified, or never recorded at all:

  • temperature = 1.0 (the oMLX default, never recorded) — cost roughly half the local reliability. "The 35B scores 0.38" really meant "0.38 at temp 1.0".
  • playpen under /var — the agent's file tool was silently refused (macOS temp dir is a "sensitive system path"), so it wrote nothing and scored a false zero, indistinguishable from a model that can't do the task. Read as a "capability wall".
  • context silently 128K, not 256K — the stack-reload hook destroyed Hermes' per-model context_length; the config file AND provenance.json both still reported 262144 while the model actually ran at half that.
  • repetition_penalty — derailed the agentic tool loop into stalls, even at the value the model's own card recommended. Model-card sampling is tuned for single-turn generation, not multi-turn agent loops.
  • lcm context_threshold = 0.35 — the real ~92K compaction ceiling (0.35 × 262144), mistaken for a residual 128K bug until traced.

How to apply:

  1. Record every parameter that could move the result: sampling (temperature / top_p / top_k / penalties), context length and the agent's compaction threshold, the serving-layer settings, the model revision hash (not just its name), and the harness settings (playpen path, timeout, stall guard). This is what provenance.json captures — and it reports the effective value, because the config file's value and the value the model actually ran at have diverged.
  2. Verify each takes effect with a cheap smoke test before the full grid: send a probe and confirm the server/agent honoured it — temp=0 → byte-identical output; settings survive a restart; the live context actually reaches the configured window. "I set it" is not "it took effect": oMLX silently STRIPS unsupported keys (e.g. min_p) and IGNORES others.
  3. A parameter whose effect you cannot observe in a smoke test is not usable in the experiment — fix the plumbing or drop the factor.

Read the full file on GitHub · 145 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 145 lines · 2,211 tokens per session scan A 8f843f8dab69

Subscribe to this mod's changes

retort CLAUDE.md is an instructions file published in the GitHub repository adrianco/retort (199 stars, last pushed 5d ago), licensed Apache-2.0. It adds 2,211 tokens to every session, about $0.0111 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other instructions, from other repositories

vscode buildNext.instructions.md

Working notes and architecture documentation for the new esbuild-based build system in build/next. Use when making changes to the new build pipeline (transpile/bundle commands, NLS plugin, source-map handling, resource copying, or self-hosting watch tasks).

microsoft/vscode · 6,785 tokens

spec-kit AGENTS.md

AGENTS.md instructions for github/spec-kit, covering agents.md, about spec kit and specify, quickstart — add a new integration in 5 steps, integration architecture and integrationmanifest — file tracking.

github/spec-kit · 7,104 tokens

codex AGENTS.md

AGENTS.md instructions for openai/codex, covering rust/codex-rs, the codex-core crate, code review rules, crate api surface and model visible context.

openai/codex · 5,182 tokens

langchain AGENTS.md

AGENTS.md instructions for langchain-ai/langchain, covering global development guidelines for the langchain monorepo, corridor security analysis, project architecture and context, monorepo structure and development tools & commands.

langchain-ai/langchain · 4,345 tokens

vscode oss-third-party-notices.instructions.md

Instructions for microsoft/vscode, covering vs code oss third-party-notices pipeline, architecture, pipeline flow in ci, applying the notice (cutover) and fallback chain (never fail the build).

microsoft/vscode · 5,001 tokens

next.js AGENTS.md

Instructions for vercel/next.js, covering next.js development guide, codebase structure, monorepo overview, core package: packages/next and other important packages.

vercel/next.js · 7,296 tokens