Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add instructions/xclaw-bot/benchmark-task-authoring/agents-mdgit clone --depth 1 https://github.com/Xclaw-bot/benchmark-task-authoringWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.06621 | $0.06621 |
| Opus 5 | $0.03311 | $0.03311 |
| Sonnet 5 | $0.01324 | $0.01324 |
| Haiku 4.5 | $0.00662 | $0.00662 |
Grade A, and why
benchmark-task-authoring AGENTS.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 400 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Authoring hard agent-benchmark tasks
Which file am I? The
AGENTS.mdentry point, read automatically by OpenAI Codex, Google Antigravity, and other AGENTS.md-aware agents. Cursor reads the same knowledge from.cursor/rules/; Claude Code fromSKILL.md. All of them are generated from one source — editSKILL.mdorreferences/, then runpython scripts/port.py.
Everything below applies whenever you are designing, building, critiquing or debugging a Terminal-Bench 2 / Harbor benchmark task.
This skill packages what was measured across 30+ Terminal-Bench 2 task slots — the ones that cleared a difficulty gate, and the many that died first. It exists because the intuition almost everyone brings to "write a hard task" is wrong in a specific, repeatable way, and the correction is cheap once you know it.
New here? Fifteen minutes, in this order
1 · Read "The one law" and the kill-list below (5 min). Do not skip to the reference files. If you internalise only one thing, make it the pre-build test: what must the agent build that the verifier could not hand it by simulating?
2 · Set up retrieval (1 min). The field manual is ~35k tokens and will not fit in one read call. Index it once and query it instead — this is the difference between a design question costing ~1k tokens and ~35k:
python scripts/dr.py index --no-embed
python scripts/dr.py ask "what trips ava_review verifier_coverage" --fast
3 · Put a board up (1 min). Before you write anything, see where your slots actually stand — and specifically which review stage each one is sitting at:
python scripts/taskdesk.py --org <your-task-org>
4 · Then, when you have a slot to work: references/ci-stages.md before your first push
(it is the difference between one three-hour CI run and three of them), and the phase table
below for everything else.
The two mistakes that cost the most, stated once so you can avoid both: designing a task where the agent checks a property instead of constructing an object, and treating CI as a test loop instead of a confirmation step.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 400 lines · 6,621 tokens per session scan A d0eafb62816f
benchmark-task-authoring AGENTS.md is an instructions file published in the GitHub repository Xclaw-bot/benchmark-task-authoring (2 stars, last pushed 18d ago), licensed MIT. It adds 6,621 tokens to every session, about $0.0331 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other instructions, from other repositories
spellbook AGENTS.md
Instructions for majiayu000/spellbook, covering spellbook agent contract, routing, scope rules, threads long-run guardrails and validation.
open-supermarkets AGENTS.md
Instructions for abracadabra50/open-supermarkets, covering agent integration guide, supported frameworks, quick integration, 1. add as skill and 2. agent calls commands.
helm copilot-instructions.md
Instructions for PetePeter/helm, covering gamepad-cli-hub — copilot instructions, project purpose, system overview, data flow pipeline and key controls.
flyto-core CLAUDE.md
Instructions for flytohub/flyto-core, covering claude notes, cross-agent handoff and shared code intelligence.
fusion-skills AGENTS.md
AGENTS.md instructions for CrowdStrike/fusion-skills, covering agents.md, what this is, prerequisites, repository structure and skills ecosystem.
aeon CLAUDE.md
Instructions for aeonfun/aeon, covering aeon, how aeon works, strategy, voice and soul file hierarchy (read in this order).