Claude Code skill: measured laws for authoring hard Terminal-Bench 2 / Harbor benchmark tasks, plus the review-pipeline map for clearing CI in one push
1 file for Claude Code, Codex and OpenCode: benchmark-task-authoring AGENTS.md — 6,621 tokens loaded in every session.
AGENTS.md A 6,621 tok These files are Xclaw-bot/benchmark-task-authoring's own configuration — they tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.