Cursor rule "00-mission" from Xclaw-bot/benchmark-task-authoring, covering the program authoring — always in force, priorities, in order, authorship — do not cross this line, measurements are never invented and read your category brief before you design anything.
Cursor rule "05-retrieval" from Xclaw-bot/benchmark-task-authoring, covering retrieve, do not re-read, cache measurements, never re-derive or re-run them, store: pipe the command's output straight in and reuse, or re-run if missing (3) or stale (4).
Cursor rule "10-hardness-gate" from Xclaw-bot/benchmark-task-authoring, covering the hardness gate, the law, why those conditions and not others, pre-build check — run before writing task files and strong probe — diagnostic, runs last.
Cursor rule "20-design-doctrine" from Xclaw-bot/benchmark-task-authoring, covering design doctrine — what is measured, not theorised, the decisive finding, dead ends — do not retry these, what the accepted corpus looks like and working hypothesis for new designs.
Cursor rule "30-harbor-format" from Xclaw-bot/benchmark-task-authoring, covering harbor format — the mechanical gate, layout, task.toml, dockerfile and before every push.
Cursor rule "40-authoring" from Xclaw-bot/benchmark-task-authoring, covering authoring the prose artifacts, instruction.md, the 4-section proposal, rubric alignment and instruction defects that recur.
Cursor rule "50-verifier" from Xclaw-bot/benchmark-task-authoring, covering verifier craft, ground truth must be unreachable, assertions, tolerances and anti-cheat.
Cursor rule "60-pipeline" from Xclaw-bot/benchmark-task-authoring, covering pipeline and iteration, flow, the difficulty gate, economics — treat runs as scarce and reading a bad result.
Instructions for Xclaw-bot/benchmark-task-authoring, covering authoring hard agent-benchmark tasks, new here? fifteen minutes, in this order, the one law, before designing anything: run the kill-list and translating your domain into a row.
Design, red-team, ship and debug hard Terminal-Bench 2 / Harbor benchmark tasks: the measured laws for what makes agents actually fail, the kill-list of dead task shapes, and how to clear all 17 review stages in one push instead of three. Use for benchmark task slots, TB2/Harbor tasks, task.toml, instruction.md, task…