Claude Code skill: measured laws for authoring hard Terminal-Bench 2 / Harbor benchmark tasks, plus the review-pipeline map for clearing CI in one push
Cursor rule "00-mission" from Xclaw-bot/benchmark-task-authoring, covering the program authoring — always in force, priorities, in order, authorship — do not cross this line, measurements are never invented and read your category brief before you design anything.
Cursor rule "05-retrieval" from Xclaw-bot/benchmark-task-authoring, covering retrieve, do not re-read, cache measurements, never re-derive or re-run them, store: pipe the command's output straight in and reuse, or re-run if missing (3) or stale (4).
Cursor rule "10-hardness-gate" from Xclaw-bot/benchmark-task-authoring, covering the hardness gate, the law, why those conditions and not others, pre-build check — run before writing task files and strong probe — diagnostic, runs last.
Cursor rule "20-design-doctrine" from Xclaw-bot/benchmark-task-authoring, covering design doctrine — what is measured, not theorised, the decisive finding, dead ends — do not retry these, what the accepted corpus looks like and working hypothesis for new designs.
Cursor rule "30-harbor-format" from Xclaw-bot/benchmark-task-authoring, covering harbor format — the mechanical gate, layout, task.toml, dockerfile and before every push.
Cursor rule "40-authoring" from Xclaw-bot/benchmark-task-authoring, covering authoring the prose artifacts, instruction.md, the 4-section proposal, rubric alignment and instruction defects that recur.
Cursor rule "50-verifier" from Xclaw-bot/benchmark-task-authoring, covering verifier craft, ground truth must be unreachable, assertions, tolerances and anti-cheat.
Cursor rule "60-pipeline" from Xclaw-bot/benchmark-task-authoring, covering pipeline and iteration, flow, the difficulty gate, economics — treat runs as scarce and reading a bad result.