Xclaw-bot/benchmark-task-authoring

Claude Code skill: measured laws for authoring hard Terminal-Bench 2 / Harbor benchmark tasks, plus the review-pipeline map for clearing CI in one push

2Stars on the repository
10Mods indexed here, across every type
18d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

Xclaw-bot/benchmark-task-authoring

Instructions file CodexOpenCode

Instructions for Xclaw-bot/benchmark-task-authoring, covering authoring hard agent-benchmark tasks, new here? fifteen minutes, in this order, the one law, before designing anything: run the kill-list and translating your domain into a row.

2 18d ago A 6,621 tokens original MIT