Xclaw-bot/benchmark-task-authoring

Claude Code skill: measured laws for authoring hard Terminal-Bench 2 / Harbor benchmark tasks, plus the review-pipeline map for clearing CI in one push

2Stars on the repository
10Mods indexed here, across every type
17d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

Xclaw-bot/benchmark-task-authoring

Skill Claude CodeCodex

Design, red-team, ship and debug hard Terminal-Bench 2 / Harbor benchmark tasks: the measured laws for what makes agents actually fail, the kill-list of dead task shapes, and how to clear all 17 review stages in one push instead of three. Use for benchmark task slots, TB2/Harbor tasks, task.toml, instruction.md, task…

2 17d ago A 207 tokens original MIT