Xclaw-bot/benchmark-task-authoring
Skill Claude CodeCodex
Design, red-team, ship and debug hard Terminal-Bench 2 / Harbor benchmark tasks: the measured laws for what makes agents actually fail, the kill-list of dead task shapes, and how to clear all 17 review stages in one push instead of three. Use for benchmark task slots, TB2/Harbor tasks, task.toml, instruction.md, task…
2 17d ago A 207 tokens
original MIT