terminal-bench skills

15 tagged terminal-bench, measured the same way as everything else here.

Browse within: evals 9rl-environments 9fable5 6long-horizon-agents 6long-horizon-tasks 6long-horizon-terminal-bench 6

create-adapter

01

harbor-framework/harbor

Skill Claude CodeCodex

Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.

4.9k +72 today A 34 tokens original Apache-2.0

rewardkit

02

harbor-framework/harbor

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

4.9k +72 changed today A 46 tokens original Apache-2.0

harbor-framework/harbor

Skill Claude CodeCodex

Create or reuse Hugging Face dataset PRs for harborframework/parity-experiments and upload Harbor parity/oracle result folders efficiently with sparse checkout, raw git pushes, and Git LFS.

4.9k +72 today A 49 tokens original Apache-2.0

create-adapter

04

zli12321/LHTB

Skill Claude CodeCodex

Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.

695 +5 5d ago A 34 tokens copy · 92% Apache-2.0

publish

05

zli12321/LHTB

Skill Claude CodeCodex

Publish a Harbor task or dataset to the registry. Use when the user wants to upload, publish, or share tasks or datasets/benchmarks on the Harbor registry.

695 +5 5d ago A 35 tokens copy · 95% Apache-2.0

rewardkit

06

zli12321/LHTB

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

695 +5 5d ago A 46 tokens copy · 91% Apache-2.0