harbor-framework/terminal-bench-1

A benchmark for LLMs on complicated tasks in the terminal

About the project

Terminal-Bench is a benchmark and execution harness for testing whether AI agents can complete difficult, end-to-end tasks in real terminal environments. It is used by developers of language-model agents and benchmarking systems, with tasks covering activities such as compiling code, training models, and setting up servers. Catalogue add-ons help agents work with the benchmark and its task suite.

These files are harbor-framework/terminal-bench-1's own configuration. They tell GitHub Copilot and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

2,564Stars on the repository
2Files it configures its agents with
1,389Tokens loaded in every session
2Agents configured

Instructions