Framework for evaluating and improving agents
Harbor is a framework for evaluating and improving AI agents and language models across reusable benchmarks and execution environments. Developers and researchers use it to run agent evaluations, create benchmarks, perform parallel experiments, and generate rollouts for reinforcement-learning optimization, with the catalogue offering skills, instructions, and an MCP integration for the framework.
Latest release v0.22.0 · 22 Aug 2026
These files are harbor-framework/harbor's own configuration. They tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 4,097 tok CLAUDE.md A 3 tok