Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains
Latest release v0.1.0 · 26 Aug 2026
These files are harbor-framework/terminal-bench-science's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 12,599 tok CLAUDE.md A 3 tok .claude/skills/convert-separate-verifier/SKILL.md F 63 tok .claude/skills/review-task/SKILL.md A 26 tok .claude/skills/update-rubric/SKILL.md A 20 tok