Instructions file CodexOpenCode
Instructions for harbor-framework/harbor, covering claude.md - harbor framework, contributing, project overview, quick start commands and install.
Instructions file CodexOpenCode
Instructions for harbor-framework/harbor, covering claude.md - harbor framework, contributing, project overview, quick start commands and install.
Instructions file
Instructions for harbor-framework/harbor, a project described as: Framework for evaluating and improving agents.
MCP server Claude CodeCodexCursor +2
MCP server "runtimeproof", hosted remotely at runtime-mcp-server, as configured in harbor-framework/harbor.
Skill Claude CodeCodex
Existing task skill that should remain after job-level skill injection.
Skill Claude CodeCodex
Write the proof file for the Harbor runtime skill injection example.
Skill Claude CodeCodex
Generate a greeting message and write it to a file.
Skill Claude CodeCodex
Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.
Skill Claude CodeCodex
Create a new Harbor task for evaluating agents. Use when the user wants to scaffold, build, or design a new task, benchmark problem, or eval. Guides through instruction writing, environment setup, verifier design (pytest vs Reward Kit vs custom), and solution scripting.
Skill Claude CodeCodex
Use when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only verification; using map-reduce; writing or reviewing ExecConfig YAML/JSON/TOML; or debugging command behavior, config validation, and job outputs.
Skill Claude CodeCodex
Publish a Harbor task or dataset to the registry. Use when the user wants to upload, publish, or share tasks or datasets/benchmarks on the Harbor registry.
Skill Claude CodeCodex
Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.
Skill Claude CodeCodex
Create or reuse Hugging Face dataset PRs for harborframework/parity-experiments and upload Harbor parity/oracle result folders efficiently with sparse checkout, raw git pushes, and Git LFS.
harbor-framework/terminal-bench-1
Instructions file GitHub Copilot
Instructions for harbor-framework/terminal-bench-1: If the pull request implements a task for the Terminal-Bench benchmark (creates files in the tasks/ directory), then please do the following review of the task.
harbor-framework/terminal-bench-1
Instructions file
Instructions for harbor-framework/terminal-bench-1, covering claude.md, architecture, core components, common development commands and running tasks.
harbor-framework/terminal-bench-science
Skill Claude CodeCodex
Convert a Harbor benchmark task from Harbor's shared verifier mode (default) to separate verifier mode. Use when the user asks to "convert this task to separate verifier", "make the verifier run in its own container", or asks about Harbor's separate-verifier environment for a specific task.
harbor-framework/terminal-bench-science
Skill Claude CodeCodex
Review a benchmark task PR — downloads all artifacts, analyzes against the rubric, launches harbor view, and enables interactive review.
harbor-framework/terminal-bench-science
Skill Claude CodeCodex
Propose new rubric criteria based on review findings — creates a branch and PR.
harbor-framework/terminal-bench-science
Instructions file CodexOpenCode
Instructions for harbor-framework/terminal-bench-science, covering terminal-bench-science, repo structure, key documentation, for agents reviewing task prs and template relationship.
harbor-framework/terminal-bench-science
Instructions file
Instructions for harbor-framework/terminal-bench-science, a project described as: Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains.
Plugin Claude Code
Plugin marketplace listing 1 plugin: harbor-skills.
Instructions file CodexOpenCode
Instructions for harbor-framework/skills, covering harbor-skills — agent quick reference, skill format quick reference, harbor documentation map, conventions and common pitfalls.
Instructions file
Instructions for harbor-framework/skills, a project described as: Public agent skills catalog for Harbor.
Skill Claude CodeCodex
Create Harbor benchmark adapters that convert external benchmark datasets into Harbor task format. Use when porting an existing benchmark to Harbor, running parity experiments, registering a dataset to the Harbor registry, or debugging adapter validation failures. Covers: adapter class interface (generatetask…
Skill Claude CodeCodex
Harbor CLI command reference and usage patterns. Covers harbor run, harbor jobs, harbor trials, harbor datasets, harbor adapters, harbor tasks, harbor view, harbor sweeps, harbor traces, harbor cache, and harbor admin commands. Use this skill whenever running Harbor evaluations, managing datasets, viewing results…