pku-liang

10 mods across 1 repository, 49 stars between them.

runtime_proof

01

pku-liang/hwe-bench

MCP server Claude CodeCodexCursor +2

MCP server "runtimeproof", hosted remotely at runtime-mcp-server, as configured in pku-liang/hwe-bench.

49 1mo ago A tokens not measured copy · 100% Apache-2.0

bundled-keep

02

pku-liang/hwe-bench

Skill Claude CodeCodex

Existing task skill that should remain after job-level skill injection.

49 1mo ago A 17 tokens copy · 100% Apache-2.0

runtime-proof

03

pku-liang/hwe-bench

Skill Claude CodeCodex

Write the proof file for the Harbor runtime skill injection example.

49 1mo ago A 15 tokens copy · 100% Apache-2.0

create-adapter

05

pku-liang/hwe-bench

Skill Claude CodeCodex

Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.

49 1mo ago A 34 tokens copy · 100% Apache-2.0

create-task

06

pku-liang/hwe-bench

Skill Claude CodeCodex

Create a new Harbor task for evaluating agents. Use when the user wants to scaffold, build, or design a new task, benchmark problem, or eval. Guides through instruction writing, environment setup, verifier design (pytest vs Reward Kit vs custom), and solution scripting.

49 1mo ago F 56 tokens copy · 100% Apache-2.0

harbor-exec

07

pku-liang/hwe-bench

Skill Claude CodeCodex

Use when working with Harbor's harbor exec CLI workflow: compiling files, directories, or globs into Harbor tasks; running map jobs; configuring artifacts and existence-only verification; using map-reduce; writing or reviewing ExecConfig YAML/JSON/TOML; or debugging command behavior, config validation, and job outputs.

49 1mo ago A 71 tokens copy · 100% Apache-2.0

publish

08

pku-liang/hwe-bench

Skill Claude CodeCodex

Publish a Harbor task or dataset to the registry. Use when the user wants to upload, publish, or share tasks or datasets/benchmarks on the Harbor registry.

49 1mo ago A 35 tokens copy · 95% Apache-2.0

rewardkit

09

pku-liang/hwe-bench

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

49 1mo ago A 46 tokens copy · 91% Apache-2.0

pku-liang/hwe-bench

Skill Claude CodeCodex

Create or reuse Hugging Face dataset PRs for harborframework/parity-experiments and upload Harbor parity/oracle result folders efficiently with sparse checkout, raw git pushes, and Git LFS.

49 1mo ago A 49 tokens copy · 100% Apache-2.0