benchmark instructions

75 tagged benchmark, measured the same way as everything else here.

Browse within: Evaluation 12dataset 5llm-evaluation 5performance 5

itayinbarr/little-coder

Instructions file CodexOpenCode

Instructions for itayinbarr/little-coder, covering little-coder, capabilities & autonomy, runtime invariants, available tools and file & shell.

2.5k 2d ago C 1,788 tokens original Apache-2.0

benchflow-ai/skillsbench

Instructions file CodexOpenCode

Instructions for benchflow-ai/skillsbench, covering skillsbench, commands, task layout, rules and key references.

1.7k 1mo ago A 692 tokens original Apache-2.0

benchflow-ai/skillsbench

Instructions file

Instructions for benchflow-ai/skillsbench: See AGENTS.md for project overview and contribution guidelines.

1.7k 1mo ago A 17 tokens original Apache-2.0

pawurb/hotpath-rs

Instructions file CodexOpenCode

Instructions for pawurb/hotpath-rs, covering agents.md, project overview, reference docs, development commands and profiling modes are combined via features, e.g.

1.7k 2d ago A 2,336 tokens original MIT

pawurb/hotpath-rs

Instructions file

Instructions for pawurb/hotpath-rs, a project described as: Quickly find bottlenecks in Rust - one profiler for CPU, memory, SQL, HTTP, I/O and async code.

1.7k 2d ago A 3 tokens copy · 100% MIT

OpenCUA CLAUDE.md

06

xlang-ai/OpenCUA

Instructions file

Instructions for xlang-ai/OpenCUA, covering claude.md, project overview, commands, environment setup and model inference dependencies.

830 3mo ago B 1,409 tokens original MIT

mobilegym AGENTS.md

07

Purewhiter/mobilegym

Instructions file CodexOpenCode

Instructions for Purewhiter/mobilegym, covering agents.md, project overview, type-checking strategy, eslint and with data graph generation.

777 4d ago A 6,877 tokens original Apache-2.0

ClawBench AGENTS.md

09

TIGER-AI-Lab/ClawBench

Instructions file CodexOpenCode

Instructions for TIGER-AI-Lab/ClawBench, covering clawbench -- agent context, what this is, project structure, setup and 2. configure at least one model.

609 9d ago A 1,679 tokens original Apache-2.0

lynote-ai/ai-text-detector

Instructions file CodexOpenCode

Instructions for lynote-ai/ai-text-detector, covering agent instructions: ai detector skill, when to use, required behavior, how to run and response format for users.

440 27d ago A 212 tokens original MIT

benchflow AGENTS.md

11

benchflow-ai/benchflow

Instructions file CodexOpenCode

Instructions for benchflow-ai/benchflow, covering benchflow, setup + test, conventions and skill catalog (.agents/skills, mirrored at .claude/skills).

335 2d ago A 1,663 tokens original Apache-2.0

benchflow CLAUDE.md

12

benchflow-ai/benchflow

Instructions file

Instructions for benchflow-ai/benchflow, a project described as: Research infra for creating RL environments, post-training, and evals.

335 2d ago A 11 tokens copy · 88% Apache-2.0

hogeheer499-commits/strix-halo-guide

Instructions file CodexOpenCode

Instructions for hogeheer499-commits/strix-halo-guide, covering agents.md, core project lens, documentation style, do-not-invent rules and benchmark claim rules.

312 2d ago A 2,053 tokens original MIT

codspeed AGENTS.md

14

CodSpeedHQ/codspeed

Instructions file CodexOpenCode

Instructions for CodSpeedHQ/codspeed, covering agents.md, common development commands, building and testing, build the project and build in release mode.

280 3d ago A 663 tokens original Apache-2.0

codspeed CLAUDE.md

15

CodSpeedHQ/codspeed

Instructions file

Instructions for CodSpeedHQ/codspeed, a project described as: CodSpeed is the all-in-one performance testing toolkit. Optimize code performance and catch regressions early.

280 3d ago A 3 tokens copy · 100% Apache-2.0

gregs1104/pgbent

Instructions file GitHub Copilot

Copilot instructions for gregs1104/pgbent: This file exists for GitHub Copilot compatibility. The canonical agent guide is AGENTS.md at the repository root.

269 16d ago A 76 tokens

pgbent AGENTS.md

17

gregs1104/pgbent

Instructions file CodexOpenCode

AGENTS.md instructions for gregs1104/pgbent, covering agent guide, architecture, workflows, conventions and where to look first.

269 16d ago A 1,168 tokens

robocurve/inspect-robots

Instructions file

Instructions for robocurve/inspect-robots, covering inspect robots — agent guide, the one big idea, layout, working here and out of scope (separate repos / plugins).

191 12d ago A 2,035 tokens original MIT

openbench AGENTS.md

19

minghinmatthewlam/openbench

Instructions file CodexOpenCode

Instructions for minghinmatthewlam/openbench, covering openbench — agent context, local execution context, what openbench is, execution ownership and product goals (the two things we are building toward).

130 5d ago A 2,496 tokens original MIT

Raidriar7170/hermes-skilleval

Instructions file CodexOpenCode

Instructions for Raidriar7170/hermes-skilleval, a project described as: Verification-gated skill routing and self-improvement harness for Hermes-style agent skills.

127 1mo ago A 363 tokens original MIT

aws-bench AGENTS.md

21

aws-bench/aws-bench

Instructions file CodexOpenCode

Instructions for aws-bench/aws-bench, covering agents.md, setup, commands, project structure and code style.

103 3d ago A 1,084 tokens original Apache-2.0

aws-bench CLAUDE.md

22

aws-bench/aws-bench

Instructions file

Instructions for aws-bench/aws-bench, a project described as: aws-bench measures how well AI agents and model combinations perform on real AWS work — diagnosing misconfigurations, provisioning infrastructure, and operating live cloud environments.

103 3d ago A 3 tokens copy · 100% Apache-2.0

PhiloLabs/agentic-vbench

Instructions file

Instructions for PhiloLabs/agentic-vbench, a project described as: AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?

89 3d ago A 3 tokens copy · 100% Apache-2.0