Instructions file CodexOpenCode
Instructions for itayinbarr/little-coder, covering little-coder, capabilities & autonomy, runtime invariants, available tools and file & shell.
75 tagged benchmark, measured the same way as everything else here.
Browse within: Evaluation 12dataset 5llm-evaluation 5performance 5
Instructions file CodexOpenCode
Instructions for itayinbarr/little-coder, covering little-coder, capabilities & autonomy, runtime invariants, available tools and file & shell.
Instructions file CodexOpenCode
Instructions for benchflow-ai/skillsbench, covering skillsbench, commands, task layout, rules and key references.
Instructions file
Instructions for benchflow-ai/skillsbench: See AGENTS.md for project overview and contribution guidelines.
Instructions file CodexOpenCode
Instructions for pawurb/hotpath-rs, covering agents.md, project overview, reference docs, development commands and profiling modes are combined via features, e.g.
Instructions file
Instructions for pawurb/hotpath-rs, a project described as: Quickly find bottlenecks in Rust - one profiler for CPU, memory, SQL, HTTP, I/O and async code.
Instructions file
Instructions for xlang-ai/OpenCUA, covering claude.md, project overview, commands, environment setup and model inference dependencies.
Instructions file CodexOpenCode
Instructions for Purewhiter/mobilegym, covering agents.md, project overview, type-checking strategy, eslint and with data graph generation.
Instructions file
Instructions for Purewhiter/mobilegym, a project described as:
Instructions file CodexOpenCode
Instructions for TIGER-AI-Lab/ClawBench, covering clawbench -- agent context, what this is, project structure, setup and 2. configure at least one model.
Instructions file CodexOpenCode
Instructions for lynote-ai/ai-text-detector, covering agent instructions: ai detector skill, when to use, required behavior, how to run and response format for users.
Instructions file CodexOpenCode
Instructions for benchflow-ai/benchflow, covering benchflow, setup + test, conventions and skill catalog (.agents/skills, mirrored at .claude/skills).
Instructions file
Instructions for benchflow-ai/benchflow, a project described as: Research infra for creating RL environments, post-training, and evals.
hogeheer499-commits/strix-halo-guide
Instructions file CodexOpenCode
Instructions for hogeheer499-commits/strix-halo-guide, covering agents.md, core project lens, documentation style, do-not-invent rules and benchmark claim rules.
Instructions file CodexOpenCode
Instructions for CodSpeedHQ/codspeed, covering agents.md, common development commands, building and testing, build the project and build in release mode.
Instructions file
Instructions for CodSpeedHQ/codspeed, a project described as: CodSpeed is the all-in-one performance testing toolkit. Optimize code performance and catch regressions early.
Instructions file GitHub Copilot
Copilot instructions for gregs1104/pgbent: This file exists for GitHub Copilot compatibility. The canonical agent guide is AGENTS.md at the repository root.
Instructions file CodexOpenCode
AGENTS.md instructions for gregs1104/pgbent, covering agent guide, architecture, workflows, conventions and where to look first.
Instructions file
Instructions for robocurve/inspect-robots, covering inspect robots — agent guide, the one big idea, layout, working here and out of scope (separate repos / plugins).
Instructions file CodexOpenCode
Instructions for minghinmatthewlam/openbench, covering openbench — agent context, local execution context, what openbench is, execution ownership and product goals (the two things we are building toward).
Instructions file CodexOpenCode
Instructions for Raidriar7170/hermes-skilleval, a project described as: Verification-gated skill routing and self-improvement harness for Hermes-style agent skills.
Instructions file CodexOpenCode
Instructions for aws-bench/aws-bench, covering agents.md, setup, commands, project structure and code style.
Instructions file
Instructions for aws-bench/aws-bench, a project described as: aws-bench measures how well AI agents and model combinations perform on real AWS work — diagnosing misconfigurations, provisioning infrastructure, and operating live cloud environments.
Instructions file CodexOpenCode
Instructions for PhiloLabs/agentic-vbench, covering agents.md, start, map, docs and code.
Instructions file
Instructions for PhiloLabs/agentic-vbench, a project described as: AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?