benchmark skills

724 tagged benchmark, measured the same way as everything else here.

Browse within: automatic 200continual-learning 200skill-generation 200skillsbench 197agent-evaluation 89ci 86llm-agents 84python-cli 83release-gate 83Evaluation 61bun 42game-development 39harness 39leaderboard 38

ChainBench/OpenChainBench

Skill Claude CodeCodex

Walks a contributor through adding a new benchmark to OpenChainBench. Covers spec format, harness contract, Prometheus scrape wiring, local validation, and PR conventions.

not rated 7 today A 39 tokens original MIT

metrillm

27

MetriLLM/metrillm

Skill Claude CodeCodex

Find the best local LLM for your machine. Tests speed, quality and RAM fit, then tells you if a model is worth running on your hardware.

not rated 6 3mo ago A 35 tokens original Apache-2.0

Yinzhanqing/embodied-eval-automation

Skill Claude CodeCodex

Plan, explain, build, run, monitor, validate, transfer, and audit reproducible embodied-model studies and batch episode collection. Use when a user wants to connect a local, SSH, or cloud GPU host; understand and compare a policy, VLA, world model, world-action model, or hybrid with a benchmark; discover and reuse…

not rated 5 1mo ago A 145 tokens original MIT

audit-post

29

DanceNitra/agora

Skill Claude CodeCodex

The end-to-end procedure for auditing one already-published post to a scientific-organization standard. It chains the adversarial skills and — critically — re-runs the auditor on the CORRECTED post to confirm it now passes clean before committing. We are a scientific organization; we do not ship missteps. Run EVERY…

not rated 3 yesterday A 0 tokens original MIT

cosmergon

30

rkocosmergon/cosmergon-agent

Skill Claude CodeCodex

Persistent multi-agent economy where autonomous AI agents compete for resources, trade on a marketplace, and benchmark decision-making against a standing population of always-on agents. Invite other agents for energy rewards. Auto-registers — no API key needed.

not rated 3 10d ago A 50 tokens original MIT

xorcise-playbooks

31

xorcise-ai/xorcise-skills

Skill Claude CodeCodex

Run a standardized XORCISE agent benchmark — pick a ready-made playbook (or build your own from existing missions), point it at any OpenHands-supported model(s), and get a polished eval-card HTML report of how each model performed across the missions. Warns you about cost before spending anything.

not rated 2 1mo ago B 66 tokens

using-shakespii

32

ai-creed/ai-shakespii

Skill Claude CodeCodex

Use when the user asks to lint, audit, test, benchmark, validate, or fix an agent skill — from a single SKILL.md frontmatter check to trigger-accuracy measurement or a corpus-wide audit of installed skills for duplication — driving the shakespii CLI (init, lint --json, test --run, bench) to resolve findings until…

not rated 2 1mo ago A 78 tokens original MIT

unreal-mcp

33

44-99/unreal-agent-benchmark

Skill Claude CodeCodex

Use this skill to perform actions inside an Unreal Engine project via a live-editor MCP connection. Trigger when the user wants to change, query, or run something in their Unreal Engine project, not for conceptual or docs questions. Concrete triggers: spawn/move/duplicate/transform actors in a level, open a .uproject…

not rated 2 1mo ago A 266 tokens

adversarial-review

34

theMobiusStrip/agentpay-guard

Skill Claude CodeCodex

Independent adversarial security review of a custody-spine change against SECURITY.md. Spawns a FRESH reviewer subagent that did NOT write the code, prompted to break the diff (cap overspend, window-slide, early release, fail-open void-return, TOCTOU, dedup/replay bypass, honest-scope overclaim, DrainBench fairness…

not rated 2 1mo ago A 143 tokens original MIT

self-improve

35

ttxs69/coding-agent-eval

Skill Claude CodeCodex

Use when the user wants to proactively push a project forward by implementing its next feature. Reads project artifacts (specs' "future work" sections, README, TODOs, recent commit trajectory, codebase gaps), infers candidate features, asks the user to pick one, then implements it on an isolated branch with…

not rated 2 2mo ago A 99 tokens

eco-max

36

sup3x/codex-eco

Skill Claude CodeCodex

Maximum-savings variant of eco - the same frugality rules with a tighter reply budget, for routine chores (rename, small fix, quick lookup, boilerplate). Prefer plain eco for hard or high-stakes work. Works in any language.

not rated 2 18d ago A 53 tokens original MIT

JoniMartin27/inferbench

Skill Claude CodeCodex

Build, launch, smoke-test and drive the InferBench FastAPI backend (:7777) and its inference-engine orchestration. Use when asked to run, start, launch, boot, smoke-test, benchmark, or drive the inferbench backend / API / engines (llama.cpp, ollama, vLLM, SGLang, TGI), or to verify an engine actually starts and runs a…

not rated 2 10d ago A 90 tokens original MIT

patchleague

38

269394628/PatchLeague

Skill Claude CodeCodex

Run, inspect, and compare multiple coding-agent CLIs on the same repository task with PatchLeague. Use when Codex should benchmark agents such as Codex, Claude Code, Gemini CLI, or OpenCode in isolated Git worktrees; verify their patches with tests; inspect saved runs or HTML reports; or preview and apply a selected…

not rated 1 1mo ago A 72 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: