benchmark skills

432 tagged benchmark, measured the same way as everything else here.

Browse within: agent-evaluation 65ci 61Evaluation 60automatic 60continual-learning 60llm-agents 60skill-generation 60python-cli 59release-gate 59skillsbench 58bun 42game-development 39harness 39leaderboard 38

pubchem-database

01

benchflow-ai/skillsbench

Skill Claude CodeCodex

Query PubChem via PUG-REST API/PubChemPy (110M+ compounds). Search by name/CID/SMILES, retrieve properties, similarity/substructure searches, bioactivity, for cheminformatics.

1.7k 1mo ago A 49 tokens original Apache-2.0

rdkit

02

benchflow-ai/skillsbench

Skill Claude CodeCodex

Cheminformatics toolkit for fine-grained molecular control. SMILES/SDF parsing, descriptors (MW, LogP, TPSA), fingerprints, substructure search, 2D/3D generation, similarity, reactions. For standard workflows with simpler interface, use datamol (wrapper around RDKit). Use rdkit for advanced control, custom…

1.7k 1mo ago A 80 tokens original Apache-2.0

gpt-multimodal

03

benchflow-ai/skillsbench

Skill Claude CodeCodex

Analyze images and multi-frame sequences using OpenAI GPT series.

1.7k 1mo ago A 18 tokens original Apache-2.0

hotpath_bump

04

pawurb/hotpath-rs

Skill Claude CodeCodex

Bump the hotpath version number across the workspace and related files. Updates crate versions in Cargo.toml files (exact patch version) and version references in the backend middleware, hotpathinit skill, and README (major.minor only). Use when the user wants to bump, bump the version, or release a new hotpath…

1.7k 2d ago A 73 tokens original MIT

hotpath_init

05

pawurb/hotpath-rs

Skill Claude CodeCodex

Configure hotpath profiling in a Rust project. Adds the hotpath dependency with feature-gated setup, instruments main with hotpath::main, functions with measure/measureall, and wraps channels, mutexes, rwlocks, streams, futures, reqwest clients, axum routers and byte-level I/O with hotpath macros. Use when the user…

1.7k 2d ago A 88 tokens original MIT

syncmeta

06

pawurb/hotpath-rs

Skill Claude CodeCodex

Sync changes from hotpath and hotpath-macros crates to their meta counterparts (hotpath-meta and hotpath-macros-meta). Use when meta crates need to be updated with recent changes.

1.7k 2d ago A 43 tokens original MIT

ceo-setup

08

suyoumo/ClawProBench

Skill Claude CodeCodex

One-time onboarding for the executive/manager commitment workflow — delegation-heavy, meeting prep, decision capture, morning and evening digests. Creates a commitments project and installs two dashboard widgets. After successful setup this skill is excluded from selection until the marker file is deleted.

823 7d ago A 60 tokens original Apache-2.0

developer-setup

09

suyoumo/ClawProBench

Skill Claude CodeCodex

One-time onboarding for the developer workflow — installs github-workflow missions, creates the commitments workspace, registers per-repo projects, writes calibration memories. After successful setup this skill is excluded from selection until the marker file is deleted.

823 7d ago A 49 tokens original Apache-2.0

portfolio

10

suyoumo/ClawProBench

Skill Claude CodeCodex

Cross-chain DeFi portfolio discovery, rebalancing suggestions, and NEAR Intent construction. Activates when the user pastes a wallet address or asks about yield/positions/rebalancing. Bootstraps a per-user "portfolio" project, aggregates positions across all the user's addresses inside one project, and offers a…

823 7d ago A 69 tokens original Apache-2.0

Purewhiter/mobilegym

Skill Claude CodeCodex

Use when designing a new benchenv task suite, adding several new tasks to an existing suite, or critiquing a task-set proposal for a mobile-gym App — before any class FooTask(...) is written under benchenv/task/.

777 4d ago A 57 tokens original Apache-2.0

testing-bench-task

12

Purewhiter/mobilegym

Skill Claude CodeCodex

Use when adding or modifying offline judge tests for benchenv tasks — specifically entries in OFFLINEJUDGEPOSITIVECASES / OFFLINEJUDGENEGATIVECASES in benchenv/tests/ /testtasks.py, or writing live tests. Triggers after a new task is added, or when tightening judge coverage.

777 4d ago A 74 tokens original Apache-2.0

Purewhiter/mobilegym

Skill Claude CodeCodex

Use when writing or modifying checkgoals() / getanswer() / App check methods in benchenv/task/, or when reviewing a draft task's judge correctness. Triggers include adding a new task, editing a judge method, or diagnosing a judge false-positive/negative.

777 4d ago A 68 tokens original Apache-2.0

ai-detector

14

lynote-ai/ai-text-detector

Skill Claude CodeCodex

Evidence-based AI-generated text risk analysis for essays, emails, reviews, articles, messages, and other prose samples.

440 27d ago A 27 tokens original MIT

lynote-ai/ai-text-detector

Skill Claude CodeCodex

Detect whether a passage shows AI-like writing signals and return an explainable risk estimate with confidence, caveats, and next steps.

440 27d ago A 33 tokens original MIT

game-ai

16

ukanwat/aaabench

Skill Claude CodeCodex

Design NPC and enemy decision-making with finite state machines, behavior trees, steering behaviors, and A pathfinding — engine-neutral algorithms that pair with the detected engine's navigation API. Use when building enemy AI, an FSM or behavior tree, steering/flocking, or pathfinding, or when the user mentions state…

378 17d ago A 85 tokens copy · 88% MIT

reference-images

17

ukanwat/aaabench

Skill Claude CodeCodex

Find and actually LOOK at real photographs — keyless image-search APIs, downloaded to disk so they render as images. Use before building any place, material, vehicle, sky or lighting condition, and again when judging your own screenshots.

378 17d ago A 49 tokens original MIT

ukanwat/aaabench

Skill Claude CodeCodex

Set up player input in Unreal Engine 5 with Enhanced Input: Input Actions, Input Mapping Contexts, modifiers and triggers, adding the mapping context, and binding actions by ETriggerEvent. Use when wiring movement/look/jump input, creating IA/IMC assets, binding in C++ or Blueprints, or when the user mentions Enhanced…

378 17d ago A 94 tokens original MIT

benchflow-ai/benchflow

Skill Claude CodeCodex

Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. Use this skill whenever the user asks to audit traj health, failed or timed-out runs, healthy pass/fail/timeout status, no-skill leakage, skill loading, reward hacking, verifier isolation, metadata completeness, token…

335 2d ago A 99 tokens original Apache-2.0

benchflow-ai/benchflow

Skill Claude CodeCodex

Find a local Claude Code or Codex session, open the BenchFlow trajectory viewer, and submit it after the user reviews it. Use this skill whenever someone pastes a BenchFlow eval prize line, wants to submit / share / contribute / upload a trajectory, set up traj upload, view a session, or pick a session to send. Also…

335 2d ago A 98 tokens original Apache-2.0

benchflow

21

benchflow-ai/benchflow

Skill Claude CodeCodex

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. Use when asked to benchmark an AI coding agent, run a benchmark suite, create tasks, view trajectories, or compare agent performance.

335 2d ago A 46 tokens original Apache-2.0

aenvironment-deploy

22

inclusionAI/AEnvironment

Skill Claude CodeCodex

Deploy sandboxed environment instances and services using AEnvironment. Use when deploying agent instances, web services, or applications to AEnvironment sandbox infrastructure. Supports three workflows - (1) Build image locally and deploy, (2) Register existing image and deploy, (3) Deploy from registered…

314 1mo ago A 89 tokens original Apache-2.0

codspeed-optimize

23

CodSpeedHQ/codspeed

Skill Claude CodeCodex

Autonomously optimize code for performance using CodSpeed benchmarks, flamegraph analysis, and iterative improvement. Use this skill whenever the user wants to make code faster, reduce CPU usage, optimize memory, improve throughput, find performance bottlenecks, or asks to 'optimize', 'speed up', 'make faster'…

280 3d ago A 113 tokens original Apache-2.0

CodSpeedHQ/codspeed

Skill Claude CodeCodex

Set up performance benchmarks and CodSpeed harness for a project. Use this skill whenever the user wants to create benchmarks, add performance tests, set up CodSpeed, configure codspeed.yml, integrate a benchmarking framework (criterion, divan, pytest-benchmark, vitest bench, go test -bench, google benchmark), or when…

280 3d ago A 120 tokens original Apache-2.0