benchmarking skills

104 tagged benchmarking, measured the same way as everything else here.

Browse within: algorithmic-efficiency 17Evaluation 16business-consulting 16claude-cowork 16claude-plugin 16strategy-consulting 163d 12blender 12HumanEval 7MMLU 5agent-evaluation 5

autorag-setup

01

Marker-Inc-Korea/AutoRAG

Skill Claude CodeCodex

Configure AutoRAG for first use or repair its single-agent model, approved document roots, retrieval indexes, datasource skills, and health checks without exposing credentials.

5.1k 4d ago A 36 tokens

autorag

02

Marker-Inc-Korea/AutoRAG

Skill Claude CodeCodex

Use an already configured AutoRAG librarian agent to search, summarize, compare, and answer questions from local document collections. Use autorag-setup for configuration or indexing changes.

5.1k 4d ago A 38 tokens

aristoteleo/PantheonOS

Skill Claude CodeCodex

End-to-end workflow for gene panel design in scRNA-seq and spatial transcriptomics, that should be STRICTLY followed: dataset understanding + smart downsampling + train/test splits, algorithmic selection (HVG/DE/RF/scGeneFit/SpaPROS), optimal sub-panel discovery (ARI vs size), biological completion with a stability…

482 2d ago A 110 tokens original BSD-2-Clause

gregs1104/pgbent

Skill Claude CodeCodexCursor

Tune Matplotlib labels on pgbent PG18 charts: scatter point placement, horizontal bar value labels (inside bar with contrast color when they fit), overlap avoidance, and PNG regeneration. Use when fixing crowded scatter or bar labels, pg18-osm-power-.png, pg18-osm-relation-power.png, pg18-osm-relation-scatter.png…

269 17d ago A 89 tokens

hegelion

05

Hmbown/Hegelion

Skill Claude CodeCodex

Dialectical reasoning and autocoding via Hegelion MCP tools.

171 5mo ago A 17 tokens original MIT

benchmark-radar

06

ktwu01/benchmark-radar

Skill Claude CodeCodex

Find, inspect, and check AI benchmark records with the Benchmark Radar CLI. Use when a request needs benchmark discovery, details, recent Radar evidence, or local data health; do not assume why the user needs the results.

123 2d ago A 48 tokens original MIT

kangarooking/X-growth-skills

Skill Claude CodeCodex

A research method for studying successful social-media accounts and their popular posts to find ideas for a distinct content approach.

60 1mo ago A 169 tokens original MIT

flameox

08

morluto/flameox

Skill Claude CodeCodex

Use Flameox to investigate runtime performance, memory, execution, scaling, GPU kernels, inference, and reliability with preserved evidence and explicit claim quality.

49 2d ago A 33 tokens original MIT

research-proof

09

tonyblu331/research-proof

Skill Claude CodeCodex

Turn vague research ideas, math-heavy claims, AI-lab style agent loops, benchmark claims, causal claims, prototype-readiness claims, design research, prompt-injection-sensitive evidence reviews, and medical-research style questions into falsifiable proof programs with fixed Claim/Verifier/Current…

45 3mo ago A 101 tokens original MIT

sota-scan

10

MerlijnW70/sota-scan

Skill Claude CodeCodex

Benchmark the current repo against the state of the art by scanning real online repos (GitHub and the wider web) in the same domain, then produce a cited capability matrix and a ranked, repo-grounded gap list. Use when the user asks "is our X top-tier", "what are we missing", "compare us to the best", "study online…

42 2mo ago A 89 tokens original MIT

patrick-toulme/harnessgym

Skill Claude CodeCodex

Optimize HarnessGym tensor-layout kernelplan.json tasks with a generated MCP server for plan validation, dev/final benchmarking, trace analysis, rollback-safe search, candidate application, history comparison, and experiment ranking. Use when a task asks to minimize bestcycles for benchmark.py or tune tensor…

40 2mo ago A 72 tokens original Apache-2.0

h100-triton-rmsnorm

12

patrick-toulme/harnessgym

Skill Claude CodeCodex

Use for the H100 Triton fused RMSNorm + SiLU gate optimization task. Provides the workflow and MCP tooling for objective runs, rollback-safe config sweeps, source diagnostics, benchmark history, and final held-out verification.

40 2mo ago A 53 tokens original Apache-2.0

workdir-smoke

13

pinecone-io/cultivar

Skill Claude CodeCodex

Placeholder skill used only to smoke-test the eval framework's workdir-capture mechanism. Not a real skill — delete once a real code-gen skill exists.

37 15d ago A 36 tokens original MIT

workdir-smoke

14

pinecone-io/cultivar

Skill Claude CodeCodex

Placeholder skill used only to smoke-test the eval framework's workdir-capture mechanism. Not a real skill — delete once a real code-gen skill exists.

37 15d ago A 36 tokens copy · 100% MIT

cultivar

15

pinecone-io/cultivar

Skill Claude CodeCodex

Drive the cultivar CLI to test whether an agent skill improves behavior — scaffold tasks, run with/without the skill across Claude/Copilot/Gemini (locally or on Modal), grade against a rubric, and read the results.

37 15d ago A 50 tokens original MIT

dotnet-performance

16

wieslawsoltes/Performance-Skill

Skill Claude CodeCodex

Evidence-first performance engineering workflow for coding agents combining portable .NET diagnostics with platform-native profiling, BenchmarkDotNet, application benchmarking, statistical validation, deployment, container, graphics, and GPU procedures.

34 1mo ago A 42 tokens original MIT

change-management

17

abinauv/business-consulting

Skill Claude CodeCodex

Comprehensive change management consulting toolkit — stakeholder mapping, ADKAR framework, Kotter's 8-Step Model, communication planning, resistance analysis, cultural assessment, readiness assessment, training needs analysis, impact assessment, and transition planning.

25 6mo ago A 49 tokens original MIT

customer-insights

18

abinauv/business-consulting

Skill Claude CodeCodex

Analyze customer behavior, map journeys, develop personas, and identify growth opportunities. Use this skill when the user mentions: customer insights, voice of customer, VoC, customer journey, journey mapping, JTBD, jobs to be done, persona, customer segmentation, churn analysis, retention, NPS, CSAT, customer…

25 6mo ago A 108 tokens original MIT

abinauv/business-consulting

Skill Claude CodeCodex

Assess digital maturity, build transformation roadmaps, evaluate AI/automation opportunities, rationalize technology stacks, and design data and cloud strategies. Use this skill when the user mentions: digital transformation, digital maturity, digital strategy, technology modernization, legacy modernization…

25 6mo ago A 115 tokens original MIT

null-epoch

20

Firespawn-Studios/tne-sdk

Skill Claude CodeCodex

Play The Null Epoch, a persistent AI agent MMO. Use when the user wants to connect an agent to Null Epoch, check game state, submit actions, play the game, or interact with the Null Epoch API. Handles authentication, state polling, action submission, and survival strategy for the Sundered Grid. Do NOT use for general…

13 3mo ago A 78 tokens original MIT

autoresearch-ml

21

proyecto26/autoresearch-ai-plugin

Skill Claude CodeCodex

Autonomous LLM training optimization with GPU support. Runs 5-minute training experiments, measures valbpb, keeps improvements or reverts — repeat forever. Use this skill when the user asks to "train a model autonomously", "optimize LLM training", "run ML experiments", "autoresearch with GPU", "optimize valbpb"…

12 1mo ago A 212 tokens original MIT

autoresearch

22

proyecto26/autoresearch-ai-plugin

Skill Claude CodeCodex

Autonomous experiment loop: edit code, commit, run benchmark, extract metrics, keep improvements or revert, repeat forever. Use this skill when the user asks to "run autoresearch", "start an experiment loop", "optimize a metric autonomously", "autonomous experiments", "benchmark loop", "keep/discard experiments"…

12 1mo ago A 209 tokens original MIT

fusion-engine

23

luckeyfaraday/fusion-engine

Skill Claude CodeCodex

Run a single prompt across multiple LLMs in parallel via OpenRouter, then synthesize their responses into one answer with a judge model. Use when the user wants multi-model fusion, ensemble LLM queries, parallel model comparison, or a second opinion across providers.

11 2mo ago A 56 tokens original MIT

swe-bench-lite

24

greynewell/mcpbr

Skill Claude CodeCodex

Quick-start command to run SWE-bench Lite evaluation with sensible defaults.

10 4mo ago A 20 tokens original MIT