benchmarks skills

36 tagged benchmarks, measured the same way as everything else here.

Browse within: agentic-evaluation 27openclaw 27

InternLM/WildClawBench

Skill Claude CodeCodex

Fetches and summarizes recent arXiv and Hugging Face papers with Agentic Paper Digest. Use when the user wants a paper digest, a JSON feed of recent papers, or to run the arXiv/HF pipeline.

515 15d ago A 54 tokens original MIT

agenticmail

02

InternLM/WildClawBench

Skill Claude CodeCodex

πŸŽ€ AgenticMail β€” Full email, SMS, storage & multi-agent coordination for AI agents. 63 tools.

515 15d ago A 28 tokens original MIT

InternLM/WildClawBench

Skill Claude CodeCodex

Booking links fail for groups. SkipUp schedules meetings with 2-50 participants via email β€” one API call coordinates across timezones automatically. Also: check status, pause, resume, or cancel requests. Async only β€” does not instant-book, access calendars, or do free/busy lookups.

515 15d ago A 66 tokens original MIT

eli-labz/Cognitive-Core-Skills

Skill Claude CodeCodex

Analyzes user queries, extracts relevant predicates, and utilizes Knowledge Catalog Search to find and rank the most relevant data entries. Engages with the user throughout the process.

163 1mo ago A 39 tokens original MIT

kb-search

05

eli-labz/Cognitive-Core-Skills

Skill Claude CodeCodex

Allows listing, searching and extracting information from local knowledge base documents for information about tables.

163 1mo ago A 20 tokens original MIT

parallel

06

mvanhorn/clawdbot-skill-parallel

Skill Claude CodeCodex

High-accuracy web research platform with 7 APIs - Search, Extract, Task (Deep Research), Chat, FindAll, Monitor, and Task Groups. Fast mode, 8 processor tiers, MCP tool calling, authenticated browsing, SSE streaming, and OpenAI-compatible chat. OpenClaw skill.

21 5mo ago A 62 tokens

rctruta/sql-benchmarks-dagster

Skill Claude CodeCodex

Build and submit a scaling benchmark experiment. Use when the goal names a scale-varying investigation (e.g. how does X scale from N to M rows, is the growth linear, or at what size does Y break).

2 1mo ago A 52 tokens original Apache-2.0

rctruta/sql-benchmarks-dagster

Skill Claude CodeCodex

Read and analyze completed benchmark experiment results. Use when an experiment status is complete and you need to compare engines, speedups, scaling, replication stability, or raw timings.

2 1mo ago A 39 tokens original Apache-2.0

artificial-analysis

09

mrloldev/artificial-analysis-skill

Skill Claude CodeCodex

Pick the right LLM or media model for a task, backed by live benchmark data from artificialanalysis.ai. Use when the user asks "which model should I use for X", "what's the best/fastest/cheapest model", "compare model A vs B", "model leaderboard", or anything about model intelligence / speed / price / context /…

2 4mo ago A 154 tokens original MIT

haico-bench-skill

10

izhaha0226/haico-bench-skill

Skill Claude CodeCodex

Use when creative or marketing work risks premature convergence on the first plausible AI output, especially when benchmarks are present and must inform structure without causing imitation.

2 4mo ago A 36 tokens original MIT