Skill Claude CodeCodex
Fetches and summarizes recent arXiv and Hugging Face papers with Agentic Paper Digest. Use when the user wants a paper digest, a JSON feed of recent papers, or to run the arXiv/HF pipeline.
36 tagged benchmarks, measured the same way as everything else here.
Browse within: agentic-evaluation 27openclaw 27
Skill Claude CodeCodex
Fetches and summarizes recent arXiv and Hugging Face papers with Agentic Paper Digest. Use when the user wants a paper digest, a JSON feed of recent papers, or to run the arXiv/HF pipeline.
Skill Claude CodeCodex
π AgenticMail β Full email, SMS, storage & multi-agent coordination for AI agents. 63 tools.
Skill Claude CodeCodex
Booking links fail for groups. SkipUp schedules meetings with 2-50 participants via email β one API call coordinates across timezones automatically. Also: check status, pause, resume, or cancel requests. Async only β does not instant-book, access calendars, or do free/busy lookups.
eli-labz/Cognitive-Core-Skills
Skill Claude CodeCodex
Analyzes user queries, extracts relevant predicates, and utilizes Knowledge Catalog Search to find and rank the most relevant data entries. Engages with the user throughout the process.
eli-labz/Cognitive-Core-Skills
Skill Claude CodeCodex
Allows listing, searching and extracting information from local knowledge base documents for information about tables.
mvanhorn/clawdbot-skill-parallel
Skill Claude CodeCodex
High-accuracy web research platform with 7 APIs - Search, Extract, Task (Deep Research), Chat, FindAll, Monitor, and Task Groups. Fast mode, 8 processor tiers, MCP tool calling, authenticated browsing, SSE streaming, and OpenAI-compatible chat. OpenClaw skill.
rctruta/sql-benchmarks-dagster
Skill Claude CodeCodex
Build and submit a scaling benchmark experiment. Use when the goal names a scale-varying investigation (e.g. how does X scale from N to M rows, is the growth linear, or at what size does Y break).
rctruta/sql-benchmarks-dagster
Skill Claude CodeCodex
Read and analyze completed benchmark experiment results. Use when an experiment status is complete and you need to compare engines, speedups, scaling, replication stability, or raw timings.
mrloldev/artificial-analysis-skill
Skill Claude CodeCodex
Pick the right LLM or media model for a task, backed by live benchmark data from artificialanalysis.ai. Use when the user asks "which model should I use for X", "what's the best/fastest/cheapest model", "compare model A vs B", "model leaderboard", or anything about model intelligence / speed / price / context /β¦
Skill Claude CodeCodex
Use when creative or marketing work risks premature convergence on the first plausible AI output, especially when benchmarks are present and must inform structure without causing imitation.