Skill Claude CodeCodex
Configure AutoRAG for first use or repair its single-agent model, approved document roots, retrieval indexes, datasource skills, and health checks without exposing credentials.
104 tagged benchmarking, measured the same way as everything else here.
Browse within: algorithmic-efficiency 17Evaluation 16business-consulting 16claude-cowork 16claude-plugin 16strategy-consulting 163d 12blender 12HumanEval 7MMLU 5agent-evaluation 5
Skill Claude CodeCodex
Configure AutoRAG for first use or repair its single-agent model, approved document roots, retrieval indexes, datasource skills, and health checks without exposing credentials.
Skill Claude CodeCodex
Use an already configured AutoRAG librarian agent to search, summarize, compare, and answer questions from local document collections. Use autorag-setup for configuration or indexing changes.
Skill Claude CodeCodex
End-to-end workflow for gene panel design in scRNA-seq and spatial transcriptomics, that should be STRICTLY followed: dataset understanding + smart downsampling + train/test splits, algorithmic selection (HVG/DE/RF/scGeneFit/SpaPROS), optimal sub-panel discovery (ARI vs size), biological completion with a stability…
Skill Claude CodeCodexCursor
Tune Matplotlib labels on pgbent PG18 charts: scatter point placement, horizontal bar value labels (inside bar with contrast color when they fit), overlap avoidance, and PNG regeneration. Use when fixing crowded scatter or bar labels, pg18-osm-power-.png, pg18-osm-relation-power.png, pg18-osm-relation-scatter.png…
Skill Claude CodeCodex
Dialectical reasoning and autocoding via Hegelion MCP tools.
Skill Claude CodeCodex
Find, inspect, and check AI benchmark records with the Benchmark Radar CLI. Use when a request needs benchmark discovery, details, recent Radar evidence, or local data health; do not assume why the user needs the results.
Skill Claude CodeCodex
A research method for studying successful social-media accounts and their popular posts to find ideas for a distinct content approach.
Skill Claude CodeCodex
Use Flameox to investigate runtime performance, memory, execution, scaling, GPU kernels, inference, and reliability with preserved evidence and explicit claim quality.
Skill Claude CodeCodex
Turn vague research ideas, math-heavy claims, AI-lab style agent loops, benchmark claims, causal claims, prototype-readiness claims, design research, prompt-injection-sensitive evidence reviews, and medical-research style questions into falsifiable proof programs with fixed Claim/Verifier/Current…
Skill Claude CodeCodex
Benchmark the current repo against the state of the art by scanning real online repos (GitHub and the wider web) in the same domain, then produce a cited capability matrix and a ranked, repo-grounded gap list. Use when the user asks "is our X top-tier", "what are we missing", "compare us to the best", "study online…
Skill Claude CodeCodex
Optimize HarnessGym tensor-layout kernelplan.json tasks with a generated MCP server for plan validation, dev/final benchmarking, trace analysis, rollback-safe search, candidate application, history comparison, and experiment ranking. Use when a task asks to minimize bestcycles for benchmark.py or tune tensor…
Skill Claude CodeCodex
Use for the H100 Triton fused RMSNorm + SiLU gate optimization task. Provides the workflow and MCP tooling for objective runs, rollback-safe config sweeps, source diagnostics, benchmark history, and final held-out verification.
Skill Claude CodeCodex
Placeholder skill used only to smoke-test the eval framework's workdir-capture mechanism. Not a real skill — delete once a real code-gen skill exists.
Skill Claude CodeCodex
Placeholder skill used only to smoke-test the eval framework's workdir-capture mechanism. Not a real skill — delete once a real code-gen skill exists.
Skill Claude CodeCodex
Drive the cultivar CLI to test whether an agent skill improves behavior — scaffold tasks, run with/without the skill across Claude/Copilot/Gemini (locally or on Modal), grade against a rubric, and read the results.
wieslawsoltes/Performance-Skill
Skill Claude CodeCodex
Evidence-first performance engineering workflow for coding agents combining portable .NET diagnostics with platform-native profiling, BenchmarkDotNet, application benchmarking, statistical validation, deployment, container, graphics, and GPU procedures.
Skill Claude CodeCodex
Comprehensive change management consulting toolkit — stakeholder mapping, ADKAR framework, Kotter's 8-Step Model, communication planning, resistance analysis, cultural assessment, readiness assessment, training needs analysis, impact assessment, and transition planning.
Skill Claude CodeCodex
Analyze customer behavior, map journeys, develop personas, and identify growth opportunities. Use this skill when the user mentions: customer insights, voice of customer, VoC, customer journey, journey mapping, JTBD, jobs to be done, persona, customer segmentation, churn analysis, retention, NPS, CSAT, customer…
Skill Claude CodeCodex
Assess digital maturity, build transformation roadmaps, evaluate AI/automation opportunities, rationalize technology stacks, and design data and cloud strategies. Use this skill when the user mentions: digital transformation, digital maturity, digital strategy, technology modernization, legacy modernization…
Skill Claude CodeCodex
Play The Null Epoch, a persistent AI agent MMO. Use when the user wants to connect an agent to Null Epoch, check game state, submit actions, play the game, or interact with the Null Epoch API. Handles authentication, state polling, action submission, and survival strategy for the Sundered Grid. Do NOT use for general…
proyecto26/autoresearch-ai-plugin
Skill Claude CodeCodex
Autonomous LLM training optimization with GPU support. Runs 5-minute training experiments, measures valbpb, keeps improvements or reverts — repeat forever. Use this skill when the user asks to "train a model autonomously", "optimize LLM training", "run ML experiments", "autoresearch with GPU", "optimize valbpb"…
proyecto26/autoresearch-ai-plugin
Skill Claude CodeCodex
Autonomous experiment loop: edit code, commit, run benchmark, extract metrics, keep improvements or revert, repeat forever. Use this skill when the user asks to "run autoresearch", "start an experiment loop", "optimize a metric autonomously", "autonomous experiments", "benchmark loop", "keep/discard experiments"…
Skill Claude CodeCodex
Run a single prompt across multiple LLMs in parallel via OpenRouter, then synthesize their responses into one answer with a judge model. Use when the user wants multi-model fusion, ensemble LLM queries, parallel model comparison, or a second opinion across providers.
Skill Claude CodeCodex
Quick-start command to run SWE-bench Lite evaluation with sensible defaults.