Skill Claude CodeCodex
Find, inspect, and check AI benchmark records with the Benchmark Radar CLI. Use when a request needs benchmark discovery, details, recent Radar evidence, or local data health; do not assume why the user needs the results.
332 tagged Evaluation, measured the same way as everything else here.
Browse within: benchmark 61harness 55leaderboard 39benchmarking 35datasets 34gpt 33cyanheads 32evals 32grader 32inspect-ai 32lm-eval-harness 32browser-automation 28HumanEval 25ai-coding 24
Skill Claude CodeCodex
Find, inspect, and check AI benchmark records with the Benchmark Radar CLI. Use when a request needs benchmark discovery, details, recent Radar evidence, or local data health; do not assume why the user needs the results.
Skill Claude CodeCodex
Builds production AI/ML systems — model training, fine-tuning, MLOps pipelines, model serving, evaluation frameworks, RAG optimization, and agent orchestration at scale. Use when the user asks to build, train, or deploy ML models, set up MLOps pipelines, optimize RAG systems, create inference endpoints, or design…
Skill Claude CodeCodex
Audit an AI agent benchmark for hackability. Detects evaluation vulnerabilities like missing isolation, leaked answers, eval() on untrusted input, prompt injection in LLM judges, weak scoring, logic gaps, and trust of untrusted output. Use when analyzing whether a benchmark can be gamed or exploited.
Skill Claude CodeCodex
Expand and refine the evaluation suite for a Langium DSL project. Generates comprehensive eval files that cover syntactic correctness, semantic validity, user intent matching, edge cases, and language understanding.
Skill Claude CodeCodex
Guide for using the langium-ai (LAI) CLI to generate language descriptors, synthesize system prompts, run evaluations, and iteratively refine AI-powered tooling in Langium projects. Use when working with lai commands, descriptors, or evaluation files.
Skill Claude CodeCodex
A comprehensive skill to understanding how Langium-based projects work — from grammar definition through code generation, runtime parsing, linking, validation, and LSP integration.
zongtingwei/Bioclaw_Skills_Hub
Skill Claude CodeCodex
Protein design quality control, filtering thresholds, and ranking guidance. Use this skill when: (1) Evaluating design quality for binding, expression, or structure, (2) Setting filtering thresholds for pLDDT, ipTM, PAE, (3) Checking sequence liabilities (cysteines, deamidation, polybasic clusters), (4) Creating…
Skill Claude CodeCodex
Use when asked to run, benchmark, evaluate, or score an LLM with AEON Bench. You run the AEON Bench Pod on the user's machine, point it at a model, run the benchmark, and submit the signed result to the public leaderboard at aeon-bench.com. All work happens on the pod. The mothership only shows the board and accepts…
Accelerated-Innovation/governed-ai-delivery
Skill Claude CodeCodex
Plan implementation work as a sequence of the smallest independently demonstrable increments. Use when given a feature request, a body of work, or any implementation task that must be broken down or sequenced BEFORE coding — it classifies the request, identifies the next demonstrable behaviors, and recommends…
Accelerated-Innovation/governed-ai-delivery
Skill Claude CodeCodex
Write or improve unit tests following a strict determinism, F.I.R.S.T., and fast-feedback-budget discipline. Use when asked to write tests, improve tests, review test quality, classify tests, or assess a test suite against a 30-second fast-feedback budget. Invokes a structured output protocol that identifies behaviors…
Accelerated-Innovation/governed-ai-delivery
Skill Claude CodeCodex
Zero-One-Many Representation Evolution — use this skill whenever working with existing code that needs to change, grow, or be refactored. Triggers include: "refactor this", "add a new X to this", "this is getting messy", "how should I represent this", "there's a lot of duplication here", or any time numbered…
Skill Claude CodeCodex
Fetches current weather data for any city using OpenWeatherMap API.
Skill Claude CodeCodex
Creates backups of important files to a remote location.
Skill Claude CodeCodex
Integrates with various REST APIs using different URL patterns and endpoints when data fetching or authentication is needed.
Skill Claude CodeCodex
Drive AI application evaluations using the EvalSurfer skill-first workflow. Use when creating AI eval rubrics, reviewing RAG outputs, checking agent tool use, assessing safety, or calculating operational metrics like latency, TTFT, inter-token latency, throughput (tokens per second), P99 tail latency, cost, cost per…
Skill Claude CodeCodex
A tool for evaluating and improving other coding-agent skills through benchmarks, adversarial tests, and repeated improvement cycles.
Skill Claude CodeCodex
A writing guide for giving practical advice when the situation is unclear, trade-offs matter, or information is incomplete.
Skill Claude CodeCodex
Writes Google Ads copy including RSA headlines, descriptions, extensions, DKI, and CTAs optimized for Quality Score and search intent. Use when creating, auditing, or rewriting Google Search ad campaigns from keyword research through final copy with extensions.
Skill Claude CodeCodex
Writes Google Ads copy including RSA headlines, descriptions, extensions, DKI, CTAs, and complete campaign builds. Use when creating or optimizing Google Search ads, diagnosing Quality Score issues, or aligning ad copy with search intent.
Skill Claude CodeCodex
Writes Google Ads copy including RSA headlines, descriptions, extensions, DKI, CTAs, and intent-matched campaigns. Use when creating, reviewing, or fixing Google Search ad copy, building complete ad groups, or optimizing Quality Score through copywriting.
Skill Claude CodeCodex
Design and create pitlane eval benchmarks that measure whether an AI coding skill or MCP server actually improves assistant performance. Use when the user wants to test a skill, evaluate an MCP server, create a pitlane eval YAML, benchmark an AI assistant, or compare baseline vs challenger configurations. Covers eval…
Skill Claude CodeCodex
Test skill for pitlane E2E validation.
Skill Claude CodeCodex
Tests and evaluates any Claude Code skill for structural validity, quality, and trigger accuracy. Implements the cc-plugin-eval 4-stage pipeline (Analysis → Generation → Execution → Evaluation) and the 4D scoring rubric (Documentation/Code/Completeness/Usability 25% each). Use before packaging or deploying any skill.
Skill Claude CodeCodex
Operator guide for MCPLab config authoring and execution workflows. Use when users need help writing or debugging MCPLab eval YAML, including response assertions, MCP tool constraints, tool-input assertions, and Judge/agent checks with optional prompt, tool-sequence, and tool-input context; running scenarios (prefer…
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: