skill-benchmark

skill-benchmark is a skill for Claude Code, Codex from OutlineDriven/odin-claude-plugin. It costs 64 tokens per session (1,744 once invoked), scanned A, original, Apache-2.0.

A benchmarking tool for scoring agent skills or comparing AI models on the same task set.

In plain words
What is it for?
Running judged evaluations, comparing candidate models, tracking inference spend, and writing benchmark reports.
Why use it?
It shows quality scores, regressions, trends, and model costs so results can be compared against a baseline.

Skill for Claude CodeCodex

Written for Claude Code and Codex: disable-model-invocation in frontmatter, but also agents/openai.yaml present. Also seen: built for gstack.

Part of the odin-agent plugin — 46 skills shipped together

Good fit Running judged evaluations, comparing candidate models, tracking inference spend, and writing benchmark…

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/outlinedriven/odin-claude-plugin/skill-benchmark
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add OutlineDriven/odin-claude-plugin --skill skill-benchmark
Clone the repo
git clone --depth 1 https://github.com/OutlineDriven/odin-claude-plugin

Made for: Claude Code, Codex.

Or install odin-agent, the plugin that ships this one along with the rest of its 46 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for skill-benchmark

README.md
[![agentmods](https://agentmods.dev/badge/skills/outlinedriven/odin-claude-plugin/skill-benchmark.svg)](https://agentmods.dev/skills/outlinedriven/odin-claude-plugin/skill-benchmark)
Your own site
<a href="https://agentmods.dev/skills/outlinedriven/odin-claude-plugin/skill-benchmark"><img src="https://agentmods.dev/badge/skills/outlinedriven/odin-claude-plugin/skill-benchmark.svg" alt="Measured on agentmods" height="20"></a>
Per session 64 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,744 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00064 $0.01744
Opus 5 $0.00032 $0.00872
Sonnet 5 $0.00013 $0.00349
Haiku 4.5 $0.00006 $0.00174

Measured 2d ago against content hash c7b3b6dd0fe6, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

skill-benchmark scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/odin-agent/skills/skill-benchmark/SKILL.md · 68 lines

How it starts

The opening of the file, as written. The whole thing — 68 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Skill benchmark

Contract

Field Bound contract
Trigger The user runs /skill-benchmark
Authority Human-only. Preview the benchmark target, judge or candidate models, rubric or task set, and estimated spend before any LLM call. No skill, code, credential, or remote mutation.
Side effect Writes benchmark artifacts under .gstack/benchmark-reports/ and incurs LLM inference spend.
Done A scored skill-quality report or model-comparison table is written and returned to the human.

Inputs

Skill-quality mode

  • --baseline: capture a scored baseline before changes. Run first on a clean branch.
  • --quick: single-pass scoring without baseline comparison.
  • --skills <name1>,<name2>: score only named skills. Omit to auto-discover from the skill directory.
  • --diff: score only skills whose files changed on the current branch.
  • --trend: show score trends from historical baseline files.
  • Judge model and rubric must be supplied or confirmed by the user before scoring begins.

Model-comparison mode

  • The task or task set to run against every candidate model (required).
  • The candidate model list (required): two or more models to compare.
  • Per-model run count or spend budget cap (optional; defaults to one run per model per task).
  • Output path for the comparison table (optional; defaults to a local artifact under .gstack/benchmark-reports/).

Procedure

  1. Determine the benchmark target from the request. If the user names skills to score or asks for skill-quality scoring, select skill-quality mode. If the user names candidate models and a task set, select model-comparison mode. Done when: the mode is selected.
  2. Create .gstack/benchmark-reports/ and .gstack/benchmark-reports/baselines/. Done when: both directories exist.
  3. Preview the benchmark plan to the user. In skill-quality mode: the skill list, judge model, rubric criteria, and estimated spend. In model-comparison mode: the candidate models, task set, per-model run count, and estimated spend. Stop and wait for confirmation before any LLM call. Done when: the user confirms the preview.
  4. Resolve and lock the benchmark scope. In skill-quality mode: if --skills is supplied, use those names; if --diff, run git diff <base>...HEAD --name-only and select skills whose files changed; otherwise auto-discover all skills in the skill directory. In model-comparison mode: fix the task set and model list; no new tasks or models may be added after this step. Done when: the skill set is resolved and non-empty, or the task set and model list are locked.
  5. Run the benchmark. In skill-quality mode: for each skill, send the skill body and the following rubric to the LLM judge; collect a 0-10 score per criterion and an overall score (mean of criteria).
    • Trigger clarity: does the trigger predicate unambiguously route the skill?
    • Procedure executability: can the procedure be followed step-by-step without ambiguity?
    • Failure recovery: are failure classes named with recovery or stop rules?
    • Output concreteness: does the output section name a concrete artifact? In model-comparison mode: for each task and each candidate model, run the task the fixed number of times; record each result with the model, task, run index, and observed cost. Done when: every skill has per-criterion and overall scores, or every model/task/run combination has a recorded result, cost, or failure marker.
  6. Score the results. In skill-quality mode: if --baseline, write per-skill per-criterion scores, timestamp, and branch to .gstack/benchmark-reports/baselines/baseline.json, report absolute scores, and stop. If a baseline exists and --baseline was not passed, compare each current score against the baseline: score drop greater than 50% of the baseline value or more than 2 points absolute is REGRESSION; score drop greater than 20% is WARNING; otherwise OK. In model-comparison mode: score or rank each result against the shared task's success criterion; use the criterion stated with the task, or ask the human for one if none is stated. Done when: every skill has a regression status (or absolute scores reported for --baseline), or each completed result has a score with no score invented without human approval.
  7. Aggregate and rank. In skill-quality mode: check each skill against the quality budget (overall score 7 or above passes, below 7 fails), compute the overall grade from the fraction of skills passing, rank skills by lowest current score, and for each failing skill name the weakest criterion and quote the judge rationale. In model-comparison mode: aggregate per-model scores across the task set into a comparison table with one row per model showing aggregate score, per-task breakdown, total observed spend, and run count. Done when: the overall grade is computed and failing skills are ranked with weakest-criterion rationale, or the table contains every model and accurately sums spend and run counts.
  8. If --trend in skill-quality mode: load historical baseline files, tabulate overall scores over time, and state whether quality is improving, stable, or degrading. Done when: the trend table is produced or --trend was not passed.
  9. Write and return the report. In skill-quality mode: write to .gstack/benchmark-reports/<date>-benchmark.md and .gstack/benchmark-reports/<date>-benchmark.json. In model-comparison mode: write the comparison table to the chosen output path (default .gstack/benchmark-reports/<date>-model-comparison.md). Present the completed report or table to the human with its saved path. Done when: the files are written and their completed contents have been returned.

Read the full file on GitHub · 68 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 68 lines · 64 tokens per session scan A c7b3b6dd0fe6

Subscribe to this mod's changes

skill-benchmark is a skill published in the GitHub repository OutlineDriven/odin-claude-plugin (35 stars, last pushed today), licensed Apache-2.0. It adds 64 tokens to every session and 1,744 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-04.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

local-ai-agents

Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…

microsoft/ai-agents-for-beginners · 200 tokens

next-cache-components-adoption

Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…

vercel/next.js · 95 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

insight-error-page

Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…

vercel/next.js · 83 tokens