Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add OutlineDriven/odin-claude-plugin --skill skill-benchmarkgit clone --depth 1 https://github.com/OutlineDriven/odin-claude-pluginWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/outlinedriven/odin-claude-plugin/skill-benchmark)<a href="https://agentmods.dev/skills/outlinedriven/odin-claude-plugin/skill-benchmark"><img src="https://agentmods.dev/badge/skills/outlinedriven/odin-claude-plugin/skill-benchmark.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00064 | $0.01744 |
| Opus 5 | $0.00032 | $0.00872 |
| Sonnet 5 | $0.00013 | $0.00349 |
| Haiku 4.5 | $0.00006 | $0.00174 |
Grade A, and why
skill-benchmark scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 68 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Skill benchmark
Contract
| Field | Bound contract |
|---|---|
| Trigger | The user runs /skill-benchmark |
| Authority | Human-only. Preview the benchmark target, judge or candidate models, rubric or task set, and estimated spend before any LLM call. No skill, code, credential, or remote mutation. |
| Side effect | Writes benchmark artifacts under .gstack/benchmark-reports/ and incurs LLM inference spend. |
| Done | A scored skill-quality report or model-comparison table is written and returned to the human. |
Inputs
Skill-quality mode
--baseline: capture a scored baseline before changes. Run first on a clean branch.--quick: single-pass scoring without baseline comparison.--skills <name1>,<name2>: score only named skills. Omit to auto-discover from the skill directory.--diff: score only skills whose files changed on the current branch.--trend: show score trends from historical baseline files.- Judge model and rubric must be supplied or confirmed by the user before scoring begins.
Model-comparison mode
- The task or task set to run against every candidate model (required).
- The candidate model list (required): two or more models to compare.
- Per-model run count or spend budget cap (optional; defaults to one run per model per task).
- Output path for the comparison table (optional; defaults to a local artifact under .gstack/benchmark-reports/).
Procedure
- Determine the benchmark target from the request. If the user names skills to score or asks for skill-quality scoring, select skill-quality mode. If the user names candidate models and a task set, select model-comparison mode. Done when: the mode is selected.
- Create
.gstack/benchmark-reports/and.gstack/benchmark-reports/baselines/. Done when: both directories exist. - Preview the benchmark plan to the user. In skill-quality mode: the skill list, judge model, rubric criteria, and estimated spend. In model-comparison mode: the candidate models, task set, per-model run count, and estimated spend. Stop and wait for confirmation before any LLM call. Done when: the user confirms the preview.
- Resolve and lock the benchmark scope. In skill-quality mode: if
--skillsis supplied, use those names; if--diff, rungit diff <base>...HEAD --name-onlyand select skills whose files changed; otherwise auto-discover all skills in the skill directory. In model-comparison mode: fix the task set and model list; no new tasks or models may be added after this step. Done when: the skill set is resolved and non-empty, or the task set and model list are locked. - Run the benchmark. In skill-quality mode: for each skill, send the skill body and the following rubric to the LLM judge; collect a 0-10 score per criterion and an overall score (mean of criteria).
- Trigger clarity: does the trigger predicate unambiguously route the skill?
- Procedure executability: can the procedure be followed step-by-step without ambiguity?
- Failure recovery: are failure classes named with recovery or stop rules?
- Output concreteness: does the output section name a concrete artifact? In model-comparison mode: for each task and each candidate model, run the task the fixed number of times; record each result with the model, task, run index, and observed cost. Done when: every skill has per-criterion and overall scores, or every model/task/run combination has a recorded result, cost, or failure marker.
- Score the results. In skill-quality mode: if
--baseline, write per-skill per-criterion scores, timestamp, and branch to.gstack/benchmark-reports/baselines/baseline.json, report absolute scores, and stop. If a baseline exists and--baselinewas not passed, compare each current score against the baseline: score drop greater than 50% of the baseline value or more than 2 points absolute is REGRESSION; score drop greater than 20% is WARNING; otherwise OK. In model-comparison mode: score or rank each result against the shared task's success criterion; use the criterion stated with the task, or ask the human for one if none is stated. Done when: every skill has a regression status (or absolute scores reported for--baseline), or each completed result has a score with no score invented without human approval. - Aggregate and rank. In skill-quality mode: check each skill against the quality budget (overall score 7 or above passes, below 7 fails), compute the overall grade from the fraction of skills passing, rank skills by lowest current score, and for each failing skill name the weakest criterion and quote the judge rationale. In model-comparison mode: aggregate per-model scores across the task set into a comparison table with one row per model showing aggregate score, per-task breakdown, total observed spend, and run count. Done when: the overall grade is computed and failing skills are ranked with weakest-criterion rationale, or the table contains every model and accurately sums spend and run counts.
- If
--trendin skill-quality mode: load historical baseline files, tabulate overall scores over time, and state whether quality is improving, stable, or degrading. Done when: the trend table is produced or--trendwas not passed. - Write and return the report. In skill-quality mode: write to
.gstack/benchmark-reports/<date>-benchmark.mdand.gstack/benchmark-reports/<date>-benchmark.json. In model-comparison mode: write the comparison table to the chosen output path (default.gstack/benchmark-reports/<date>-model-comparison.md). Present the completed report or table to the human with its saved path. Done when: the files are written and their completed contents have been returned.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 68 lines · 64 tokens per session scan A c7b3b6dd0fe6
skill-benchmark is a skill published in the GitHub repository OutlineDriven/odin-claude-plugin (35 stars, last pushed today), licensed Apache-2.0. It adds 64 tokens to every session and 1,744 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-04.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
insight-error-page
Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…