eval-skills

eval-skills is a skill for Claude Code, Codex from FlorianBruniaux/claude-code-plugins. It costs 67 tokens per session (2,909 once invoked), scanned A, original, MIT.

A skill that audits project skills and reports on their instructions, metadata, tool permissions, effort levels, and content quality.

In plain words
What is it for?
Reviewing skill collections, checking frontmatter, recommending effort levels, and producing a quality report.
Why use it?
It helps find incomplete or poorly scoped skills before they cause inconsistent agent behavior or are shipped to a team.

Skill for Claude CodeCodex

Part of the code-quality plugin — 7 skills, 8 commands, 6 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/florianbruniaux/claude-code-plugins/eval-skills
Any agent
npx skills add FlorianBruniaux/claude-code-plugins --skill eval-skills
Clone the repo
git clone --depth 1 https://github.com/FlorianBruniaux/claude-code-plugins

Made for: Claude Code, Codex.

Or install code-quality, the plugin that ships this one along with the rest of its 7 skills, 8 commands, 6 agents.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for eval-skills

README.md
[![agentmods](https://agentmods.dev/badge/skills/florianbruniaux/claude-code-plugins/eval-skills.svg)](https://agentmods.dev/skills/florianbruniaux/claude-code-plugins/eval-skills)
Your own site
<a href="https://agentmods.dev/skills/florianbruniaux/claude-code-plugins/eval-skills"><img src="https://agentmods.dev/badge/skills/florianbruniaux/claude-code-plugins/eval-skills.svg" alt="Measured on agentmods" height="20"></a>
Per session 67 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,909 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00067 $0.02909
Opus 5 $0.00034 $0.01455
Sonnet 5 $0.00013 $0.00582
Haiku 4.5 $0.00007 $0.00291

Measured 4d ago against content hash dbec825a7063, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

eval-skills scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Enumerates other installed skillslowAgent snooping

Other skills' SKILL.md files reveal prompts, capabilities and secrets that should be invisible to peers.

find .claude/skills -name "SKILL.md" 2>/dev/null

Downgraded: this mod is about security review, or the phrase is quoted, so it is likely naming the pattern rather than instructing it.

plugins/code-quality/skills/eval-skills/SKILL.md · 269 lines

How it starts

The opening of the file, as written. The whole thing — 269 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Skill Evaluator

Discover all skills in the project, score them across 6 criteria, and infer the appropriate effort level based on content analysis.

When to Use

  • New project: run once to establish baseline quality
  • Before committing a skill to a team repo
  • After bulk-importing skills from another project
  • When adding effort fields for the first time
  • When a skill doesn't auto-trigger and you want to diagnose why

What Gets Audited

All SKILL.md files and flat .md files found in:

  • .claude/skills/**
  • ~/.claude/skills/** (if requested)
  • .claude/commands/** (legacy flat files, still valid)
  • Any path passed as argument: /eval-skills ./my-skills-dir

Valid Frontmatter Fields

Claude Code skills follow the agentskills.io open standard, extended with Claude Code-specific fields. Flag any field not in this table as unsupported.

agentskills.io spec fields

Field Required Notes
name No Display label shown in skill lists. The command name always comes from the directory name, not this field.
description Recommended Combined with when_to_use, truncated at 1,536 chars in context. First paragraph used if omitted.
when_to_use No Additional trigger phrases and example requests. Appended to description in context; counts toward the 1,536-char cap.
allowed-tools No Tools usable without per-use approval while the skill is active. Space-separated string or YAML list (both valid).
license No License identifier (agentskills.io spec)
compatibility No Compatibility constraints (agentskills.io spec)
metadata No Arbitrary metadata object (agentskills.io spec)

Claude Code extension fields

Field Required Notes
argument-hint No Hint shown during autocomplete. Example: [issue-number] or [filename] [format]
arguments No Named positional args for $name substitution. Space-separated string or YAML list. Names map to positions in order.
disable-model-invocation No true = user-only invocation. Removes skill from Claude's context and prevents preloading in subagents.
user-invocable No false = hides skill from / menu but Claude can still auto-invoke it.
disallowed-tools No Tools blocked while this skill is active. Cleared after the current message.
model No Override model for this skill's turn only. Reverts to session model on next prompt.
effort No Thinking effort: low, medium, high, xhigh, max. Overrides session effort for the turn.
context No fork = runs skill in an isolated subagent context. The skill body becomes the subagent's prompt.
agent No Which subagent type to use when context: fork is set. Options: Explore, Plan, general-purpose, or any custom agent in .claude/agents/.
hooks No Skill-scoped lifecycle hooks. Same format as session hooks.
paths No Glob patterns that limit when Claude auto-loads this skill. Same format as path-specific rules.
shell No Shell for !backtick injection: bash (default) or powershell.

Read the full file on GitHub · 269 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 269 lines · 67 tokens per session scan A dbec825a7063

Subscribe to this mod's changes

eval-skills is a skill published in the GitHub repository FlorianBruniaux/claude-code-plugins (40 stars, last pushed 2d ago), licensed MIT. It adds 67 tokens to every session and 2,909 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 1 finding (enumerates other installed skills). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

vc-autopilot

Emit and validate the provisional goal block for Autopilot Mode. Owns the 9-field format and resume detection from a pasted goal block.

withkynam/vibecode-pro-max-kit · 34 tokens

vc-problem-solving

Apply systematic problem-solving techniques when stuck. Use for complexity spirals, innovation blocks, recurring patterns, assumption constraints, simplification cascades, scale uncertainty.

withkynam/vibecode-pro-max-kit · 36 tokens

report

Read the delivery log and say which rules actually fire, which never have, and what to prune or fix. Use when the user asks whether ballast is doing anything, wants to clean up their rule catalog, or on a periodic review.

svy04/ballast · 49 tokens

squid-self-improve

Analyze developer corrections from the current coding session and persist lessons learned as rules in AGENTS.md files or memory. Use at the end of a session after the developer corrected your work, when they say "squid-self-improve", ask to capture what was learned, or ask you to reflect on mistakes and extract…

iusztinpaul/squid · 73 tokens

prompt-template-guide

查看、创建、修改、删除或排查 DeterminFlow Prompt Template 与 system prompt section 时必须加载此技能;也适用于 section 顺序、workflowonly/chatonly、cache break、自定义 templatevariables、系统变量渲染以及 Agent Definition 的 prompttemplate 绑定。.

alikon-art/DeterminFlow · 62 tokens

script-library-guide

创建、更新、删除、引用或排查 DeterminFlow Script Library 脚本时必须加载此技能;也适用于 Workflow Script 节点、inline 与 library 选择、SCRIPT.md、scriptargv、共享 workspace、WFVAR/scriptout 输出协议、Plugin 脚本 owner 冲突与 Task 身份冻结。.

alikon-art/DeterminFlow · 76 tokens