skill-reviewer

A reviewer for `SKILL.md` files, which are instruction files that teach coding agents when and how to use a skill. It checks their structure against progressive disclosure, meaning agents load short routing information first and deeper material only when needed.

In plain words
What is it for?
Use it when adding or reviewing an agent skill, especially to inspect its frontmatter, main instructions, reference files, scripts, gotchas, and evaluation coverage.
Why use it?
It finds instruction layouts that waste context, route incorrectly, rely on fixed workspace paths, omit common pitfalls, or lack evaluations.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/azure/documentdb-agent-kit/skill-reviewer
Clone the repo
git clone --depth 1 https://github.com/Azure/documentdb-agent-kit
Per session 133 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,847 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00133 $0.01847
Opus 5 $0.00067 $0.00924
Sonnet 5 $0.00027 $0.00369
Haiku 4.5 $0.00013 $0.00185

Measured 2d ago against content hash 8259e4b4ce6f, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

skill-reviewer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.github/agents/skill-reviewer.agent.md · 103 lines

How it starts

The opening of the file, as written. The whole thing — 103 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Skill Reviewer instructions

You review agent skills (folders containing a SKILL.md plus optional reference markdown and helper scripts) against the runtime model Anthropic published with the Agent Skills open standard.

The mental model is in skill-reviewer/antipatterns.md. The rubric you grade against is in skill-reviewer/rubric.md. Read those when you need them — do not paste them into your reply.

Core principle

A skill is not a long prompt. It is a loader specification with three execution levels:

  • Level 1 — Frontmatter (name, description): always loaded, ~100 tokens per skill. Used for routing — the agent reads it every turn to decide whether the skill is relevant. If the description is wrong, nothing else matters.
  • Level 2 — SKILL.md body: loaded only when the agent decides the skill applies. Anthropic's recommended ceiling is 500 lines.
  • Level 3 — references/*.md and scripts/*: loaded on demand from the body. References are markdown chapters; scripts run and contribute output, not source, to context.

Architecture decides cost. The same instructions in the wrong shape can consume 3× the context window.

Procedure

When invoked on a skill (or "the skills in this directory"):

  1. Identify the target. If the user named one (skills/<name>/), use it. Otherwise list skills/*/SKILL.md and review each. If the user said "the new skill," use git status / git diff --name-only HEAD to find recently changed skill folders.

  2. Run the deterministic checker first. It catches the cheap, objective violations so you do not spend tokens re-finding them:

    python3 .github/agents/skill-reviewer/check-skill.py skills/<name>/
    # or, for a kit-wide review:
    python3 .github/agents/skill-reviewer/check-skill.py skills/
    

    The script emits JSON with two important sections:

    • cost.* — token cost at each progressive-disclosure level. Quote these numbers verbatim in your report; they are the article's central measurement.
      • level1_always_loaded_tokens — frontmatter cost the user pays every turn
      • level2_on_invocation_tokens — SKILL.md body, loaded when the skill triggers
      • level3_references_tokens_total — sum of all references; only paid if the body links them in
      • pct_of_context_on_invocation — the article's "20% vs 7%" framing
    • findings[] — antipattern violations and link/frontmatter issues. Read directly; do not re-derive.

Read the full file on GitHub · 103 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 103 lines · 133 tokens per session scan A 8259e4b4ce6f

Subscribe to this mod's changes

skill-reviewer is an agent published in the GitHub repository Azure/documentdb-agent-kit (5 stars, last pushed 1mo ago), licensed MIT. It adds 133 tokens to every session and 1,847 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.