Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/typefox/agent-skills/skill-evalsnpx skills add TypeFox/agent-skills --skill skill-evalsgit clone --depth 1 https://github.com/TypeFox/agent-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/typefox/agent-skills/skill-evals)<a href="https://agentmods.dev/skills/typefox/agent-skills/skill-evals"><img src="https://agentmods.dev/badge/skills/typefox/agent-skills/skill-evals.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00110 | $0.08813 |
| Opus 5 | $0.00055 | $0.04407 |
| Sonnet 5 | $0.00022 | $0.01763 |
| Haiku 4.5 | $0.00011 | $0.00881 |
Grade B, and why
skill-evals scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Reads agent configuration directoriesmediumAgent snooping
.claude/, .codex/, .gemini/ hold keys, settings and other credentials a mod has no legitimate need for.
- **Out-of-process headless sessions** (Claude Code: `claude -p`; Codex: `codex exec`): each run is its own fresh process, which assembles project memory from disk at its own start — so with the root files off disk nothi How it starts
The opening of the file, as written. The whole thing — 201 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evaluating a skill with the eval loop
A skill is a bet that some instructions make an agent's output better. The loop is how you collect on the bet: run real prompts with the skill and against a baseline without it (or against the previous version), grade both against the same bar, and read the delta — what the skill costs in time and tokens versus what it buys in pass rate. Repeat until the delta stops moving.
This skill is the spine of that loop — the order of operations, the two gates that guard it, and the two points where only the user can supply what you need. The mechanics — exact file schemas, the grader and analyzer subagents, the aggregation script — live in skill-creator, which this skill drives rather than restates. When you need a schema or a command, follow the pointer into skill-creator.
One deliberate deviation from skill-creator. skill-creator surfaces results through an HTML eval viewer plus a generated benchmark.md and review.md. skill-evals does not: it runs skill-creator's scripts only to process and aggregate the JSON (grading.json, benchmark.json, timing.json), then consolidates everything human-facing into a single REPORT.md (Activity 7). Do not launch the eval viewer or rely on the scripts' Markdown/HTML output — REPORT.md is the one artifact the user reviews.
Preflight — clear both gates before any activity
Gate 1: skill-creator must be available
Locate the installed skill-creator skill and read its SKILL.md. Note its directory as SKILL_CREATOR; every mechanic below is addressed relative to it — SKILL_CREATOR/references/schemas.md, SKILL_CREATOR/scripts/, SKILL_CREATOR/agents/.
If skill-creator is not installed, stop here. skill-evals is an orchestration layer and cannot run without it. Tell the user to install it (it lives in the anthropics/skills repository) and re-run. Done when SKILL_CREATOR resolves to a readable directory, or you have stopped with that instruction.
What ships with it
10 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- evals/evals.json 16 KB
- evals/files/dsl-samples/basicMath.calc 127 B
- evals/files/dsl-samples/blog.dmodel 271 B
- evals/files/dsl-samples/datatypes.dmodel 123 B
- evals/files/dsl-samples/trafficlight.statemachine 362 B
- evals/files/dsl-skills/arithmetics-dsl/SKILL.md 1.3 KB
- evals/files/dsl-skills/domainmodel-dsl-v1/SKILL.md 890 B
- evals/files/dsl-skills/domainmodel-dsl-v2/SKILL.md 3.5 KB
- evals/files/dsl-skills/statemachine-dsl/SKILL.md 3.8 KB
- references/claude-code-workflow-tool.md 9.4 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 201 lines · 110 tokens per session scan B de10a8900ddf
skill-evals is a skill published in the GitHub repository TypeFox/agent-skills (5 stars, last pushed 5d ago), licensed MIT. It adds 110 tokens to every session and 8,813 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it B with 1 finding (reads agent configuration directories). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…