Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add opendatahub-io/agent-eval-harness --skill eval-analyzegit clone --depth 1 https://github.com/opendatahub-io/agent-eval-harnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/opendatahub-io/agent-eval-harness/eval-analyze)<a href="https://agentmods.dev/skills/opendatahub-io/agent-eval-harness/eval-analyze"><img src="https://agentmods.dev/badge/skills/opendatahub-io/agent-eval-harness/eval-analyze/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/opendatahub-io/agent-eval-harness/eval-analyze"><img src="https://agentmods.dev/badge/skills/opendatahub-io/agent-eval-harness/eval-analyze.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 2 findings, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium Agent Snooping · line 82 Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.Fix: Remove all code or instructions that list or read other skills' files or directories. Skills should operate independently; cross-skill access is a privilege escalation.
- medium Agent Snooping · line 120 Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.Fix: Remove all code or instructions that list or read other skills' files or directories. Skills should operate independently; cross-skill access is a privilege escalation.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00130 | $0.04362 |
| Opus 5 | $0.00065 | $0.02181 |
| Sonnet 5 | $0.00026 | $0.00872 |
| Haiku 4.5 | $0.00013 | $0.00436 |
Grade A, and why
eval-analyze scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 287 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You generate eval.yaml — the configuration that /eval-run needs. You either:
- Analyze a skill (default): Read the skill deeply (including sub-skills), explore test cases, generate config for testing the skill
- Custom analysis (
--prompt): Execute a custom analysis prompt that defines what to evaluate and how
The core principle: observe, don't assume. Every field name, file pattern, and directory path in the generated eval.yaml must come from reading actual files. If you can't point to a specific file or field you observed, don't put it in the config.
Step 0: Parse Arguments and Discover Layout
| Argument | Required | Default | Description |
|---|---|---|---|
--skill <name> |
no | auto-detect | Which skill to analyze |
--config <path> |
no | auto-discover | Output path for the config |
--prompt <path> |
no | none | Custom analysis prompt (for non-skill evals) |
--update |
no | false | Fill in missing sections only, preserve user edits |
--assess |
no | false | Assess all skills and recommend which ones need evals |
If $ARGUMENTS contains --assess: skip the rest of Step 0 and Step 1. Go directly to Batch Assessment (between Step 0 and Step 1). Do not run state.py init, do not discover configs, do not look for a single skill.
Otherwise, proceed with the normal flow:
mkdir -p tmp
python3 ${CLAUDE_SKILL_DIR}/scripts/agent_eval/state.py init tmp/analyze-config.yaml \
skill=<skill> prompt=<prompt> config=<config> update=<true/false>
Config Location Discovery
If --config provided, use that path. Otherwise run python3 ${CLAUDE_SKILL_DIR}/../../scripts/discover.py and decide:
- No configs: Create
eval.yamlat project root - One root config and
--skilltargets a different eval than the existing one: offer to reorganize intoeval/layout. If the user accepts, runpython3 ${CLAUDE_SKILL_DIR}/scripts/reorganize.py --eval-name <name>(add--project-root <path>if not the cwd). If declined, ask where to put the new config. - Nested/flat layout already exists: place the new config at
eval/<skill-name>/eval.yaml(nested) or alongside existing flat configs
What ships with it
12 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- prompts/analyze-skill.md 14 KB
- prompts/assess-skills.md 3.8 KB
- prompts/generate-eval-md.md 1.8 KB
- references/eval-yaml-template.md 41 KB
- references/judge-prompt-template.md 5.3 KB
- scripts/agent_eval 19 B
- scripts/assess_skills.py 7.5 KB runs code
- scripts/find_skills.py 9.3 KB runs code
- scripts/list_builtins.py 431 B runs code
- scripts/reorganize.py 1.4 KB runs code
- scripts/resolve_prompt.py 2.8 KB runs code
- scripts/validate_eval.py 41 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago Changed · -1 lines 43102de5627e
- 10d ago First seen · 288 lines · 130 tokens per session scan A bfc6a6d9ac86
eval-analyze is a skill published in the GitHub repository opendatahub-io/agent-eval-harness (40 stars, last pushed 7d ago), licensed Apache-2.0. It adds 130 tokens to every session and 4,362 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
insight-error-page
Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…
next-cache-components-optimizer
Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…
next-partial-prefetching-adoption
Turn on Partial Prefetching in a Next.js app and work through the insights it surfaces. Use when the user wants to enable or adopt Partial Prefetching, flip the partialPrefetching flag, opt routes in with export const prefetch = 'partial', audit Link prefetch={true} behavior, preserve existing prefetched UI with…