Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add closedloop-ai/claude-plugins --skill eval-cachegit clone --depth 1 https://github.com/closedloop-ai/claude-pluginsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/closedloop-ai/claude-plugins/eval-cache)<a href="https://agentmods.dev/skills/closedloop-ai/claude-plugins/eval-cache"><img src="https://agentmods.dev/badge/skills/closedloop-ai/claude-plugins/eval-cache.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00082 | $0.00487 |
| Opus 5 | $0.00041 | $0.00244 |
| Sonnet 5 | $0.00016 | $0.00097 |
| Haiku 4.5 | $0.00008 | $0.00049 |
Grade A, and why
eval-cache scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Eval Cache
Check whether a prior simple-mode evaluation can be reused, avoiding a redundant plan-evaluator launch when the plan has not changed.
When to Use
Activate this skill at the start of Phase 1.3 (Simple Mode Evaluation), before launching @code:plan-evaluator. If the cache is fresh, skip the evaluator entirely and use the cached result.
Usage
Run the cache check script:
bash ${CLAUDE_SKILL_DIR}/scripts/check_eval_cache.sh <WORKDIR>
Interpreting Output
The script prints one of two structured results to stdout:
Cache Hit
EVAL_CACHE_HIT
simple_mode: true|false
selected_critics: [critic1, critic2, ...]
summary: <cached evaluation summary>
Action: Parse simple_mode and selected_critics from the output. Skip launching @code:plan-evaluator and proceed with the cached values as if the evaluator had just returned them.
Cache Miss
EVAL_CACHE_MISS
reason: <why the cache is stale or missing>
Action: Launch @code:plan-evaluator as normal. The evaluator will write a fresh plan-evaluation.json that subsequent iterations can cache from.
How Freshness Works
The script uses file modification timestamps:
- If
plan-evaluation.jsondoes not exist: miss - If
plan.jsonis newer thanplan-evaluation.json: miss (plan was modified since last evaluation) - If
plan-evaluation.jsonis newer thanplan.json: hit (evaluation is still valid)
This correctly handles the case where a user modifies the plan while the workflow is paused: editing plan.json updates its mtime, invalidating the cached evaluation.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today First seen · 60 lines · 82 tokens per session scan A 2f28602bce1c
eval-cache is a skill published in the GitHub repository closedloop-ai/claude-plugins (103 stars, last pushed yesterday), licensed Apache-2.0. It adds 82 tokens to every session and 487 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-07.
Other skills, from other repositories
pr-reviewer
Reviews a diff or security scope read-only using evidence-tiered findings, structural and context-error rubrics, and repository review policy. Use when asked to "review my changes", "structural review", "review for AI patterns", or "security audit". For applying fixes use tidy; for UI defects use ui-design.
autoship
Runs a changesets npm release through the version PR, CI publish, and registry verification. Use when asked to "release this package", "autoship", "merge Version Packages", or diagnose a release that did not publish. For feature PRs use pr-creator or pr-babysitter.
scaffold-cli
Scaffolds a TypeScript CLI and npm package with the house toolchain, dual tsdown outputs, CLI contracts, changesets, and publishing templates. Use when asked to "scaffold a CLI" or "start an npm package". For an existing package release use autoship; for existing API ergonomics use dx-audit.
claw-mux
Control cmux terminal topology and I/O — send commands to panes, read output, split layouts, monitor logs, orchestrate multi-pane workflows. Requires cmux environment.
report-manager
Manage and refine vision-powers reports: list, open, delete, search, and refine sections. Use when asked to list, open, delete, search, or update generated HTML reports.
fetch-sitemap
Extract URLs from an XML sitemap with optional regex filtering.