Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/lxb12123/agent-plugin-kit/evalgit clone --depth 1 https://github.com/lxb12123/agent-plugin-kitWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00013 | $0.00461 |
| Opus 5 | $0.00006 | $0.00230 |
| Sonnet 5 | $0.00003 | $0.00092 |
| Haiku 4.5 | $0.00001 | $0.00046 |
Grade A, and why
eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
/eval — Evaluate a skill
Run a skill's bundled evals/ cases (e.g. for skills/review) to see whether it meets the bar.
Flow
- Determine the skill directory to evaluate,
skills/<name>(ask the user, or take one already present in the current project). - List the cases:
Outputsnode "${CLAUDE_PLUGIN_ROOT}/lib/cli.mjs" eval skills/<name>[{name, input, expect}, ...]. - For each case: set up the scenario as described by
input(create a temporary fixture / git repo if needed), run the skill, and capture its output text. - Collect all outputs into a single JSON
{ "<caseName>": "<output>" }and write it to a temp file, e.g.runs.json.- For cases with
expect.rubric: you (as the LLM judge) decide whether the output satisfies the rubric description, written as an object{ "<caseName>": { "output": "...", "rubric": true/false } }. The engine counts the deterministic assertions together with your rubric verdict toward pass/fail.
- For cases with
- Score:
Outputsnode "${CLAUDE_PLUGIN_ROOT}/lib/cli.mjs" eval skills/<name> --runs runs.json{total, passed, failed, cases:[{name, pass, failures}]}. - Relay the report to the user; for each case with a non-empty
failures, point out which expectation (contains / notContains / matches) was not met.
Notes
- The deterministic assertions in
expect(contains / notContains / matches regex) are scored by the engine, and pass/fail is decided by it. - When subjective quality (rubric) judgment is needed, you may add commentary, but it does not change the deterministic conclusion.
- Evaluation is read-only on the user's project; it does not write.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 33 lines · 13 tokens per session scan A c15c54b3279e
eval is a command published in the GitHub repository lxb12123/agent-plugin-kit (1 stars, last pushed 2mo ago), licensed MIT. It adds 13 tokens to every session and 461 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
template
Manage issue templates for streamlined issue creation.
design-review
Workflow recipe — review a design end-to-end, ending in measured numbers rather than adjectives, by chaining 4 skills.
setup-pm-skills
Onboard a new user — find out what they do, recommend the right bundles & top skills, and set up a project CONTEXT.md so every skill is tailored to them.
security-review
CWE 기반 보안 검토 + STRIDE 위협 모델링 (v6 - effort:max 강제).
auto-browse
Auto-browse — learn, optimize, and graduate browser operations or web data-mining workflows.
release-harn
Run the tag-first Harn release workflow.