Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/zalom/plastic/skill-evaluatingnpx skills add zalom/plastic --skill skill-evaluatinggit clone --depth 1 https://github.com/zalom/plasticWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00099 | $0.01323 |
| Opus 5 | $0.00049 | $0.00661 |
| Sonnet 5 | $0.00020 | $0.00265 |
| Haiku 4.5 | $0.00010 | $0.00132 |
Grade A, and why
plastic-skill-evaluating scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 142 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evaluating Skills
Eval methodology for Plastic skills, based on agentskills.io and Anthropic's eval guide.
Gotchas
- Assertions written before observing output are almost always wrong — run the eval first, observe actual output, THEN write assertions
- Near-miss negative test cases are the most valuable — prompts that share keywords with should-trigger cases but need a different skill entirely
- Select the best skill iteration by validation pass rate, not the last one
- Grade outcomes, not execution paths — if the agent solved the task via an unexpected route but produced correct output, that is a pass
- Same skill can behave differently across agent frameworks — test on each target agent (Claude Code, Hermes, OpenClaw, Codex)
Procedure
Step 1: Choose eval scope
Determine what you are evaluating:
- Description triggering — does the agent activate the right skill for a given prompt? Tests the description field effectiveness.
- Output quality — does the skill produce correct results when activated? Tests the skill body and references.
- Convention compliance — does the output follow Plastic conventions?
Read
references/convention-checks.mdfor the full assertion library.
Multiple scopes can apply to the same skill. Start with the scope that addresses your immediate concern, add others as needed.
Step 2: Design test cases
Create evals/evals.json in the skill being evaluated. Copy the starter
template from assets/eval-template.json in this skill.
For description triggering:
- Write ~20 queries: 8-10 should-trigger, 8-10 should-not-trigger
- Split 60/40 into train and validation sets (proportional mix in each)
- Include near-miss negatives that share keywords but need a different skill
- In
expected_output, describe whether the skill should or should not activate and why
For output quality:
- Start with 2-3 test cases, expand after first results
- Use realistic user prompts with varied phrasing, detail level, and formality
- In
expected_output, describe what correct output looks like — not exact text - Use
filesarray for any input files the test needs
What ships with it
4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 142 lines · 99 tokens per session scan A 3b41538c478a
plastic-skill-evaluating is a skill published in the GitHub repository zalom/plastic (10 stars, last pushed 3d ago), licensed MIT. It adds 99 tokens to every session and 1,323 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
rulesync
Generates and syncs AI rule configuration files (.cursorrules, CLAUDE.md, copilot-instructions.md) across 20+ coding tools from a single source. Use when syncing AI rules, running rulesync commands, importing or generating rule files, or managing shared AI coding configurations.
agent-workspace-linux
Use when a task needs an isolated hidden Linux desktop or workspace-owned browser: GUI app QA, web/browser/shopping automation, sandboxed app observation, or stale workspace cleanup. Routes agent-workspace-linux MCP tools on demand. Does NOT apply to host desktop/Chrome control, generic MCP setup, or pure code/file…
ss-component
Generate a new UI component following the StyleSeed design conventions.
loongsuite-pilot-insight
基于 LoongSuite Pilot / AI Coding Agent 日志生成事件洞察、组织洞察、数据质量、研发效能和 AI Native 使用类 SLS 报表时使用;包含 AI Coding 事件表语义,以及团队报表可选的部门维表、deptuser 组织关系、指标口径和公共 CTE,通常与 sls-dashboard-builder 一起使用。.
map-review
Interactive 4-section code review using monitor, predictor, and evaluator agents plus the user and maintainer role reviewers on current changes. Use when reviewing a diff, PR, or staged work before merge. Do NOT use to plan or implement; use map-plan or map-efficient.
map-fast
Minimal workflow for small, low-risk changes — no planning, no learning.