Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/evol-ai/skillcompass/eval-comparegit clone --depth 1 https://github.com/Evol-ai/SkillCompassWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00753 |
| Opus 5 | $0.00000 | $0.00377 |
| Sonnet 5 | $0.00000 | $0.00151 |
| Haiku 4.5 | $0.00000 | $0.00075 |
Grade A, and why
eval-compare scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 80 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/eval-compare — Version Comparison
Locale: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per SKILL.md's Global UX Rules. Dimension labels: see the canonical table in SKILL.md.
Arguments
<version-a>(required): File path or{skill-name}@{version}identifier.<version-b>(required): File path or{skill-name}@{version}identifier.--internal/--ci(optional): Skip interactive choice prompts; output results and exit.
Steps
Step 1: Resolve Versions
For each argument:
- If it's a file path: use the Read tool to load the file directly.
- If it's a
name@versionidentifier: look up.skill-compass/{name}/snapshots/{version}.mdusing the Read tool. - If version not found: output
"Version not found: {identifier}"and stop.
Cross-skill check: If both arguments use name@version syntax and the skill names differ, warn the user that they are comparing different skills ({name_a} vs {name_b}) and the result may lack meaningful reference value, then present the choice:
[Continue comparing / Cancel]
If the user chooses Cancel, stop.
Step 2: Check Cached Results
For each version, check .skill-compass/{name}/manifest.json for cached evaluation results. Use the Read tool to load the manifest.
If cached results exist (matching content_hash): use cached scores. If not: run eval-skill flow on the version to generate fresh results.
Step 3: Compare
Generate a side-by-side comparison:
Version Comparison: sql-optimizer
| Dimension | v1.0.0 | v1.0.0-evo.2 | Delta |
|-----------------|--------|--------------|--------|
| D1 Structure | 6 | 7 | ↑ +1 |
| D2 Trigger | 3 | 6 | ↑ +3 * |
| D3 Security | 2 | 7 | ↑ +5 * |
| D4 Functional | 4 | 4 | → 0 |
| D5 Comparative | 3 | 3 | → 0 |
| D6 Uniqueness | 7 | 7 | → 0 |
|-----------------|--------|--------------|--------|
| Overall | 38 | 52 | ↑ +14 |
| Verdict | FAIL | CAUTION | |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 80 lines · 0 tokens per session scan A ffe4f93f44d6
eval-compare is a command published in the GitHub repository Evol-ai/SkillCompass (215 stars, last pushed 4mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 753 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
story-long-scan
长篇网文扫榜。分析起点、番茄、晋江等平台排行数据,提炼市场趋势。.
audit_docs
Audit all project metadata files for stale counts, version mismatches, broken references, and missing entries. Reports discrepancies without auto-fixing.
merge_session
Merge session branch(es) into main via rebase + fast-forward push. Use from inside a session worktree pane to land your work, or with --all to batch merge all sessions from the main repo.
cherry_pick_pr
Cherry-pick one or more commits onto a new branch from main and open a PR. Useful when a commit landed on the wrong branch or you want to split a multi-commit branch into separate PRs.
cti-report
Render case deliverables — relationship graph (PNG/SVG/Mermaid) and a polished PDF/DOCX assessment. Usage: /cti-report [--graph|--pdf].
reskin
Extract a measured design signature (type ramp, accent + its budget, grid unit, radius, layout) from a reference image or site, write it to signature.json, and drive the build to match it — "steal this vibe" as a spec, not pixels.