Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/evol-ai/skillcompass/eval-evolvegit clone --depth 1 https://github.com/Evol-ai/SkillCompassWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.02199 |
| Opus 5 | $0.00000 | $0.01099 |
| Sonnet 5 | $0.00000 | $0.00440 |
| Haiku 4.5 | $0.00000 | $0.00220 |
Grade A, and why
eval-evolve scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 205 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/eval-evolve — Optional Plugin-Assisted Multi-Round Evolution via Ralph Loop
Locale: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per SKILL.md's Global UX Rules. Dimension labels: see the canonical table in SKILL.md.
Arguments
<path>(required): Path to the SKILL.md file to evolve.--max-iterations <n>(optional): Max improvement rounds. Default: 6.--target-score <n>(optional): Stop when overall_score >= n. Default: 70.--internal(optional): Skip all interactive prompts. Used when this command is called programmatically by another command or script.
Prerequisites
-
Recommended model: Claude Opus 4.6 (
claude-opus-4-6). Multi-round evolution requires consistent scoring across iterations to detect genuine improvements vs noise. Weaker models may cause the evolution loop to oscillate rather than converge. -
This command requires the ralph-wiggum plugin. If not installed, present the user with a choice before attempting any plugin call:
┌─ Plugin required: ralph-wiggum ───────────────────────┐ │ This command depends on the ralph-wiggum plugin to │ │ run the multi-round evolution loop. │ │ │ │ [Install ralph-wiggum plugin] [Cancel] │ └────────────────────────────────────────────────────────┘- If the user chooses Install ralph-wiggum plugin: run
claude plugin install ralph-wiggum@claude-code-pluginsand continue. - If the user chooses Cancel: stop immediately with no further action.
- If
--internalis passed, skip the prompt and run the install command directly. - Security note: No third-party code is fetched or executed without explicit user consent. The plugin install only proceeds after the user actively selects "Install".
- If the user chooses Install ralph-wiggum plugin: run
What This Command Does
Generates and executes a /ralph-loop invocation that chains /eval-skill → /eval-improve automatically until the skill reaches PASS verdict (or hits the iteration limit). This is a power-user workflow, not the default path for normal evaluations.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 205 lines · 0 tokens per session scan A 8fd010d541a7
eval-evolve is a command published in the GitHub repository Evol-ai/SkillCompass (216 stars, last pushed 4mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 2,199 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
story-long-scan
长篇网文扫榜。分析起点、番茄、晋江等平台排行数据,提炼市场趋势。.
audit_docs
Audit all project metadata files for stale counts, version mismatches, broken references, and missing entries. Reports discrepancies without auto-fixing.
merge_session
Merge session branch(es) into main via rebase + fast-forward push. Use from inside a session worktree pane to land your work, or with --all to batch merge all sessions from the main repo.
cherry_pick_pr
Cherry-pick one or more commits onto a new branch from main and open a PR. Useful when a commit landed on the wrong branch or you want to split a multi-commit branch into separate PRs.
cti-report
Render case deliverables — relationship graph (PNG/SVG/Mermaid) and a polished PDF/DOCX assessment. Usage: /cti-report [--graph|--pdf].
reskin
Extract a measured design signature (type ramp, accent + its budget, grid unit, radius, layout) from a reference image or site, write it to signature.json, and drive the build to match it — "steal this vibe" as a spec, not pixels.