Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/evol-ai/skillcompass/eval-securitygit clone --depth 1 https://github.com/Evol-ai/SkillCompassWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00625 |
| Opus 5 | $0.00000 | $0.00313 |
| Sonnet 5 | $0.00000 | $0.00125 |
| Haiku 4.5 | $0.00000 | $0.00063 |
Grade A, and why
eval-security scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 71 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/eval-security — Standalone Security Scan
Locale: All templates in this spec are written in English. Detect the user's language from the session and translate user-facing text at display time per SKILL.md's Global UX Rules. Dimension labels: see the canonical table in SKILL.md.
Arguments
<path>(required): Path to the SKILL.md file to scan.--verbose(optional): Show detailed findings including low severity.
Steps
Step 1: Load Target
Parse arguments. Use the Read tool to load the target SKILL.md file.
Step 2: L0 Built-in Scan
Use the Read tool to load {baseDir}/prompts/d3-security.md. Execute all 7 L0 check categories against the target skill content. Record findings.
Step 3: L1/L2 External Tools
Use the Read tool to load {baseDir}/shared/tool-instructions.md. Follow the L1 whitelist detection procedure: for each tool, use the Bash tool to check if installed, and invoke if found. Then check .skill-compass/config.json for L2 custom tools and invoke those.
Step 4: Aggregate
Merge all findings from L0 + L1 + L2. Deduplicate by (location, check_type), keeping highest severity. Add source field to each finding.
Step 5: Output
Output the D3 section of the evaluation result (conforming to the security portion of schemas/eval-result.json):
{
"dimension": "D3",
"dimension_name": "security",
"score": 8,
"max": 10,
"pass": true,
"findings": [],
"tools_used": ["builtin"],
"details": "..."
}
If --verbose is not set: omit findings with severity "low" from display (still count them in score).
After printing the result:
-
Findings exist AND neither
--internalnor--ciis set: print a status line then present choices:⚠ {N} security issue(s) found. [Fix security issues / View details / Done]- Fix security issues — invoke the fix workflow to address reported findings.
- View details — re-display all findings including those hidden by verbosity rules.
- Done — exit with no further action.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 71 lines · 0 tokens per session scan A 165b4023aeb6
eval-security is a command published in the GitHub repository Evol-ai/SkillCompass (215 stars, last pushed 4mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 625 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
story-long-scan
长篇网文扫榜。分析起点、番茄、晋江等平台排行数据,提炼市场趋势。.
audit_docs
Audit all project metadata files for stale counts, version mismatches, broken references, and missing entries. Reports discrepancies without auto-fixing.
merge_session
Merge session branch(es) into main via rebase + fast-forward push. Use from inside a session worktree pane to land your work, or with --all to batch merge all sessions from the main repo.
cherry_pick_pr
Cherry-pick one or more commits onto a new branch from main and open a PR. Useful when a commit landed on the wrong branch or you want to split a multi-commit branch into separate PRs.
cti-report
Render case deliverables — relationship graph (PNG/SVG/Mermaid) and a polished PDF/DOCX assessment. Usage: /cti-report [--graph|--pdf].
reskin
Extract a measured design signature (type ramp, accent + its budget, grid unit, radius, layout) from a reference image or site, write it to signature.json, and drive the build to match it — "steal this vibe" as a spec, not pixels.