Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/aaronsb/claude-code-config/ways-testsnpx skills add aaronsb/claude-code-config --skill ways-testsgit clone --depth 1 https://github.com/aaronsb/claude-code-configWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/aaronsb/claude-code-config/ways-tests)<a href="https://agentmods.dev/skills/aaronsb/claude-code-config/ways-tests"><img src="https://agentmods.dev/badge/skills/aaronsb/claude-code-config/ways-tests.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00049 | $0.06917 |
| Opus 5 | $0.00024 | $0.03459 |
| Sonnet 5 | $0.00010 | $0.01383 |
| Haiku 4.5 | $0.00005 | $0.00692 |
Grade A, and why
ways-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 702 lines — stays where its author put it; the contents beside it link to each section on GitHub.
ways-tests: Way Matching & Vocabulary Tool
Test how well a way matches sample prompts, analyze vocabulary for gaps, and validate frontmatter.
Usage
/ways-tests score <way> "prompt" # Score one way against a prompt
/ways-tests score-all "prompt" # Rank all ways against a prompt
/ways-tests suggest <way> # Analyze vocabulary gaps
/ways-tests suggest <way> --apply # Update vocabulary in-place
/ways-tests suggest --all [--apply] # Analyze/update all ways
/ways-tests lint <way> # Validate frontmatter
/ways-tests lint --all # Validate all ways
/ways-tests check <check> "context" # Test check scoring curve
/ways-tests check-all "context" # Rank all checks against context
/ways-tests tree <path> # Analyze progressive disclosure tree structure
/ways-tests budget <path> # Token cost analysis for a way tree
/ways-tests jaccard <tree> # Sibling vocabulary isolation for a tree
/ways-tests jaccard <way1> <way2> # Vocabulary overlap between two specific ways
/ways-tests crowding "prompt" # Detect vocabulary crowding across all ways
/ways-tests compare <path1> <path2> # Side-by-side tree metrics comparison
/ways-tests metrics # Show tree disclosure metrics for current session
/ways-tests embed-status # Embedding engine health dashboard
/ways-tests embed-score <way> "prompt" # Cosine similarity score for one way
/ways-tests embed-score-all "prompt" # Cosine similarity ranking across all ways
Engine Hierarchy
The matching pipeline has three tiers. Always report which tier is active before presenting scores — the numbers mean different things across tiers:
Tier 1 — Embedding (primary, ~20ms batch)
ways embed: cosine similarity, 0–1 scale
Threshold field: embed_threshold (per-way, default: 0.35)
Fires when: cosine(query, way) >= embed_threshold
Tier 2 — BM25 (fallback when embedding model missing)
ways match: BM25 relevance score, unbounded scale
Threshold field: threshold (per-way, default: 2.0)
Fires when: BM25_score >= threshold
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 702 lines · 49 tokens per session scan A 94f8b02b1e21
ways-tests is a skill published in the GitHub repository aaronsb/claude-code-config (18 stars, last pushed 5mo ago), licensed MIT. It adds 49 tokens to every session and 6,917 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
achieving-cmmc-level-2-compliance
Prepare a defense-contractor environment for CMMC Level 2 certification: scope CUI and FCI, implement the 110 NIST SP 800-171 Rev 2 security requirements across 14 families, compute the SPRS score with the DoD Assessment Methodology, manage a compliant POA&M, and ready the organization for a C3PAO assessment. Use when…
task-planning-arch
计算任务 gap 并产出下一步可执行子任务 List[TaskSpec];gap 已闭返回空数组。对齐 arch 场景(架构师名册/技术栈概览/双视角分析)确定式分解——按根目标交付物集合 + donechildren 查表(参照 task-planning storage 特例,非自由 LLM 分解)。.
task-search
在框架预查的候选 bot 集里决出执行者(who)与协作方式(how),返回 4 态 SearchResult(HITSINGLE/HITGROUP/HITMULTIBOTS/MISS)。对齐案例剧本确定式映射。.
escalation-governance
Assess whether to escalate models. Use when evaluating reasoning depth.
agt-policy-authoring
Create and validate a minimal AGT Copilot CLI policy tailored to the repository being inspected.
silicon-throne
硅基圣座总控核心,以神谕式统御把零散任务压成 末将-first 之文明工程.