Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/bjcoombs/ai-native-toolkit/assessnpx skills add bjcoombs/ai-native-toolkit --skill assessgit clone --depth 1 https://github.com/bjcoombs/ai-native-toolkitWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/bjcoombs/ai-native-toolkit/assess)<a href="https://agentmods.dev/skills/bjcoombs/ai-native-toolkit/assess"><img src="https://agentmods.dev/badge/skills/bjcoombs/ai-native-toolkit/assess.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00113 | $0.12473 |
| Opus 5 | $0.00056 | $0.06236 |
| Sonnet 5 | $0.00023 | $0.02495 |
| Haiku 4.5 | $0.00011 | $0.01247 |
Grade B, and why
assess scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Asks for rootmediumPrivilege escalation
A mod that escalates privileges can change anything on the machine, not only the project.
command -v apt >/dev/null && sudo apt install -y scc \ How it starts
The opening of the file, as written. The whole thing — 499 lines — stays where its author put it; the contents beside it link to each section on GitHub.
AI Readiness Assessment + Complexity Hotspot
Three artefacts in one pass against a target repo:
- Layered contract assessment - 0-8 score across navigability, runtime liveness, code design, linters, architecture tests, CI, coverage, review bots, and AI project management.
- Complexity hotspot SVG - Codecov-style treemap of the code. Size = LOC. Colour = cyclomatic complexity. Saturation = recent git churn. Vivid red = complex AND active = riskiest to change.
- Doc navigability SVG - a node-graph of the docs. Structure = connectivity (centre = entry, rim = unreachable, dashed ring = orphan); colour = staleness (vivid red = a frozen doc beside churning code = a lying map); size = file length. Folds navigability and the decaying-map signal into one artifact.
Both SVGs are colour-blind-safe by default (OrRd ramp, no red-green).
All land as files inside the target repo. The skill always writes them locally; after writing, ask the user whether to open a PR in the target repo with the artefacts.
The model: truth-pressure, not presence
Read this before scoring - it changes how you score. Across every layer, the real signal is never presence. It is whether a thing is under active pressure to stay true:
- Tests keep behaviour honest (CI fails when it's wrong).
- Retros / feedback loops keep the process honest (Layer 8 scores whether retros are carried out, not merely present).
- Maintenance keeps docs honest (a wiki tracked against code churn).
- Telemetry / liveness keeps relevance honest (is this code actually exercised).
So AI-readiness is the degree to which a codebase's self-descriptions are kept honest, not the degree to which scaffolding exists. Score artefacts on maintenance pressure, not existence. A stale-but-present doc scores at or below absent: missing makes the agent go look; confidently-stale makes it navigate fast to a wrong, current-looking conclusion.
The 9 layers (0-8) fall into three bands, ordered by dependency - what must hold for the next band to mean anything:
What ships with it
60 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- pyproject.toml 3.2 KB
- references/actions-schema.md 5.0 KB
- references/consent-lifecycle.md 6.4 KB
- references/monorepo-scoping.md 3.9 KB
- references/uninstall.md 4.3 KB
- scripts/assess_core.py 73 KB runs code
- scripts/assess_emit_workflow.py 3.6 KB runs code
- scripts/assess_finalize.py 22 KB runs code
- scripts/assess_gate.py 11 KB runs code
- scripts/assess_report.py 19 KB runs code
- scripts/complexity-treemap.py 47 KB runs code
- scripts/doc-graph-svg.py 23 KB runs code
- scripts/lib/__init__.py 446 B runs code
- scripts/lib/accretion_ratchet.py 15 KB runs code
- scripts/lib/agent_instructions_grader.py 17 KB runs code
- scripts/lib/agent_ops.py 5.6 KB runs code
- scripts/lib/anomaly_detector.py 2.9 KB runs code
- scripts/lib/archetype.py 16 KB runs code
- scripts/lib/assess_config.py 14 KB runs code
- scripts/lib/badge.py 7.4 KB runs code
- scripts/lib/change_coupling.py 17 KB runs code
- scripts/lib/ci_workflow.py 5.1 KB runs code
- scripts/lib/coupling_analysis.py 11 KB runs code
- scripts/lib/coverage_report.py 6.6 KB runs code
- scripts/lib/decline_markers.py 7.5 KB runs code
- scripts/lib/doc_complexity_join.py 16 KB runs code
- scripts/lib/doc_graph.py 43 KB runs code
- scripts/lib/doc_provenance.py 9.2 KB runs code
- scripts/lib/doc_staleness.py 19 KB runs code
- scripts/lib/git_churn.py 16 KB runs code
- scripts/lib/interactivity.py 4.4 KB runs code
- scripts/lib/jvm_capabilities.py 17 KB runs code
- scripts/lib/keyhole_signals.py 53 KB runs code
- scripts/lib/liveness_scan.py 28 KB runs code
- scripts/lib/ownership_parser.py 19 KB runs code
- scripts/lib/promissory_markers.py 19 KB runs code
- scripts/lib/raw_source.py 6.5 KB runs code
- scripts/lib/README.md 29 KB
- scripts/lib/stats_diff.py 3.2 KB runs code
- scripts/lib/structure_drift.py 30 KB runs code
- scripts/lib/structure_graph.py 21 KB runs code
- scripts/lib/test_focus.py 8.8 KB runs code
- scripts/lib/test_pressure/__init__.py 2.7 KB runs code
- scripts/lib/test_pressure/aggregate.py 3.9 KB runs code
- scripts/lib/test_pressure/common.py 1.8 KB runs code
- scripts/lib/test_pressure/heuristics.py 18 KB runs code
- scripts/lib/test_pressure/mutation.py 18 KB runs code
- scripts/lib/treemap_render.py 13 KB runs code
- scripts/lib/understanding_analysis.py 11 KB runs code
- scripts/lib/vault_queries.py 9.5 KB runs code
- scripts/lib/wiki_writer.py 20 KB runs code
- templates/assess-gate.yml.template 2.2 KB
- templates/hotspot.md.template 547 B
- templates/index.md.template 815 B
- templates/log_entry.md.template 352 B
- tests/__init__.py 0 B runs code
- tests/conftest.py 3.7 KB runs code
- tests/fixtures/.gitkeep 0 B
- tests/fixtures/bad_instructions.md 210 B
- tests/fixtures/coverage.xml 855 B
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 499 lines · 113 tokens per session scan B a0052aa49e76
assess is a skill published in the GitHub repository bjcoombs/ai-native-toolkit (30 stars, last pushed 1mo ago), licensed Apache-2.0. It adds 113 tokens to every session and 12,473 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it B with 1 finding (asks for root). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…