Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
/plugin marketplace add air-gapped/skillsnpx agentmods add plugins/air-gapped/skills/inference-cachegit clone --depth 1 https://github.com/air-gapped/skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/plugins/air-gapped/skills/inference-cache)<a href="https://agentmods.dev/plugins/air-gapped/skills/inference-cache"><img src="https://agentmods.dev/badge/plugins/air-gapped/skills/inference-cache.svg" alt="Measured on agentmods" height="20"></a>Grade A, and why
inference-cache scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
{
"name": "inference-cache",
"source": "./",
"description": "Inference KV-cache and transport suite — LMCache multiprocess (MP) standalone-server mode (DaemonSet + Deployment K8s pattern, ZMQ connector, L1 + L2 NIXL/POSIX/GDS/HF3FS/fs/s3/mooncake adapters) and NVIDIA NIXL transfer library (UCX/GDS/Mooncake/libfabric/HF3FS/S3 plugins, agent API, telemetry) used by Dynamo/vLLM/SGLang. Pairs with the vllm-caching skill in the `vllm` suite and sglang-hicache in the `sglang` suite.",
"version": "0.20260819.12",
"author": {
"name": "Jörgen"
},
"license": "MIT",
"skills": [
"./.claude/skills/lmcache-mp",
"./.claude/skills/nvidia-nixl"
],
"strict": false,
"category": "inference",
"tags": [
"kv-cache",
"lmcache",
"nixl",
"vllm",
"sglang",
"dynamo",
"offload",
"prefix-caching",
"disaggregated-prefill",
"kubernetes"
]
}What it installs
The manifest is a name and a version. 2 skills travel with it, and installing the plugin installs all of them — 426 tokens a session between them. Each is measured on its own page, and each can be installed alone.
What ships with it
1 file beside marketplace.json#inference-cache in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 28 lines scan A 7694d9c29de2
inference-cache is a plugin published in the GitHub repository air-gapped/skills (5 stars, last pushed 3d ago), licensed MIT. Its token cost is not measured: this kind of file is read by the harness, not the model. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other plugins, from other repositories
rag-core
Document indexing and hybrid retrieval with heading-boundary chunking, FAISS vector store, and PageIndex tree filtering.
rag-reviewer
Agentic PR review: hybrid RAG + code graph via MCP, review skills for Claude Code.
veritasreason
Full-stack knowledge graph skills: semantic extraction, decision intelligence, context graphs, reasoning, explainability, ontology, provenance, deduplication, visualization, and multi-format export.
graph-source
Connect a property graph (knowledge graph, GraphRAG corpus, Apache AGE / openCypher data) to Skardi and query it through SQL, end to end: provision the AGE backend with a least-privilege role, declare the type: graph source and its views in context YAML, verify registration health, write correct queries (cypherquery.
Claude Dev Assistant marketplace
An intelligent development workflow featuring Socratic interviewing, dual-agent debate, expert team coordination, and RAG system design. Built on Claude Code for seamless AI-assisted development.
nextjs
Official Next.js skills: adopt and optimize Cache Components, adopt Partial Prefetching, and verify runtime behavior against a running dev server.