Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add AnthonyAlcaraz/agentic-graph-rag-skills --skill skill-quality-evaluatorgit clone --depth 1 https://github.com/AnthonyAlcaraz/agentic-graph-rag-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/anthonyalcaraz/agentic-graph-rag-skills/skill-quality-evaluator)<a href="https://agentmods.dev/skills/anthonyalcaraz/agentic-graph-rag-skills/skill-quality-evaluator"><img src="https://agentmods.dev/badge/skills/anthonyalcaraz/agentic-graph-rag-skills/skill-quality-evaluator/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/anthonyalcaraz/agentic-graph-rag-skills/skill-quality-evaluator"><img src="https://agentmods.dev/badge/skills/anthonyalcaraz/agentic-graph-rag-skills/skill-quality-evaluator.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00158 | $0.02522 |
| Opus 5 | $0.00079 | $0.01261 |
| Sonnet 5 | $0.00032 | $0.00504 |
| Haiku 4.5 | $0.00016 | $0.00252 |
Grade A, and why
skill-quality-evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 184 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Skill Quality Evaluator
Overview
RAG and gateways (the two sibling Ch6 skills) solve the routing problem: which skill matches this task? They do not solve the quality problem: is this skill worth trusting once matched? Raw skill libraries accumulate junk. As a repository scales from dozens to hundreds to thousands of skills, retrieval quality degrades because the agent pulls low-quality or unsafe skills alongside the useful ones.
The chapter grounds this in two research anchors:
- SkillsBench (DAIR.AI). Curated skills improve task performance by 16.2%. Self-generated skills — produced by agents without human review — show zero improvement over baseline. Quality matters more than quantity, and 2-3 focused skills per task is the optimal number.
- SkillNet (2026). A repository of 200,000+ reusable skills, each rated across five quality dimensions. Agents using SkillNet-rated skills achieved 40% higher average rewards and 30% fewer execution steps across ALFWorld, WebShop, and ScienceWorld — held across multiple backbone models, confirming quality-rated retrieval is a model-agnostic infrastructure layer.
The five dimensions (Table 6-1) and the composite weighting:
| Dimension | Weight | Maps to |
|---|---|---|
| safety | 2.0 | security scanning |
| completeness | 1.0 | dependency checking |
| executability | 2.0 | hallucinated-tool-call risk |
| maintainability | 1.0 | composability |
| cost_awareness | 1.0 | token economics (Ch8) |
Composite = weighted mean with safety and executability at double weight, so a skill that hallucinates tool calls (low executability) or runs unconstrained shell commands (low safety) ranks low regardless of how well it documents.
When to Use
- A skill library has grown past ~20 entries and mixed provenance (vendor, team-authored, agent-generated)
- Retrieval is surfacing plausible-but-junk skills alongside good ones
- You need a deployment knob (
min_quality) that encodes organizational risk appetite (research 0.4 → production healthcare 0.8+) - You are integrating quality as a second ranking signal on top of relevance
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 184 lines · 158 tokens per session scan A aff8c6117eef
skill-quality-evaluator is a skill published in the GitHub repository AnthonyAlcaraz/agentic-graph-rag-skills (10 stars, last pushed 2mo ago), licensed MIT. It adds 158 tokens to every session and 2,522 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
graphify
Use for any question about a codebase, its architecture, file relationships, or project content — especially when graphify-out/ exists, where the question should be treated as a graphify query first. Turns any input (code, docs, papers, images, videos) into a persistent knowledge graph with god nodes, community…
lemmalog
Externalize working memory and logical state into the lemmalog Datalog engine (MCP). Use for ANY multi-step task where state should outlive one context window or span agents: long investigations, debugging sessions, audits, multi-agent searches, systematic explorations, planning with many interdependent constraints…
jurisd-research
Expert Australian/NZ legal research and AGLC4 citation using the jurisd MCP server. Use when finding cases or legislation (AustLII), looking up a provision offline, formatting or resolving citations, building a pinpoint, tracing who-cites-what, or producing an AGLC4 bibliography. Triggers on case law, legislation…
repo_search
Search repository text through a deterministic first-class fak command.
container-manager-kg-ingestion
Snapshot a host's Docker/Podman/Swarm inventory into the epistemic-graph knowledge graph as typed OWL nodes via the container-manager-mcp MCP server — containers, images, volumes, networks, swarm services and nodes, with their :usesImage / :runsOn / :builtFrom links. Use when the agent must record live container state…
cortex-design
Use this skill to generate well-branded interfaces and assets for Cortex, either for production or throwaway prototypes/mocks/etc. Contains essential design guidelines, colors, type, fonts, assets, and UI kit components for prototyping.