Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add AnthonyAlcaraz/agentic-graph-rag-skills --skill evolution-taxonomy-classifiergit clone --depth 1 https://github.com/AnthonyAlcaraz/agentic-graph-rag-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/anthonyalcaraz/agentic-graph-rag-skills/evolution-taxonomy-classifier)<a href="https://agentmods.dev/skills/anthonyalcaraz/agentic-graph-rag-skills/evolution-taxonomy-classifier"><img src="https://agentmods.dev/badge/skills/anthonyalcaraz/agentic-graph-rag-skills/evolution-taxonomy-classifier/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/anthonyalcaraz/agentic-graph-rag-skills/evolution-taxonomy-classifier"><img src="https://agentmods.dev/badge/skills/anthonyalcaraz/agentic-graph-rag-skills/evolution-taxonomy-classifier.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00206 | $0.03093 |
| Opus 5 | $0.00103 | $0.01546 |
| Sonnet 5 | $0.00041 | $0.00619 |
| Haiku 4.5 | $0.00021 | $0.00309 |
Grade A, and why
evolution-taxonomy-classifier scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 189 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evolution Taxonomy Classifier
Overview
Your agent has diagnosed a reasoning failure. The execution graph shows where it went wrong and the cognitive-fault isolator has classified the breakdown. Now what? The temptation is to jump straight to a fix: tweak a prompt, retrain an adapter, add a guardrail. But self-evolution spans more than a single lever: it is a four-dimensional design space, and pulling the wrong lever wastes compute, introduces regressions, or both.
This skill locates any proposed evolution on the four Gao et al. axes:
- WHAT evolves —
model(weights or prompts; needs execution graphs for causal tracing) |context(retrieval; this IS graph evolution: rewire edges, merge nodes, prune subgraphs) |tool(rewires the tool subgraph by reweighting task-type-to-tool edges on observed success) |architecture(graph surgery on the workflow graph itself). - WHEN it fires —
intra_test_time(within one request; must be sub-second; Reflect-Retry-Reward) |inter_test_time(between requests; can afford fine-tuning or graph restructuring; SEAL overnight, semantic backpropagation). - HOW the agent learns —
reward_based(scalar signals: InfoGain, user satisfaction) |imitation_based(copy successful trajectories) |population_based(maintain variants, select fittest). - WHERE it applies —
general_purpose(all tasks) |domain_specialized(one vertical, e.g. cascade failures in microservice topologies).
Every axis requires graph structure to operate. As the chapter states: the
graph is not optional infrastructure here, it is the substrate that makes any
of these axes operable. Alshikh's production research reinforces the point: the
first methodology is GNN-inspired, where "each adaptation becomes a traceable
node." The classifier attaches that graph-dependency rationale to every value
it assigns, and route_failure maps a diagnosed failure to its axis per Table
7-1.
When to Use
- AFTER a diagnostic report exists and you are deciding which evolution lever to pull
- To sanity-check a proposed evolution: is this really model evolution, or is it context evolution wearing a model-evolution costume?
- To route a diagnosed failure type (FORMAT, REASONING, KNOWLEDGE) to its primary axis, timing, and mechanism
- When designing a self-evolution loop and you need the four-axis vocabulary to keep intra-test-time and inter-test-time paths distinct
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 189 lines · 206 tokens per session scan A 528dfaa72e62
evolution-taxonomy-classifier is a skill published in the GitHub repository AnthonyAlcaraz/agentic-graph-rag-skills (10 stars, last pushed 1mo ago), licensed MIT. It adds 206 tokens to every session and 3,093 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
graphify
Use for any question about a codebase, its architecture, file relationships, or project content — especially when graphify-out/ exists, where the question should be treated as a graphify query first. Turns any input (code, docs, papers, images, videos) into a persistent knowledge graph with god nodes, community…
lemmalog
Externalize working memory and logical state into the lemmalog Datalog engine (MCP). Use for ANY multi-step task where state should outlive one context window or span agents: long investigations, debugging sessions, audits, multi-agent searches, systematic explorations, planning with many interdependent constraints…
jurisd-research
Expert Australian/NZ legal research and AGLC4 citation using the jurisd MCP server. Use when finding cases or legislation (AustLII), looking up a provision offline, formatting or resolving citations, building a pinpoint, tracing who-cites-what, or producing an AGLC4 bibliography. Triggers on case law, legislation…
repo_search
Search repository text through a deterministic first-class fak command.
container-manager-kg-ingestion
Snapshot a host's Docker/Podman/Swarm inventory into the epistemic-graph knowledge graph as typed OWL nodes via the container-manager-mcp MCP server — containers, images, volumes, networks, swarm services and nodes, with their :usesImage / :runsOn / :builtFrom links. Use when the agent must record live container state…
cortex-design
Use this skill to generate well-branded interfaces and assets for Cortex, either for production or throwaway prototypes/mocks/etc. Contains essential design guidelines, colors, type, fonts, assets, and UI kit components for prototyping.