Getting it into your agent
It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.
git clone --depth 1 https://github.com/siddiqss/semantic-seo-suitenpx agentmods add skills/siddiqss/semantic-seo-suite/semantic-site-auditorWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/siddiqss/semantic-seo-suite/semantic-site-auditor)<a href="https://agentmods.dev/skills/siddiqss/semantic-seo-suite/semantic-site-auditor"><img src="https://agentmods.dev/badge/skills/siddiqss/semantic-seo-suite/semantic-site-auditor/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/siddiqss/semantic-seo-suite/semantic-site-auditor"><img src="https://agentmods.dev/badge/skills/siddiqss/semantic-seo-suite/semantic-site-auditor.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00111 | $0.01099 |
| Opus 5 | $0.00056 | $0.00549 |
| Sonnet 5 | $0.00022 | $0.00220 |
| Haiku 4.5 | $0.00011 | $0.00110 |
Grade A, and why
semantic-site-auditor scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 85 lines — stays where its author put it; the contents beside it link to each section on GitHub.
semantic-site-auditor
Turn a live site + its topical map into a prioritised, evidence-backed fix list. Every
finding carries provenance (measured from crawl, derived from embeddings) — no
invented scores. This is also the strongest pre-sales artifact for services work: run it
on a prospect's domain before a call.
Read first: ../../framework/topical-map-theory.md (gap/section logic),
../../framework/internal-linking-rules.md (orphans), ../../framework/eav-modeling.md
(drift = claimed vs perceived identity).
Preconditions
brands/<slug>/config.yaml(tier).- A
topical-map.jsonto audit against. If none exists, run topical-map-builder first (you can't measure coverage without a target map). entity-profile.json(for the claimed-identity centroid used by drift).
Workflow
-
Crawl + extract the site into
brands/<slug>/data/crawl/:python ../../scripts/crawl_sitemap.py --domain <domain> --max-pages 500 --out /tmp/urls.json python ../../scripts/extract_page_content.py --urls /tmp/urls.json --out-dir brands/<slug>/data/crawlRobustness:
crawl_sitemap.pyfalls back to a same-domain BFS when there's no sitemap, respects robots.txt, and caps pages/depth. For JS-rendered sites, extraction may be thin — note that in the report rather than treating missing content as a gap. -
Run the analysis engine:
python ../../scripts/audit_site.py \ --map brands/<slug>/topical-map.json --crawl-dir brands/<slug>/data/crawl \ --entity-profile brands/<slug>/entity-profile.json \ --locked brands/<slug>/locked-facts.json --brand "<Brand>" \ --out brands/<slug>/audits/<date>-full.mdIt matches each page to its nearest map node (embedding cosine) and reports:
- Coverage gaps — nodes with no matching page (
derived). - Cannibalization — page pairs with high similarity + same intent, and any node
hit by multiple pages (
derived). - Entity drift — pages farthest from the core-section centroid, i.e. content
pulling the site away from its claimed identity (
derived). At T0 (no embeddings), do this qualitatively and label itasserted— never emit a fake distance. - Orphans — pages with no incoming internal links (
measuredfrom the crawl link graph). - Micro-audit — per-page lint (naked stats, fluff, unanswered question headings) reusing the draft rules.
- Coverage gaps — nodes with no matching page (
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 85 lines · 111 tokens per session scan A 751c3af7575b
semantic-site-auditor is a skill published in the GitHub repository siddiqss/semantic-seo-suite (9 stars, last pushed 2mo ago), licensed MIT. It adds 111 tokens to every session and 1,099 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
fire-your-seo-agency
A procedure for improving how a website appears in search engines and how AI answer systems find and cite it. It covers search, answer-engine, generative-AI, and Naver visibility.
geo-loop
Run one bounded eGEOagents loop iteration over a workspace domain - read the charter and fresh collector data, do ONE unit of work, write substrate artifacts, append one Timeline entry and one LOG line. Use for loop mode, /geo:loop, scheduled GEO runs, or continuous monitoring.
content-scoring
Score content against the 10 GEO criteria with evidence and prioritized fixes. Use when users ask to score, rate, evaluate, or estimate ranking strength.
schema-generator
Generate JSON-LD schema markup for pages and content types with an implementation checklist. Use when users ask for schema, structured data, rich snippets, or markup.
competitive-analysis
Analyze AI-search competitors for a query and recommend ranking strategy. Use when users ask competitor analysis, who ranks, or competitive landscape.
validation-doctor
Check Brave Search and Chrome DevTools MCP availability and provide exact setup snippets. Use when validation dependencies are missing or uncertain.