Borrowing it
Nothing to install: this file belongs to sparq-org/sparq. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/sparq-org/sparq/main/.claude/skills/hdt-format/SKILL.mdgit clone --depth 1 https://github.com/sparq-org/sparqWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/sparq-org/sparq/hdt-format)<a href="https://agentmods.dev/skills/sparq-org/sparq/hdt-format"><img src="https://agentmods.dev/badge/skills/sparq-org/sparq/hdt-format/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/sparq-org/sparq/hdt-format"><img src="https://agentmods.dev/badge/skills/sparq-org/sparq/hdt-format.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00110 | $0.01442 |
| Opus 5 | $0.00055 | $0.00721 |
| Sonnet 5 | $0.00022 | $0.00288 |
| Haiku 4.5 | $0.00011 | $0.00144 |
Grade A, and why
hdt-format scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 93 lines — stays where its author put it; the contents beside it link to each section on GitHub.
HDT — Header Dictionary Triples
[OPUS-4.8] Authored for the active HDT work. Ground truth: the sparq-hdt crate
(crates/sparq-hdt/src/lib.rs, Cargo.toml; open work in beads — bd list -l area:sparq-hdt)
and the upstream hdt crate (KonradHoeffner/hdt, MIT). Verify crate API against the pinned
version before writing code.
What HDT is
HDT (rdfhdt.org, W3C member submission) is the de-facto binary archive format for RDF. A single self-contained file is split into three components — hence the name:
- H — Header. RDF metadata about the dataset itself (VoID statistics, format, provenance) as plain triples. Queryable; only the head of the stream needs decoding to read it.
- D — Dictionary. All distinct terms, stored compressed. The standard layout is a FourSectionDictionary: four sections (subjects-only, predicates, objects-only, and shared subject∧object terms), each Plain Front Coded (PFC) — terms are sorted and each stores only the byte-length of its common prefix with the previous term plus its suffix. A term maps to a dense integer id per section.
- T — Triples. A BitmapTriples structure: the triples in SPO order as a bitmap-compressed adjacency list (Log64/Plain integer sequences + a bitmap marking the end of each subject's / predicate's adjacency run). Lookups use rank/select over the bitmaps to navigate from a subject id to its predicate run to its object run in O(1)-ish per step.
The payoff: files are a fraction of the size of even gzipped N-Triples, and they load without text parsing — no tokenizing, no UTF-8 re-validation per triple.
How sparq-hdt uses it (the actual integration)
The crate wraps the maintained hdt crate's reader rather than
reimplementing the binary format. Key design facts (in lib.rs / Cargo.toml):
- Id-level translation. Each distinct HDT dictionary id is decompressed to its
term string once, interned into the sparq
Dict, and the mapping memoized in a flat per-section table. So the term set is never materialized twice and per-triple work is three array lookups (s/p/o id → sparq Id). This is the fast-decode lever: do not route throughsophiaterms —default-features = falsedeliberately drops the sophia adapter. .hdt.gzby magic bytes. GZipped containers are detected by magic bytes (0x1f 0x8b), not file name, in every entry point, and decompressed on the fly with a streamingMultiGzDecoder(flate2, pure-Rust backend) — same magic-byte convention as the rest of the workspace.- Header access.
header(path)/header_reader(reader)expose the dataset metadata triples as a queryable sparqGraph, decoding only the control-info + header sections (a few KB). The wrapped crate parses the header internally but keeps the field private, so these re-read those sections directly. - Entry point:
sparq_hdt::load("dataset.hdt") -> Result<Graph, Error>.ErrorisIo | Hdt | Term.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 93 lines · 110 tokens per session scan A 260494543dda
hdt-format is a skill published in the GitHub repository sparq-org/sparq (12 stars, last pushed today), licensed MIT. It adds 110 tokens to every session and 1,442 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
create-pr
Creates a GitHub PR with a Linear-ticket-prefixed title and a decision-led, narrative description for prisma-next. Use when the user wants to create a pull request, open a PR, or submit changes for review.
schema-exploration
Lists tables, describes columns and data types, identifies foreign key relationships, and maps entity relationships in a database. Use when the user asks about database schema, table structure, column types, what tables exist, ERD, foreign keys, or how entities relate.
ha-data-stores
Map of Hope Agent's local data stores and safe read-only query workflow. Use when the user asks where Hope Agent stores data, wants to inspect sessions/messages/memory/logs/background jobs/knowledge indexes/settings, asks the model to query local app data, or debugging requires checking persisted state. Trigger…
supabase
Supabase / PostgREST Row-Level-Security playbook — pull the anon (or leaked servicerole) key out of the frontend JS, map tables from the auto-generated OpenAPI spec, test anonymous RLS READ disclosures (PII/secret leaks), and anonymous RLS WRITE abuse (insert/update/delete — e.g. forging…
nornicdb-cypher-queries
Pick fast, predictable Cypher query shapes in NornicDB — point lookups, batch retrieval, pagination, search, traversal, batched UNWIND/MERGE writes, cleanup, multi-tenant isolation. Use when writing or reviewing Cypher whose latency or throughput matters; maps user intent to the executor's hot-path query templates.
dsql
Build with Aurora DSQL — manage schemas, execute queries, handle migrations, diagnose query plans, diagnose cluster performance, load data, and develop applications with a serverless, distributed SQL database. Covers IAM auth, multi-tenant patterns, MySQL-to-DSQL and PostgreSQL-to-DSQL schema conversion, foreign key…