Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add xberg-io/tree-sitter-language-pack --skill chunking-for-llmsgit clone --depth 1 https://github.com/xberg-io/tree-sitter-language-packWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/xberg-io/tree-sitter-language-pack/chunking-for-llms)<a href="https://agentmods.dev/skills/xberg-io/tree-sitter-language-pack/chunking-for-llms"><img src="https://agentmods.dev/badge/skills/xberg-io/tree-sitter-language-pack/chunking-for-llms/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/xberg-io/tree-sitter-language-pack/chunking-for-llms"><img src="https://agentmods.dev/badge/skills/xberg-io/tree-sitter-language-pack/chunking-for-llms.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00059 | $0.00692 |
| Opus 5 | $0.00030 | $0.00346 |
| Sonnet 5 | $0.00012 | $0.00138 |
| Haiku 4.5 | $0.00006 | $0.00069 |
Grade A, and why
chunking-for-llms scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 81 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Syntax-aware chunking for LLMs
Splitting code on a fixed byte or line count cuts functions in half and
strips context. ts-pack process <file> --chunk-size <bytes> splits on
syntactic boundaries (whole functions, classes, blocks) so each chunk is a
coherent unit, and emits them in the JSON chunks array.
Quick recipe
# ~2 KB chunks aligned to syntax boundaries
ts-pack process src/app.ts --chunk-size 2000
--chunk-size is a maximum size in bytes. The splitter packs whole
syntactic units up to that bound; an oversized single construct becomes its
own chunk rather than being cut. Chunks are added to the normal process
JSON output under chunks.
Picking a size
- Match the downstream model's token budget. A rough rule: bytes ÷ 4 ≈
tokens for code, so
--chunk-size 4000is on the order of ~1k tokens. - Larger chunks preserve more local context but fit fewer per request.
- Leave headroom for the prompt, the surrounding messages, and the response — do not size chunks to the full context window.
Combining with extraction
Chunking composes with the other process features, so you can attach
structure metadata to each request:
ts-pack process src/service.py --structure --chunk-size 3000 \
| jq '{chunks: (.chunks | length), functions: (.structure | length)}'
Chunk output
chunks is a list of code-chunk objects in the process JSON. Each chunk
carries its source text plus span information (line/byte offsets), so you
can cite or re-locate a chunk back in the original file. Iterate the array
to feed an LLM one coherent unit at a time:
ts-pack process big_module.py --chunk-size 2500 \
| jq -c '.chunks[]'
SDK equivalent
The SDK exposes chunking through the process config: set the
chunk_max_size field (in bytes) on ProcessConfig — the same value the
CLI's --chunk-size flag sets. ProcessConfig is a frozen dataclass, so
pass it to the constructor:
from tree_sitter_language_pack import process, ProcessConfig
config = ProcessConfig("python", chunk_max_size=2500)
result = process(source_code, config)
for chunk in result.chunks: # ProcessResult is an object, not a dict
send_to_llm(chunk.content)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 81 lines · 59 tokens per session scan A 59db72751343
chunking-for-llms is a skill published in the GitHub repository xberg-io/tree-sitter-language-pack (464 stars, last pushed today), licensed MIT. It adds 59 tokens to every session and 692 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
crawlberg
Crawl, scrape, and convert websites to Markdown using the local crawlberg CLI and its MCP server. Use when the user wants to fetch a page, follow links across a domain, enumerate URLs, or drive a real browser. Covers installation, the subcommands (scrape, crawl, map, interact, batch-scrape, batch-crawl, download…
automating-the-browser
Use when extracting a page needs scripted interaction first — click, type, press a key, scroll, wait, screenshot, or run JS before capturing the DOM. Covers crawlberg interact URL --actions with the real action schema, result shape, limits, and external-CDP options.
crawling-a-site
Use when the user wants to follow links across a domain and capture every reachable page as Markdown. Covers crawlberg crawl with depth, page caps, concurrency, rate limiting, domain scoping, robots, and output selection.
headless-fallback
Use when a static fetch returns nothing useful and the page needs a real browser. Covers --browser-mode auto|always|never, external CDP via --browser-endpoint, symptoms of JS-only pages and WAF blocks, and the performance cost.
scraping-html-to-markdown
Use when the user wants a single page rendered as clean Markdown plus structured metadata. Covers crawlberg scrape URL, JSON vs Markdown output, what metadata is returned, and how to handle JS-heavy pages.
serving-the-api
Use when the user wants a long-running HTTP service for scrape/crawl/map instead of one-shot CLI calls or the MCP server — for example wiring crawlberg into other apps over REST. Covers crawlberg serve, the Firecrawl-v1-compatible endpoints, --host/--port, and when to prefer it.