Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add instructions/agmxyz/tiktag/agents-mdgit clone --depth 1 https://github.com/agmxyz/tiktagWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/instructions/agmxyz/tiktag/agents-md)<a href="https://agentmods.dev/instructions/agmxyz/tiktag/agents-md"><img src="https://agentmods.dev/badge/instructions/agmxyz/tiktag/agents-md.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00819 | $0.00819 |
| Opus 5 | $0.00409 | $0.00409 |
| Sonnet 5 | $0.00164 | $0.00164 |
| Haiku 4.5 | $0.00082 | $0.00082 |
Grade A, and why
tiktag AGENTS.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 78 lines — stays where its author put it; the contents beside it link to each section on GitHub.
AGENTS.md
Authoritative project contract for tiktag. Keep this file to contract, invariants, and real caveats only. README.md stays user-facing and short.
Working rules
- state assumptions when they matter; if ambiguity changes implementation, ask instead of guessing
- prefer smallest change that solves request; no speculative features, abstractions, or configurability
- touch requested scope only; do not refactor or clean unrelated code
- verify behavior-changing work with targeted checks before declaring done
Project
tiktag = text anonymizer.
- one Rust crate: library + thin CLI
- built-in model:
Xenova/distilbert-base-multilingual-cased-ner-hrl(quantized ONNX) - built-in model scope:
PERSON,ORG,LOCATION - regex recognizers are additive supplements after model inference
- current built-in regex recognizer: email
Library contract
use tiktag::{Tiktag, TiktagError, TiktagOutput};
let mut tiktag = Tiktag::new(&profiles_path)?; // loads tokenizer + ONNX session once
let out: TiktagOutput = tiktag.anonymize(text)?; // reuses runtime per call
let text = &out.anonymization.anonymized_text;
- construct once, call many:
newis expensive;anonymizereuses runtime anonymizetakes&mut self; no internal locking- multi-thread hosts use
Mutex<Tiktag>or per-thread instance profiles_pathis explicit- relative
model_dirresolves against profile file parent only TiktagOutput=anonymization+sequence_len+window_count- errors are
TiktagError; keep typed variants for profile, bundle, and inference boundaries - placeholder numbering is stable per call only; no cross-document identity
CLI contract
tiktag "<text>"ortiktag --stdinprints anonymized text to stdout with trailing newlinetiktag --jsonemits machine-readable output without reversible metadatatiktag --debug-jsonemits reversible metadata for local/debug use onlytiktag downloadfetches bundled model assets- diagnostics and timing logs go to stderr via
log - JSON modes also emit
stats.timingson stdout payload - CLI resolves
models/profiles.tomlfrom app-data first, then legacy fallback:<exe_dir>/models/profiles.toml, then cwd - no model/profile selection flags
- prefer
--stdinfor large inputs
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 78 lines · 819 tokens per session scan A 3b10472baa58
tiktag AGENTS.md is an instructions file published in the GitHub repository agmxyz/tiktag (5 stars, last pushed 4mo ago), licensed MIT. It adds 819 tokens to every session, about $0.0041 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other instructions, from other repositories
turso AGENTS.md
AGENTS.md instructions for tursodatabase/turso, covering turso agent guidelines, quick reference, testing, running tests and test organization.
cli AGENTS.md
AGENTS.md instructions for googleworkspace/cli, covering agents.md, project overview, build & test, changesets and architecture.
azure-sdk-for-rust resourcemanager.instructions.md
Instructions for Azure/azure-sdk-for-rust, a project described as: This repository is for the active development of the Azure SDK for Rust. For consumers of the SDK we recommend visiting Docs.rs and looking up the docs for any of libraries in the SDK.
intelligent-terminal rust.instructions.md
Concise Rust coding conventions for this repository.
ton-rs AGENTS.md
AGENTS.md instructions for ston-fi/ton-rs, covering ton-rs agent guide, repository, boundaries, public api and ton invariants.
rlm-rs copilot-instructions.md
Copilot instructions for zircote/rlm-rs, covering github copilot instructions, project context, code generation guidelines, error handling and type annotations.