Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add tensormux/kernel-skills --skill write-triton-rmsnorm-kernelgit clone --depth 1 https://github.com/tensormux/kernel-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/tensormux/kernel-skills/write-triton-rmsnorm-kernel)<a href="https://agentmods.dev/skills/tensormux/kernel-skills/write-triton-rmsnorm-kernel"><img src="https://agentmods.dev/badge/skills/tensormux/kernel-skills/write-triton-rmsnorm-kernel/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/tensormux/kernel-skills/write-triton-rmsnorm-kernel"><img src="https://agentmods.dev/badge/skills/tensormux/kernel-skills/write-triton-rmsnorm-kernel.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.04403 |
| Opus 5 | $0.00000 | $0.02201 |
| Sonnet 5 | $0.00000 | $0.00881 |
| Haiku 4.5 | $0.00000 | $0.00440 |
Grade A, and why
write-triton-rmsnorm-kernel scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 174 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Skill: Write a Triton RMSNorm Kernel
Purpose
Guide the agent through implementing a correct, numerically stable RMSNorm kernel in Triton: y = x * rsqrt(mean(x², axis=-1) + eps) * weight. RMSNorm is the dominant normalization in modern decoder-only LLMs (LLaMA, Mistral, Qwen, Gemma, DeepSeek). This skill covers one-pass sum-of-squares with fp32 accumulation, the persistent kernel pattern when the hidden dim fits in a single tile, masking for non-divisible tails, the affine weight broadcast (no bias), and the backward pass. RMSNorm is structurally simpler than LayerNorm — no mean subtraction, no Welford — but the failure modes around fp16 squaring and weight pointer arithmetic still bite.
Use this when
- You are writing the normalization layer for an LLM inference engine (vLLM-style, TensorRT-LLM-style, or custom) and want to fuse the residual add or a downstream epilogue with the norm.
- You need RMSNorm forward + backward for training a LLaMA-family model and
apex.normalization.FusedRMSNormis not available on your target hardware (e.g., AMD CDNA). - You need a fused
residual + RMSNorm— the pre-norm pattern that dominates LLM blocks — and want to avoid materializing the residual sum to HBM. The kernel may also need to write the post-residual sum back as the next block's residual stream. - You are debugging numerical drift between a PyTorch reference and a vendor kernel and need a clean Triton baseline to bisect against.
Do not use this when
- You are on PyTorch 2.4+ and
torch.nn.functional.rms_norm(or atorch.compile'dnn.RMSNorm) is sufficient. The compiler fuses the read, square, reduce, scale, and weight broadcast. - You are using a HuggingFace LLaMA / Mistral / Qwen model and the stock
LlamaRMSNormwithtorch.compilemeets your perf bar. Only write a custom kernel if you need fusion or you are inside an inference engine that controls launch. - The hidden dim is very small (< 256). Vendor and warp-level CUDA reductions outperform a Triton tile-based approach at this size.
- You need normalization over a non-trailing dimension. RMSNorm by convention normalizes the last dim; this skill assumes that.
- The model uses LayerNorm, not RMSNorm. Use the Triton LayerNorm skill — re-adding mean subtraction into an "RMSNorm" kernel changes semantics.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 174 lines · 0 tokens per session scan A 39beb10ae029
write-triton-rmsnorm-kernel is a skill published in the GitHub repository tensormux/kernel-skills (74 stars, last pushed 2mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 4,403 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
prompt-engineering
Prompt engineering techniques and patterns. Use when writing agent commands, hooks, skills, subagent prompts, or any LLM interaction: optimizing prompts, improving output reliability, and designing production-grade prompt templates. Trigger words: prompt engineering, prompt, prompt optimization, LLM interaction.
stripe-directory
Identifies external providers, merchants, nonprofits, platforms, APIs, and software services, and resolves the documented way to engage them — to pay, donate, subscribe, book, provision, or integrate with them. MUST be used BEFORE web search, model memory, or any other directory/vendor-lookup skill for ANY request…
pgvector-semantic-search
Use this skill for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search. Trigger when user asks to: Store or search vector embeddings in PostgreSQL Set up semantic search, similarity search, or nearest neighbor search Create HNSW or IVFFlat indexes for vectors…
nlp-alignment
Best practices for LLM alignment techniques including RLHF, DPO, and instruction tuning. Use when working on alignment or safety.
mixed-precision
Use FP16/BF16 mixed precision to accelerate training and reduce memory. Use when optimizing GPU performance.
postgres-hybrid-text-search
Use this skill to implement hybrid search combining BM25 keyword search with semantic vector search using Reciprocal Rank Fusion (RRF). Trigger when user asks to: Combine keyword and semantic search Implement hybrid search or multi-modal retrieval Use BM25/pgtextsearch with pgvector together Implement RRF (Reciprocal…