Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add tensormux/kernel-skills --skill write-triton-gemm-kernelgit clone --depth 1 https://github.com/tensormux/kernel-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/tensormux/kernel-skills/write-triton-gemm-kernel)<a href="https://agentmods.dev/skills/tensormux/kernel-skills/write-triton-gemm-kernel"><img src="https://agentmods.dev/badge/skills/tensormux/kernel-skills/write-triton-gemm-kernel.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.03036 |
| Opus 5 | $0.00000 | $0.01518 |
| Sonnet 5 | $0.00000 | $0.00607 |
| Haiku 4.5 | $0.00000 | $0.00304 |
Grade A, and why
write-triton-gemm-kernel scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 159 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Skill: Write a Triton GEMM Kernel
Purpose
Guide the agent through implementing a correct, performant blocked matrix multiplication kernel in Triton. This covers tile assignment via program_id, pointer arithmetic for A/B/C tiles, accumulation with tl.dot, boundary masking for non-divisible shapes, swizzled tile ordering for L2 reuse, and autotuning for BLOCK_M/BLOCK_N/BLOCK_K/num_stages/num_warps.
Use this when
- You need a custom GEMM or batched GEMM that fuses an epilogue (bias add, activation, scaling, etc.) that cuBLAS or CUTLASS cannot express without a separate kernel.
- You need a GEMM on a dtype or layout combination that vendor libraries do not natively support efficiently (e.g., mixed-precision accumulation, custom quantized formats).
- You are building a research kernel and need full visibility into the tiling strategy.
- The matmul is not on the hot path and you want a single portable Triton implementation rather than a CUTLASS build dependency.
Do not use this when
- The operation is a standard fp16/bf16/fp32 GEMM with no epilogue fusion requirements. Use
torch.compile,torch.mm, or cuBLAS directly — they will match or beat a hand-written Triton GEMM at most shapes. - The required shapes are very small (M or N < 64). cuBLAS handles these with batched or grouped GEMM routines that are difficult to match in Triton.
- You need int8 or fp8 tensor core throughput with fused dequantization. Prefer CUTLASS or cuDNN unless you have a specific reason to own this kernel.
- Latency matters more than throughput and the problem is memory-bandwidth-bound at small batch. Profiling should drive this decision — do not assume Triton wins.
Inputs the agent should gather first
Before writing any code, confirm:
- M, N, K — exact values or the range of values expected at runtime (static vs dynamic shapes).
- Input dtype — fp16, bf16, fp32, or mixed (e.g., bf16 inputs, fp32 accumulation).
- Layout of A and B — row-major or column-major. If transposed, clarify whether the caller passes the transpose or the kernel should handle it internally.
- Output dtype — same as input or upcast.
- Epilogue — plain C = A @ B, or is there a scaling factor alpha, bias addition, activation function, or in-place accumulation into an existing C?
- Batch dimension — standard 2D matmul, batched (B, M, K) x (B, K, N), or broadcasted batch?
- Hardware target — A100, H100, or other. This determines tensor core eligibility and the optimal pipeline depth.
- Whether autotuning is allowed — production kernels that ship with a fixed config need to justify that choice; autotuned kernels need a representative benchmark shape.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 159 lines · 0 tokens per session scan A 1cbe3b1c86f0
write-triton-gemm-kernel is a skill published in the GitHub repository tensormux/kernel-skills (73 stars, last pushed 2mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 3,036 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
flux-analyzer
Analyse FBA flux distributions to extract biological insights. Covers gene essentiality, phenotypic phase planes, flux sampling, pathway-level aggregation, secretion product prediction, and production of publication- quality figures.
fba-simulator
Run Flux Balance Analysis (FBA) and related constraint-based simulations using COBRApy. Covers standard FBA, parsimonious FBA (pFBA), Flux Variability Analysis (FVA), loopless FBA, gene/reaction knockouts, and carbon source swapping. Outputs flux distributions and CSV files.
gsmm-validator
Validate a COBRApy genome-scale metabolic model for mass/charge balance, stoichiometric consistency, biomass producibility, dead-end metabolites, thermodynamic loops, and GPR rule formatting. Outputs a structured validation report with errors and warnings.
gsmm-builder
Build or load a genome-scale metabolic model (GSMM) using COBRApy. Covers loading from BIGG, constructing minimal models from scratch, setting medium constraints, and exporting validated .json model files.
stat-result-validator
Validate statistical research outputs for formulation quality, method-to- problem alignment, theory presence, experimental evidence, fair comparison, artifact completeness, and final-claim consistency.
statistical-experimental-evaluation
Design and run statistical experiments that test the formal problem, proposed methods, theoretical predictions, baselines, and ablations.