write-triton-gemm-kernel

write-triton-gemm-kernel is a skill for Claude Code, Codex from tensormux/kernel-skills. It costs 0 tokens per session (3,036 once invoked), scanned A, original, MIT.

A guide for implementing blocked matrix multiplication in Triton, where rectangular pieces of two matrices are multiplied and combined. It also covers adapting the calculation for fused follow-up operations.

In plain words
What is it for?
Use it for custom or batched matrix multiplication kernels, especially when combining multiplication with steps such as adding a bias or applying an activation.
Why use it?
It helps when standard libraries cannot efficiently support a needed data type, layout, or fused operation, while keeping the implementation tunable.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it for custom or batched matrix multiplication kernels, especially when combining multiplication with steps such as adding a bias or applying an activation.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/tensormux/kernel-skills/write-triton-gemm-kernel
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add tensormux/kernel-skills --skill write-triton-gemm-kernel
Clone the repo
git clone --depth 1 https://github.com/tensormux/kernel-skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for write-triton-gemm-kernel

README.md
[![agentmods](https://agentmods.dev/badge/skills/tensormux/kernel-skills/write-triton-gemm-kernel.svg)](https://agentmods.dev/skills/tensormux/kernel-skills/write-triton-gemm-kernel)
Your own site
<a href="https://agentmods.dev/skills/tensormux/kernel-skills/write-triton-gemm-kernel"><img src="https://agentmods.dev/badge/skills/tensormux/kernel-skills/write-triton-gemm-kernel.svg" alt="Measured on agentmods" height="20"></a>
Per session 0 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 3,036 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00000 $0.03036
Opus 5 $0.00000 $0.01518
Sonnet 5 $0.00000 $0.00607
Haiku 4.5 $0.00000 $0.00304

Measured 8d ago against content hash 1cbe3b1c86f0, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-07, from the pricing page.

Security

Grade A, and why

write-triton-gemm-kernel scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/triton/write-triton-gemm-kernel/SKILL.md · 159 lines

How it starts

The opening of the file, as written. The whole thing — 159 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Skill: Write a Triton GEMM Kernel

Purpose

Guide the agent through implementing a correct, performant blocked matrix multiplication kernel in Triton. This covers tile assignment via program_id, pointer arithmetic for A/B/C tiles, accumulation with tl.dot, boundary masking for non-divisible shapes, swizzled tile ordering for L2 reuse, and autotuning for BLOCK_M/BLOCK_N/BLOCK_K/num_stages/num_warps.


Use this when

  • You need a custom GEMM or batched GEMM that fuses an epilogue (bias add, activation, scaling, etc.) that cuBLAS or CUTLASS cannot express without a separate kernel.
  • You need a GEMM on a dtype or layout combination that vendor libraries do not natively support efficiently (e.g., mixed-precision accumulation, custom quantized formats).
  • You are building a research kernel and need full visibility into the tiling strategy.
  • The matmul is not on the hot path and you want a single portable Triton implementation rather than a CUTLASS build dependency.

Do not use this when

  • The operation is a standard fp16/bf16/fp32 GEMM with no epilogue fusion requirements. Use torch.compile, torch.mm, or cuBLAS directly — they will match or beat a hand-written Triton GEMM at most shapes.
  • The required shapes are very small (M or N < 64). cuBLAS handles these with batched or grouped GEMM routines that are difficult to match in Triton.
  • You need int8 or fp8 tensor core throughput with fused dequantization. Prefer CUTLASS or cuDNN unless you have a specific reason to own this kernel.
  • Latency matters more than throughput and the problem is memory-bandwidth-bound at small batch. Profiling should drive this decision — do not assume Triton wins.

Inputs the agent should gather first

Before writing any code, confirm:

  1. M, N, K — exact values or the range of values expected at runtime (static vs dynamic shapes).
  2. Input dtype — fp16, bf16, fp32, or mixed (e.g., bf16 inputs, fp32 accumulation).
  3. Layout of A and B — row-major or column-major. If transposed, clarify whether the caller passes the transpose or the kernel should handle it internally.
  4. Output dtype — same as input or upcast.
  5. Epilogue — plain C = A @ B, or is there a scaling factor alpha, bias addition, activation function, or in-place accumulation into an existing C?
  6. Batch dimension — standard 2D matmul, batched (B, M, K) x (B, K, N), or broadcasted batch?
  7. Hardware target — A100, H100, or other. This determines tensor core eligibility and the optimal pipeline depth.
  8. Whether autotuning is allowed — production kernels that ship with a fixed config need to justify that choice; autotuned kernels need a representative benchmark shape.

Read the full file on GitHub · 159 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 8d ago First seen · 159 lines · 0 tokens per session scan A 1cbe3b1c86f0

Subscribe to this mod's changes

write-triton-gemm-kernel is a skill published in the GitHub repository tensormux/kernel-skills (73 stars, last pushed 2mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 3,036 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

flux-analyzer

Analyse FBA flux distributions to extract biological insights. Covers gene essentiality, phenotypic phase planes, flux sampling, pathway-level aggregation, secretion product prediction, and production of publication- quality figures.

aiming-lab/AutoResearchClaw · 44 tokens

fba-simulator

Run Flux Balance Analysis (FBA) and related constraint-based simulations using COBRApy. Covers standard FBA, parsimonious FBA (pFBA), Flux Variability Analysis (FVA), loopless FBA, gene/reaction knockouts, and carbon source swapping. Outputs flux distributions and CSV files.

aiming-lab/AutoResearchClaw · 69 tokens

gsmm-validator

Validate a COBRApy genome-scale metabolic model for mass/charge balance, stoichiometric consistency, biomass producibility, dead-end metabolites, thermodynamic loops, and GPR rule formatting. Outputs a structured validation report with errors and warnings.

aiming-lab/AutoResearchClaw · 52 tokens

gsmm-builder

Build or load a genome-scale metabolic model (GSMM) using COBRApy. Covers loading from BIGG, constructing minimal models from scratch, setting medium constraints, and exporting validated .json model files.

aiming-lab/AutoResearchClaw · 45 tokens

stat-result-validator

Validate statistical research outputs for formulation quality, method-to- problem alignment, theory presence, experimental evidence, fair comparison, artifact completeness, and final-claim consistency.

aiming-lab/AutoResearchClaw · 36 tokens

statistical-experimental-evaluation

Design and run statistical experiments that test the formal problem, proposed methods, theoretical predictions, baselines, and ablations.

aiming-lab/AutoResearchClaw · 31 tokens