triton skills

8 tagged triton, measured the same way as everything else here.

Browse within: cuda 6

kernel-KBS

01

fmh66/kernel-opt-agent

Skill Claude CodeCodex

Corpus-backed GPU kernel knowledge base for CUDA, Triton, CuTe, CUTLASS, and Ampere/Hopper/Blackwell kernel research. Use when the user needs to search merged kernel PR pages, inspect PR diff/provenance artifacts, find KernelWiki synthesis pages, query blog/doc/contest notes, or retrieve evidence-backed implementation…

14 3mo ago A 109 tokens

kernel-benchmark

02

fmh66/kernel-opt-agent

Skill Claude CodeCodex

Standalone kernel benchmarking skill for cuda-cpp, cutlass, cute-dsl, and triton implementations. Use when the user wants to compare a custom CUDA/CUTLASS .cu kernel or CuTe DSL/Triton .py kernel against selectable PyTorch eager, torch.compile, or FlashInfer baselines, validate correctness, measure execution time with…

14 3mo ago A 90 tokens

kernel-loop

03

fmh66/kernel-opt-agent

Skill Claude CodeCodex

Iterative GPU kernel optimization orchestrator for CUDA/CUTLASS/CuTe DSL/Triton kernels. Use for measured, one-change-at-a-time optimization loops with correctness, NCU profiling, KBS evidence, hypothesis discipline, hard iteration gates, final benchmarking, and a traceable report.

14 3mo ago A 63 tokens

troycheng/cuda-kernel-optimizer

Skill Claude CodeCodex

Use when optimizing, tuning, diagnosing, or profiling CUDA, CUTLASS, Triton, PyTorch, vLLM, TensorRT-LLM, or another GPU workload; when assessing an NCU, Nsys, or PyTorch Profiler report; or when the test workload, correctness checks, measurement path, or target environment is incomplete.

7 5d ago A 77 tokens original MIT