slowlyC

4 mods across 1 repository, 162 stars between them.

cuda-skill

01

slowlyC/agent-gpu-skills

Skill Claude CodeCodex

Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Use for direct CUDA C++ or PTX work, and for framework tasks only when they need NVIDIA ISA, API, architecture, or tool facts. Triggers include inline PTX, WMMA, WGMMA, TMA…

162 24d ago A 127 tokens original MIT

cutlass-skill

02

slowlyC/agent-gpu-skills

Skill Claude CodeCodex

Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. Use when the task explicitly involves CUTLASS/CuTe/CuTeDSL, cute::Layout, cute::Tensor, TiledMMA, TiledCopy, CollectiveBuilder, CollectiveMainloop, CollectiveEpilogue, GemmUniversal, KernelSchedule…

162 24d ago A 137 tokens original MIT

tilelang-skill

03

slowlyC/agent-gpu-skills

Skill Claude CodeCodex

Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. Use when the task explicitly involves tilelang, tilelang.language, @tilelang.jit, @T.primfunc, T.Kernel, T.copy, T.gemm, TileLang Profiler, Carver, TileLang passes, or TileLang…

162 24d ago A 148 tokens original MIT

triton-skill

04

slowlyC/agent-gpu-skills

Skill Claude CodeCodex

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Use when the task explicitly involves triton.jit, triton.language, tl., Gluon, TensorDescriptor, Triton autotune, TritonGPU/MLIR lowering, tritonkernels, or converting a…

162 24d ago A 116 tokens original MIT