KernelFlow-ops/cuda-optimized-skill
Skill Claude CodeCodex
Iteratively optimize a CUDA/CUTLASS/Triton kernel against a reference implementation using ncu-guided reasoning. Use this skill whenever the user asks to optimize, speed up, or improve the performance of a .cu kernel (CUDA or CUTLASS) or a Triton/Python kernel file, especially when they provide a baseline operator and…
202 4mo ago A 198 tokens
original MIT