vipshop

7 mods across 1 repository, 1.3k stars between them.

vipshop/cache-dit

Skill Claude CodeCodex

High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism (VAE-P), generate CLI, installation, testing workflow, and detailed references. Use when adding support for a new diffusion…

1.3k 4d ago A 82 tokens original Apache-2.0

cuda-cpp-kernel

02

vipshop/cache-dit

Skill Claude CodeCodex

Use when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems or Nsight Compute; or reasoning about Tensor Core instructions, shared memory, bank conflicts, occupancy, async copy, TMA, WGMMA, and…

1.3k 4d ago A 85 tokens original Apache-2.0

cute-dsl-kernel

03

vipshop/cache-dit

Skill Claude CodeCodex

Use when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or rewriting an existing CUDA or C++ operator into CuTe DSL while preserving correctness and performance expectations.

1.3k 4d ago A 63 tokens original Apache-2.0

cutlass-cpp-kernel

04

vipshop/cache-dit

Skill Claude CodeCodex

Use when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM schedules, or CuTe headers; or analyzing template configuration, tiling, memory movement, and kernel structure for Hopper or Blackwell GPUs.

1.3k 4d ago A 77 tokens original Apache-2.0

operator-migration

05

vipshop/cache-dit

Skill Claude CodeCodex

Use when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing operator registration and public wrappers; wiring build and packaging for optional extensions; or reviewing an operator migration plan. Guides…

1.3k 4d ago A 94 tokens original Apache-2.0

vipshop/cache-dit

Skill Claude CodeCodex

Use when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression tests, or reviewing a PTQ integration plan. Uses the SVDQ PTQ integration only as a style and coverage reference. Do not copy the SVDQ…

1.3k 4d ago A 75 tokens original Apache-2.0

triton-kernel

07

vipshop/cache-dit

Skill Claude CodeCodex

Write optimized Triton GPU kernels for deep learning operations. Covers the full spectrum from basic vector ops to Flash Attention, persistent matmul, fused normalization, quantized GEMM, and memory-efficient patterns.

1.3k 4d ago A 45 tokens original Apache-2.0