A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
Cache-DiT is a PyTorch inference engine for DiT models, with caching, parallel processing, quantization, and CPU offloading. It is for running diffusion-model pipelines with support for systems and hardware such as Diffusers, SGLang Diffusion, vLLM-Omni, ComfyUI, NVIDIA GPUs, Ascend NPUs, and AMD GPUs.
Latest release v1.5.1 · 1 Sept 2026
These files are vipshop/cache-dit's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
.github/skills/cache-dit-model-integration/SKILL.md A 82 tok .github/skills/cuda-cpp-kernel/SKILL.md A 85 tok .github/skills/cute-dsl-kernel/SKILL.md A 63 tok .github/skills/cutlass-cpp-kernel/SKILL.md A 77 tok .github/skills/operator-migration/SKILL.md A 94 tok .github/skills/ptq-workflow-integration/SKILL.md A 75 tok .github/skills/triton-kernel/SKILL.md A 45 tok