vipshop/cache-dit

A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.

About the project

Cache-DiT is a PyTorch inference engine for DiT models, with caching, parallel processing, quantization, and CPU offloading. It is for running diffusion-model pipelines with support for systems and hardware such as Diffusers, SGLang Diffusion, vLLM-Omni, ComfyUI, NVIDIA GPUs, Ascend NPUs, and AMD GPUs.

Latest release v1.5.1 · 1 Sept 2026

These files are vipshop/cache-dit's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

1,273Stars on the repository
7Files it configures its agents with
Tokens loaded in every session
1Agent configured

Skills