cache-dit

Cache-DiT: A PyTorch-native Inference Engine with Cache, Parallelism, Quantization and CPU Offload for DiTs.

These details have not been verified by PyPI

Project links

Project description

⚡️🎉A PyTorch-native Inference Engine with Cache,
Parallelism, Quantization and CPU Offload for DiTs

🤗Why Cache-DiT❓❓Cache-DiT is built on top of the 🤗Diffusers library and now supports nearly ALL DiTs from Diffusers. It provides hybrid cache acceleration (DBCache, TaylorSeer, SCM, etc.) and comprehensive parallelism optimizations, including Context Parallelism, Tensor Parallelism, hybrid 2D or 3D parallelism, and dedicated extra parallelism support for Text Encoder, VAE, and ControlNet.

Cache-DiT is compatible with compilation, CPU Offloading, and quantization, fully integrates with SGLang Diffusion, vLLM-Omni, TensorRT-LLM, ComfyUI, and runs natively on NVIDIA GPUs, Ascend NPUs and AMD GPUs. Cache-DiT is fast, easy to use, and flexible for various DiTs (online docs at 📘cache-dit.io).

⚡️9x speedup by Cache-DiT with Cache, Context Parallelism and Compilation

🚀Quick Start: Cache, Parallelism and Quantization

First, you can install the cache-dit from PyPI or install from source:

uv pip install -U cache-dit # PyPI, stable release.
uv pip install git+https://github.com/vipshop/cache-dit.git # latest.

Then, try to accelerate your DiTs with just ♥️one line♥️ of code ~

>>> import cache_dit
>>> from diffusers import DiffusionPipeline
>>> pipe = DiffusionPipeline.from_pretrained(...).to("cuda")
>>> cache_dit.enable_cache(pipe) # Cache Acceleration with One-line code.
>>> from cache_dit import DBCacheConfig, ParallelismConfig
>>> cache_dit.enable_cache( # Or, Hybrid Cache Acceleration + Parallelism.
...   pipe, cache_config=DBCacheConfig(), # w/ default
...   parallelism_config=ParallelismConfig(ulysses_size=2))
>>> from cache_dit import DBCacheConfig, ParallelismConfig, QuantizeConfig
>>> cache_dit.enable_cache( # Or, Hybrid Cache + Parallelism + Quantization.
...   pipe, cache_config=DBCacheConfig(), # w/ default
...   parallelism_config=ParallelismConfig(ulysses_size=2),
...   quantize_config=QuantizeConfig(quant_type=...))
>>> output = pipe(...) # Then, just call the pipe as normal.

🚀Quick Start: SVDQuant (W4A4) PTQ/DQ workflow

First, install Cache-DiT with SVDQuant support (Experimental):

# Required: CUDA 13.0+, PyTorch 2.11+, Ubuntu 22.04+.
uv pip install -U cache-dit-cu13 # PyPI, stable release.
CACHE_DIT_BUILD_SVDQUANT=1 uv pip install -e ".[quantization]" # latest.

Then, try to quantize your model with just ♥️a few lines♥️ of code ~

>>> from cache_dit import QuantizeConfig
>>> pipe = DiffusionPipeline.from_pretrained(...).to("cuda")
>>> # DQ: "...{dtype}_r{rank}_dq", PTQ: "...{dtype}_r{rank}"
>>> pipe.transformer = cache_dit.quantize(
...   pipe.transformer, quant_config=QuantizeConfig(
...   # ✅ NO calibration needed for SVDQ DQ in Cache-DiT! 🎉
...   quant_type="svdq_{int4|nvfp4}_r{32|64|128|256|...}_dq",
...   svdq_kwargs={"smooth_strategy": "few_shot"})) 
>>> output = pipe(...) # Then, just call the pipe as normal.

🚀Quick Start: Bucket-style Layerwise CPU Offload

Bucket-style Layerwise Offload w/ nearly zero (<5%🎉) latency overhead ~

>>> import cache_dit
>>> cache_dit.layerwise_offload(
...   pipe, # nn.Module: pipe, transformer, text_encoder, etc.
...   onload_device="cuda",
...   offload_device="cpu",
...   async_transfer=True,
...   transfer_buckets=4,
...   persistent_buckets=64,
...   persistent_bins=8,
...   prefetch_limit=True,
...   max_copy_streams=4,
...   max_inflight_prefetch_bytes="8gib")
>>> output = pipe(...) # Then, just call the pipe as normal.

For more advanced features, please refer to our online documentation at 📘cache-dit.io.

🌐Community Integration

©️Acknowledgements

Special thanks to vipshop's Computer Vision AI Team for supporting testing and deployment of this project. We learned and reused codes from: Diffusers, SGLang, vLLM-Omni, Nunchaku, xDiT and TaylorSeer.

©️Citations

@misc{cache-dit@2025,
  title={Cache-DiT: A PyTorch-native Inference Engine with Cache, Parallelism, Quantization and CPU Offload for DiTs.},
  url={https://github.com/vipshop/cache-dit.git},
  note={Open-source software available at https://github.com/vipshop/cache-dit.git},
  author={DefTruth, vipshop.com, etc.},
  year={2025}
}

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

This version

1.3.12

Jun 9, 2026

1.3.11

Jun 4, 2026

1.3.10

Jun 4, 2026

1.3.9

May 27, 2026

1.3.8

May 25, 2026

1.3.7

May 12, 2026

1.3.6

May 11, 2026

1.3.5

Mar 30, 2026

1.3.4

Mar 27, 2026

1.3.3

Mar 26, 2026

1.3.2

Mar 26, 2026

1.3.1

Mar 25, 2026

1.3.0

Mar 11, 2026

1.2.3

Feb 26, 2026

1.2.2

Feb 10, 2026

1.2.1

Feb 2, 2026

1.2.0

Jan 16, 2026

1.1.10

Dec 31, 2025

1.1.9

Dec 22, 2025

1.1.8

Dec 10, 2025

1.1.7

Dec 6, 2025

1.1.6

Dec 5, 2025

1.1.5

Dec 5, 2025

1.1.4

Nov 28, 2025

1.1.3

Nov 28, 2025

1.1.2

Nov 24, 2025

1.1.1

Nov 19, 2025

1.1.0

Nov 18, 2025

1.0.16

Nov 17, 2025

1.0.15

Nov 13, 2025

1.0.14

Nov 11, 2025

1.0.13

Nov 7, 2025

1.0.12

Nov 7, 2025

1.0.11

Nov 5, 2025

1.0.10

Oct 30, 2025

1.0.9

Oct 24, 2025

1.0.8

Oct 22, 2025

1.0.7

Oct 22, 2025

1.0.6

Oct 20, 2025

1.0.5

Oct 15, 2025

1.0.4

Oct 14, 2025

1.0.3

Oct 12, 2025

1.0.2

Oct 10, 2025

1.0.1

Sep 26, 2025

1.0.0

Sep 25, 2025

0.3.3

Sep 23, 2025

0.3.2

Sep 22, 2025

0.3.1

Sep 19, 2025

0.3.0

Sep 17, 2025

0.2.37

Sep 17, 2025

0.2.36

Sep 16, 2025

0.2.34

Sep 12, 2025

0.2.33

Sep 10, 2025

0.2.32

Sep 8, 2025

0.2.31

Sep 8, 2025

0.2.30

Sep 5, 2025

0.2.29

Sep 4, 2025

0.2.28

Sep 3, 2025

0.2.27

Sep 1, 2025

0.2.26

Aug 29, 2025

0.2.25

Aug 28, 2025

0.2.24

Aug 26, 2025

0.2.23

Aug 25, 2025

0.2.22

Aug 25, 2025

0.2.21

Aug 22, 2025

0.2.20

Aug 21, 2025

0.2.19

Aug 20, 2025

0.2.18

Aug 20, 2025

0.2.17

Aug 19, 2025

0.2.16

Aug 15, 2025

0.2.15

Aug 11, 2025

0.2.14

Aug 5, 2025

0.2.13

Jul 30, 2025

0.2.12

Jul 24, 2025

0.2.11

Jul 21, 2025

0.2.10

Jul 17, 2025

0.2.9

Jul 13, 2025

0.2.8

Jul 11, 2025

0.2.7

Jul 10, 2025

0.2.6

Jul 9, 2025

0.2.5

Jul 9, 2025

0.2.4

Jul 3, 2025

0.2.3

Jul 1, 2025

0.2.2

Jun 30, 2025

0.2.1

Jun 22, 2025

0.2.0

Jun 20, 2025

0.1.8

Jun 20, 2025

0.1.7

Jun 18, 2025

0.1.6

Jun 18, 2025

0.1.5

Jun 18, 2025

0.1.3

Jun 17, 2025

0.1.2

Jun 17, 2025

0.1.1

Jun 17, 2025

0.1.1.dev2 pre-release

Jun 16, 2025

0.1.0

Jun 17, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

cache_dit-1.3.12-py3-none-any.whl (478.1 kB view details)

Uploaded Jun 9, 2026 Python 3

File details

Details for the file cache_dit-1.3.12-py3-none-any.whl.

File metadata

Download URL: cache_dit-1.3.12-py3-none-any.whl
Upload date: Jun 9, 2026
Size: 478.1 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.1.0 CPython/3.12.3

File hashes

Hashes for cache_dit-1.3.12-py3-none-any.whl
Algorithm	Hash digest
SHA256	`75eeab88bc7166fba87660610b02fd8fd38c319be58dee801598f9e27cabcffe`
MD5	`fc81099ab2fb98bb0db390bb40258759`
BLAKE2b-256	`64be55ed889b5b555811e9eba28822d9fdaa8d6be88dcec643341c0b438c2b7c`

See more details on using hashes here.

cache-dit 1.3.12

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Project description

⚡️🎉A PyTorch-native Inference Engine with Cache,
Parallelism, Quantization and CPU Offload for DiTs

🚀Quick Start: Cache, Parallelism and Quantization

🚀Quick Start: SVDQuant (W4A4) PTQ/DQ workflow

🚀Quick Start: Bucket-style Layerwise CPU Offload

🌐Community Integration

©️Acknowledgements

©️Citations

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Release history Release notifications | RSS feed

Download files

Source Distributions

Built Distribution

File details

File metadata

File hashes

cache-dit 1.3.12

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Project description

⚡️🎉A PyTorch-native Inference Engine with Cache, Parallelism, Quantization and CPU Offload for DiTs

🚀Quick Start: Cache, Parallelism and Quantization

🚀Quick Start: SVDQuant (W4A4) PTQ/DQ workflow

🚀Quick Start: Bucket-style Layerwise CPU Offload

🌐Community Integration

©️Acknowledgements

©️Citations

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Release history Release notifications | RSS feed

Download files

Source Distributions

Built Distribution

File details

File metadata

File hashes

⚡️🎉A PyTorch-native Inference Engine with Cache,
Parallelism, Quantization and CPU Offload for DiTs