3 projects
ffpa-attn
🤖FFPA: Fast and Memory-Efficient Exact Attention for Large Headdim.
cache-dit-cu13
Cache-DiT: A PyTorch-native Inference Engine with Cache, Parallelism, Quantization and CPU Offload for DiTs.
cache-dit
Cache-DiT: A PyTorch-native Inference Engine with Cache, Parallelism, Quantization and CPU Offload for DiTs.