Skip to main content

fake-flash-attention ⚡️

PyPI version

A drop-in, pure-Python shim for the flash-attn package. It redirects all FlashAttention calls to PyTorch's native scaled_dot_product_attention (SDPA).

[!TIP] Successfully tested with the music generation model Stable Audio 3.

Why is this necessary?

Modern Large Language Models (LLMs) and popular libraries (like Hugging Face Transformers) often have hard-coded dependencies on the flash-attn package. However, the official flash-attn library has strict requirements:

  • NVIDIA GPU only: Requires Turing, Ampere, Ada, or Hopper architectures (e.g., RTX 20/30/40, A100, H100).
  • No Support for Older GPUs: Common GPUs like the NVIDIA T4 (standard in Google Colab) or GTX 10-series cards cannot run official FlashAttention kernels.
  • No CPU Support: Official flash-attn cannot be installed or run in CPU-only environments.
  • Complex Compilation: The build process is heavy and requires specific CUDA toolkit versions.

fake-flash-attention solves this by:

  1. API Parity: It exports the exact same functions (e.g., flash_attn_func) so that libraries don't crash with an ImportError.
  2. Hardware Portability: It leverages PyTorch's scaled_dot_product_attention, which is highly optimized and works on T4, older GPUs, and CPUs.
  3. Instant Setup: It is a pure-Python package with no C++/CUDA compilation required.

Installation

pip install fake-flash-attention

Note: If installing from source:

pip install .

Usage

If a library or script requires flash-attn, install this package. Existing code will work transparently:

from flash_attn import flash_attn_func
import torch

q, k, v = torch.randn(1, 12, 256, 64), torch.randn(1, 12, 256, 64), torch.randn(1, 12, 256, 64)
# This now uses PyTorch SDPA under the hood!
output = flash_attn_func(q, k, v, causal=True)

Supported Features

  • ✅ flash_attn_func
  • ✅ flash_attn_varlen_func
  • ✅ flash_attn_qkvpacked_func / kvpacked
  • ✅ FlashAttention-2 API compatibility
  • ✅ Device-agnostic (CPU, CUDA, MPS)

Metadata

Release files for fake-flash-attention 2.6.3.post2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fake-flash-attention 2.6.3.post2
File Size Uploaded
fake_flash_attention-2.6.3.post2.tar.gz 15.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for fake-flash-attention 2.6.3.post2
File Interpreter ABI Platform
fake_flash_attention-2.6.3.post2-py3-none-any.whl Python 3 none any Details

Total release size: 32.2 kB

Release files / fake_flash_attention-2.6.3.post2.tar.gz

Download URL fake_flash_attention-2.6.3.post2.tar.gz
Size 15.8 kB
Tags Source
SHA-256 checksum
How to use checksums
7d63f169b57ea456e7daa18da9e687b2641c4ef79ab51d5c165df5afece800d2
BLAKE2b-256 checksum
How to use checksums
155081956a78c6f4b87c2721912dfea23401b8c5cc42c6ff8fc46e709d55f710
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release files / fake_flash_attention-2.6.3.post2-py3-none-any.whl

Download URL fake_flash_attention-2.6.3.post2-py3-none-any.whl
Size 16.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5018a5216e2e8b6768b07e057021ceb404d7c0a881367a1bef58c1f1ea4ca594
BLAKE2b-256 checksum
How to use checksums
ffeccbe8eb070c88b2c3057e42f2097f5c430038ed493d813f407aa49ef4959f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release history Release notifications | RSS feed

This release

2.6.3.post2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page