Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

FlashAttention-4 (CuTeDSL)

FlashAttention-4 is a CuTeDSL-based implementation of FlashAttention for Hopper and Blackwell GPUs.

Installation

pip install flash-attn-4

If you're on CUDA 13, install with the cu13 extra for best performance:

pip install "flash-attn-4[cu13]"

Usage

from flash_attn.cute import flash_attn_func, flash_attn_varlen_func

out = flash_attn_func(q, k, v, causal=True)

Development

git clone https://github.com/Dao-AILab/flash-attention.git
cd flash-attention
pip install -e "flash_attn/cute[dev]"       # CUDA 12.x
pip install -e "flash_attn/cute[dev,cu13]"  # CUDA 13.x (e.g. B200)
pytest tests/cute/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

flash_attn_4-4.0.0b23.tar.gz (358.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

flash_attn_4-4.0.0b23-py3-none-any.whl (382.7 kB view details)

Uploaded Python 3

File details

Details for the file flash_attn_4-4.0.0b23.tar.gz.

File metadata

  • Download URL: flash_attn_4-4.0.0b23.tar.gz
  • Upload date:
  • Size: 358.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for flash_attn_4-4.0.0b23.tar.gz
Algorithm Hash digest
SHA256 cb7bed7ac2fec9952a0c58a8d04c64d8cbcd245998d7ea6846c6a0ef21c996fe
MD5 f4c4cf1f8d0bc63873a93b262f23fd06
BLAKE2b-256 1bf7b59668ea11fb527af875bab1a49da31b703106b842726e0fe185a21cd0aa

See more details on using hashes here.

Provenance

The following attestation bundles were made for flash_attn_4-4.0.0b23.tar.gz:

Publisher: publish-fa4.yml on Dao-AILab/flash-attention

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file flash_attn_4-4.0.0b23-py3-none-any.whl.

File metadata

  • Download URL: flash_attn_4-4.0.0b23-py3-none-any.whl
  • Upload date:
  • Size: 382.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for flash_attn_4-4.0.0b23-py3-none-any.whl
Algorithm Hash digest
SHA256 570d5803dc1e2ea8bce6c85ab411c7341efdf0be18a1a720d172f39a52b3dcbd
MD5 14f6ca6f274803c2914c86e6e3cf098d
BLAKE2b-256 6fd0cfc49b44b86f4807cfa59516bbdfd63e5b2c7572e3c63a64d0508157ba84

See more details on using hashes here.

Provenance

The following attestation bundles were made for flash_attn_4-4.0.0b23-py3-none-any.whl:

Publisher: publish-fa4.yml on Dao-AILab/flash-attention

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page