Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

FlashAttention-4 (CuTeDSL)

FlashAttention-4 is a CuTeDSL-based implementation of FlashAttention for Hopper and Blackwell GPUs.

Installation

pip install flash-attn-4

If you're on CUDA 13, install with the cu13 extra for best performance:

pip install "flash-attn-4[cu13]"

Usage

from flash_attn.cute import flash_attn_func, flash_attn_varlen_func

out = flash_attn_func(q, k, v, causal=True)

Development

git clone https://github.com/Dao-AILab/flash-attention.git
cd flash-attention
pip install -e "flash_attn/cute[dev]"       # CUDA 12.x
pip install -e "flash_attn/cute[dev,cu13]"  # CUDA 13.x (e.g. B200)
pytest tests/cute/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

flash_attn_4-4.0.0b29.tar.gz (377.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

flash_attn_4-4.0.0b29-py3-none-any.whl (402.7 kB view details)

Uploaded Python 3

File details

Details for the file flash_attn_4-4.0.0b29.tar.gz.

File metadata

  • Download URL: flash_attn_4-4.0.0b29.tar.gz
  • Upload date:
  • Size: 377.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for flash_attn_4-4.0.0b29.tar.gz
Algorithm Hash digest
SHA256 ad79d2029dbb0c7fb09b92d989caf72cfa207cedb09d58b281caab23bed077bc
MD5 7c2c1ed5d42f065f8216fe62bae3440c
BLAKE2b-256 67c8b8c59af3cc8cb90f22cb323dc79db2c19dbd6e6c0f0514bd7213f907e6a5

See more details on using hashes here.

Provenance

The following attestation bundles were made for flash_attn_4-4.0.0b29.tar.gz:

Publisher: publish-fa4.yml on Dao-AILab/flash-attention

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file flash_attn_4-4.0.0b29-py3-none-any.whl.

File metadata

  • Download URL: flash_attn_4-4.0.0b29-py3-none-any.whl
  • Upload date:
  • Size: 402.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for flash_attn_4-4.0.0b29-py3-none-any.whl
Algorithm Hash digest
SHA256 85dd4b0e98ca38f5681a3214b2f9f230638030e3b868a50f3873336fc98d7537
MD5 7a3688e677ef89d34ac7d3f02aee1398
BLAKE2b-256 74f683590de66aa8389adac280ad6a555b53064d74866d06c1def7e3397e1a82

See more details on using hashes here.

Provenance

The following attestation bundles were made for flash_attn_4-4.0.0b29-py3-none-any.whl:

Publisher: publish-fa4.yml on Dao-AILab/flash-attention

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page