Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

FlashAttention-4 (CuTeDSL)

FlashAttention-4 is a CuTeDSL-based implementation of FlashAttention for Hopper and Blackwell GPUs.

Installation

pip install flash-attn-4

If you're on CUDA 13, install with the cu13 extra for best performance:

pip install "flash-attn-4[cu13]"

Usage

from flash_attn.cute import flash_attn_func, flash_attn_varlen_func

out = flash_attn_func(q, k, v, causal=True)

Development

git clone https://github.com/Dao-AILab/flash-attention.git
cd flash-attention
pip install -e "flash_attn/cute[dev]"       # CUDA 12.x
pip install -e "flash_attn/cute[dev,cu13]"  # CUDA 13.x (e.g. B200)
pytest tests/cute/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

flash_attn_4-4.0.0b25.tar.gz (373.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

flash_attn_4-4.0.0b25-py3-none-any.whl (398.8 kB view details)

Uploaded Python 3

File details

Details for the file flash_attn_4-4.0.0b25.tar.gz.

File metadata

  • Download URL: flash_attn_4-4.0.0b25.tar.gz
  • Upload date:
  • Size: 373.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for flash_attn_4-4.0.0b25.tar.gz
Algorithm Hash digest
SHA256 dacae8f0a3744eef7d896b7d1cb0fcb21817969d8a8eb471a0062cf701fcaa96
MD5 5a71a1032f8a9056e4205f60f04cbea8
BLAKE2b-256 3583c037363e1116ebcb40c9121e2d57a65532a492240d345297f3807af65bd4

See more details on using hashes here.

Provenance

The following attestation bundles were made for flash_attn_4-4.0.0b25.tar.gz:

Publisher: publish-fa4.yml on Dao-AILab/flash-attention

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file flash_attn_4-4.0.0b25-py3-none-any.whl.

File metadata

  • Download URL: flash_attn_4-4.0.0b25-py3-none-any.whl
  • Upload date:
  • Size: 398.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for flash_attn_4-4.0.0b25-py3-none-any.whl
Algorithm Hash digest
SHA256 2e6c00179c5e017bc9db81df410487485229e3be744fd3a0dc87cf3b0d0d76f3
MD5 5fe88ebce666a04e9f8ab89e439f9333
BLAKE2b-256 9c06a9c2723e0f286b5f1b10d6072af3900b04c06e4dee60201c8cc0718eeb27

See more details on using hashes here.

Provenance

The following attestation bundles were made for flash_attn_4-4.0.0b25-py3-none-any.whl:

Publisher: publish-fa4.yml on Dao-AILab/flash-attention

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page