transfusion-pytorch

Transfusion in Pytorch

These details have not been verified by PyPI

Project links

Repository

Project description

Transfusion - Pytorch

Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI.

In this repo, we will substitute diffusion with flow matching given the success of Flux from Black Forest Labs (but will keep the original paper title given Transflow does not have the same ring). This repository will also attempt to extend to any number of modalities.

Appreciation

Pranoy for adding classifier free guidance!

Install

$ pip install transfusion-pytorch

Usage

One modality, say images

from torch import randint, randn
from transfusion_pytorch import Transfusion

model = Transfusion(
    num_text_tokens = 256,
    dim_latent = 384,
    modality_default_shape = (4,),  # fallback, in the case the language model did not produce a valid modality shape
    transformer = dict(
        dim = 512,
        depth = 8
    )
)

# any torch.long is text, torch.float is modalities

text_and_images = [
    [randint(0, 256, (16,)), randn(4, 384), randint(0, 256, (8,)), randn(6, 384)],
    [randint(0, 256, (16,)), randn(7, 384), randint(0, 256, (5,)), randn(2, 384), randint(0, 256, (9,))]
]

loss = model(text_and_images)

loss.backward()

# after much training

one_multimodal_sample = model.sample()

Multiple different modalities

from torch import randint, randn
from transfusion_pytorch import Transfusion

model = Transfusion(
    num_text_tokens = 256,
    dim_latent = (384, 192),                 # specify multiple latent dimensions
    modality_default_shape = ((4,), (2,)),   # default shapes for first and second modality
    transformer = dict(
        dim = 512,
        depth = 8
    )
)

# then for the Tensors of type float, you can pass a tuple[int, Tensor] and specify the modality index in the first position

# any torch.long is text, torch.float is modalities

text_images_and_audio = [
    [randint(0, 256, (16,)), (0, randn(4, 384)), randint(0, 256, (8,)), (1, randn(6, 192))],
    [randint(0, 256, (16,)), randn(7, 384), randint(0, 256, (5,)), (1, randn(2, 192)), randint(0, 256, (9,))]
]

loss = model(text_images_and_audio)

loss.backward()

# after much training

one_multimodal_sample = model.sample()

Automatically taking care of encoding and decoding of images

import torch
from torch import nn, randint, randn
from transfusion_pytorch import Transfusion, print_modality_sample

mock_encoder = nn.Conv2d(3, 384, 3, padding = 1)
mock_decoder = nn.Conv2d(384, 3, 3, padding = 1)

model = Transfusion(
    num_text_tokens = 12,
    dim_latent = 384,
    channel_first_latent = True,
    modality_default_shape = (4, 4),
    modality_encoder = mock_encoder,
    modality_decoder = mock_decoder,
    transformer = dict(
        dim = 512,
        depth = 8
    )
)

text_and_images = [
    [
        randint(0, 12, (16,)),  # 16 text tokens
        randn(3, 8, 8),         # (8 x 8) 3 channeled image
        randint(0, 12, (8,)),   # 8 text tokens
        randn(3, 7, 7)          # (7 x 7) 3 channeled image
    ],
    [
        randint(0, 12, (16,)),  # 16 text tokens
        randn(3, 8, 5),         # (8 x 5) 3 channeled image
        randint(0, 12, (5,)),   # 5 text tokens
        randn(3, 2, 16),        # (2 x 16) 3 channeled image
        randint(0, 12, (9,))    # 9 text tokens
    ]
]

loss = model(text_and_images)

loss.backward()

# after much training

one_multimodal_sample = model.sample()

print_modality_sample(one_multimodal_sample)

To pretrain on language first, just pass in your text as type Int['batch seq']

import torch
from transfusion_pytorch import Transfusion

model = Transfusion(
    num_text_tokens = 256,
    dim_latent = 384,
    transformer = dict(
        dim = 512,
        depth = 8,
    )
).cuda()

text = torch.randint(0, 256, (2, 1024)).cuda()

loss = model(text)
loss.backward()

# after much training

sampled = model.generate_text_only(text[:, :1], 1024)

Examples

To run any of the examples train_{example_name}.py in the project root, simply install dependencies first as so

$ pip install .[examples]

If you run into some weird error with safetensors, run this too

$ pip install -U diffusers transformers accelerate scipy ftfy safetensors

Citations

@inproceedings{Zhou2024TransfusionPT,
    title  = {Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model},
    author = {Chunting Zhou and Lili Yu and Arun Babu and Kushal Tirumala and Michihiro Yasunaga and Leonid Shamis and Jacob Kahn and Xuezhe Ma and Luke Zettlemoyer and Omer Levy},
    year   = {2024},
    url    = {https://api.semanticscholar.org/CorpusID:271909855}
}

@misc{Rubin2024,
    author  = {Ohad Rubin},
    url     = {https://medium.com/@ohadrubin/exploring-weight-decay-in-layer-normalization-challenges-and-a-reparameterization-solution-ad4d12c24950}
}

@article{Nguyen2024MinPS,
    title   = {Min P Sampling: Balancing Creativity and Coherence at High Temperature},
    author  = {Minh Nguyen and Andrew Baker and Andreas Kirsch and Clement Neo},
    journal = {ArXiv},
    year    = {2024},
    volume  = {abs/2407.01082},
    url     = {https://api.semanticscholar.org/CorpusID:270870613}
}

@article{Bao2022AllAW,
    title   = {All are Worth Words: A ViT Backbone for Diffusion Models},
    author  = {Fan Bao and Shen Nie and Kaiwen Xue and Yue Cao and Chongxuan Li and Hang Su and Jun Zhu},
    journal = {2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    year    = {2022},
    pages   = {22669-22679},
    url     = {https://api.semanticscholar.org/CorpusID:253581703}
}

@inproceedings{Zhao2024MonoFormerOT,
    title     = {MonoFormer: One Transformer for Both Diffusion and Autoregression},
    author    = {Chuyang Zhao and Yuxing Song and Wenhao Wang and Haocheng Feng and Errui Ding and Yifan Sun and Xinyan Xiao and Jingdong Wang},
    year      = {2024},
    url       = {https://api.semanticscholar.org/CorpusID:272832492}
}

@article{Yang2024ConsistencyFM,
    title   = {Consistency Flow Matching: Defining Straight Flows with Velocity Consistency},
    author  = {Ling Yang and Zixiang Zhang and Zhilong Zhang and Xingchao Liu and Minkai Xu and Wentao Zhang and Chenlin Meng and Stefano Ermon and Bin Cui},
    journal = {ArXiv},
    year    = {2024},
    volume  = {abs/2407.02398},
    url     = {https://api.semanticscholar.org/CorpusID:270878436}
}

@inproceedings{Zhou2024ValueRL,
    title   = {Value Residual Learning For Alleviating Attention Concentration In Transformers},
    author  = {Zhanchao Zhou and Tianyi Wu and Zhiyun Jiang and Zhenzhong Lan},
    year    = {2024},
    url     = {https://api.semanticscholar.org/CorpusID:273532030}
}

@inproceedings{Duvvuri2024LASERAW,
    title   = {LASER: Attention with Exponential Transformation},
    author  = {Sai Surya Duvvuri and Inderjit S. Dhillon},
    year    = {2024},
    url     = {https://api.semanticscholar.org/CorpusID:273849947}
}

@inproceedings{Dong2024HymbaAH,
    title   = {Hymba: A Hybrid-head Architecture for Small Language Models},
    author  = {Xin Dong and Y. Fu and Shizhe Diao and Wonmin Byeon and Zijia Chen and Ameya Mahabaleshwarkar and Shih-Yang Liu and Matthijs Van Keirsbilck and Min-Hung Chen and Yoshi Suhara and Yingyan Lin and Jan Kautz and Pavlo Molchanov},
    year    = {2024},
    url     = {https://api.semanticscholar.org/CorpusID:274166163}

@article{Zhu2024HyperConnections,
    title   = {Hyper-Connections},
    author  = {Defa Zhu and Hongzhi Huang and Zihao Huang and Yutao Zeng and Yunyao Mao and Banggu Wu and Qiyang Min and Xun Zhou},
    journal = {ArXiv},
    year    = {2024},
    volume  = {abs/2409.19606},
    url     = {https://api.semanticscholar.org/CorpusID:272987528}
}

@article{Zhu2025FracConnectionsFE,
    title   = {Frac-Connections: Fractional Extension of Hyper-Connections},
    author  = {Defa Zhu and Hongzhi Huang and Jundong Zhou and Zihao Huang and Yutao Zeng and Banggu Wu and Qiyang Min and Xun Zhou},
    journal = {ArXiv},
    year    = {2025},
    volume  = {abs/2503.14125},
    url     = {https://api.semanticscholar.org/CorpusID:277104144}
}

@misc{li2025basicsletdenoisinggenerative,
    title   = {Back to Basics: Let Denoising Generative Models Denoise},
    author  = {Tianhong Li and Kaiming He},
    year    = {2025},
    eprint  = {2511.13720},
    archivePrefix = {arXiv},
    primaryClass = {cs.CV},
    url     = {https://arxiv.org/abs/2511.13720},
}

@misc{chefer2026self,
    title   = {Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis},
    author  = {Hila Chefer and Patrick Esser and Dominik Lorenz and Dustin Podell and Vikash Raja and Vinh Tong and Antonio Torralba and Robin Rombach},
    year    = {2026},
    url     = {https://bfl.ai/research/self-flow},
    note    = {Preprint}
}

Project details

These details have not been verified by PyPI

Project links

Repository

Release history Release notifications | RSS feed

This version

0.17.0

Mar 8, 2026

0.16.3

Jan 27, 2026

0.16.2

Jan 23, 2026

0.16.1

Jan 19, 2026

0.16.0

Jan 19, 2026

0.15.3

Jan 8, 2026

0.15.2

Jan 6, 2026

0.15.1

Jan 5, 2026

0.15.0

Jan 5, 2026

0.14.5

Dec 4, 2025

0.14.4

Dec 4, 2025

0.14.2

Dec 4, 2025

0.14.1

Dec 4, 2025

0.14.0

Dec 4, 2025

0.12.0

Oct 13, 2025

0.11.0

Jun 18, 2025

0.10.5

May 13, 2025

0.10.4

May 10, 2025

0.10.3

May 8, 2025

0.10.2

Mar 18, 2025

0.10.1

Feb 1, 2025

0.10.0

Feb 1, 2025

0.9.5

Jan 16, 2025

0.9.4

Jan 11, 2025

0.9.3

Jan 7, 2025

0.9.2

Jan 6, 2025

0.9.1

Jan 4, 2025

0.9.0

Jan 4, 2025

0.8.0

Dec 29, 2024

0.7.2

Dec 26, 2024

0.7.1

Dec 26, 2024

0.7.0

Dec 23, 2024

0.6.7

Dec 4, 2024

0.6.6

Dec 4, 2024

0.6.5

Dec 4, 2024

0.6.4

Dec 3, 2024

0.6.3

Dec 2, 2024

0.6.2

Dec 2, 2024

0.6.1

Dec 2, 2024

0.6.0

Dec 2, 2024

0.5.6

Dec 2, 2024

0.5.5

Nov 26, 2024

0.5.4

Nov 25, 2024

0.5.3

Nov 25, 2024

0.5.2

Nov 25, 2024

0.5.1

Nov 24, 2024

0.5.0

Nov 24, 2024

0.4.16

Nov 23, 2024

0.4.15

Nov 23, 2024

0.4.14

Nov 21, 2024

0.4.12

Nov 21, 2024

0.4.11

Nov 20, 2024

0.4.10

Nov 20, 2024

0.4.9

Nov 20, 2024

0.4.8

Nov 19, 2024

0.4.7

Nov 19, 2024

0.4.6

Nov 19, 2024

0.4.5

Nov 19, 2024

0.4.4

Nov 18, 2024

0.4.3

Nov 18, 2024

0.4.1

Nov 18, 2024

0.4.0

Nov 17, 2024

0.3.15

Nov 17, 2024

0.3.14

Nov 17, 2024

0.3.12

Nov 16, 2024

0.3.11

Nov 16, 2024

0.3.10

Nov 16, 2024

0.3.9

Nov 16, 2024

0.3.8

Nov 16, 2024

0.3.7

Nov 15, 2024

0.3.6

Nov 11, 2024

0.3.5

Nov 10, 2024

0.3.4

Nov 5, 2024

0.3.3

Nov 5, 2024

0.3.2

Nov 4, 2024

0.3.1

Oct 31, 2024

0.3.0

Oct 31, 2024

0.2.3

Oct 16, 2024

0.2.2

Oct 16, 2024

0.2.1

Oct 16, 2024

0.2.0

Oct 16, 2024

0.1.19

Oct 15, 2024

0.1.18

Oct 15, 2024

0.1.17

Oct 15, 2024

0.1.16

Oct 15, 2024

0.1.15

Oct 15, 2024

0.1.14

Oct 15, 2024

0.1.12

Oct 13, 2024

0.1.11

Oct 7, 2024

0.1.10

Oct 7, 2024

0.1.9

Oct 7, 2024

0.1.8

Oct 7, 2024

0.1.7

Oct 7, 2024

0.1.6

Oct 1, 2024

0.1.5

Sep 29, 2024

0.1.4

Sep 29, 2024

0.1.2

Sep 29, 2024

0.1.1

Sep 29, 2024

0.1.0

Sep 28, 2024

0.0.45

Sep 28, 2024

0.0.44

Sep 28, 2024

0.0.43

Sep 26, 2024

0.0.42

Sep 17, 2024

0.0.41

Sep 15, 2024

0.0.40

Sep 15, 2024

0.0.39

Sep 13, 2024

0.0.38

Sep 13, 2024

0.0.37

Sep 11, 2024

0.0.36

Sep 11, 2024

0.0.35

Sep 11, 2024

0.0.34

Sep 10, 2024

0.0.33

Sep 9, 2024

0.0.32

Sep 9, 2024

0.0.31

Sep 9, 2024

0.0.30

Sep 9, 2024

0.0.29

Sep 8, 2024

0.0.28

Sep 8, 2024

0.0.27

Sep 8, 2024

0.0.26

Sep 8, 2024

0.0.25

Sep 8, 2024

0.0.24

Sep 7, 2024

0.0.20

Sep 7, 2024

0.0.19

Sep 7, 2024

0.0.18

Sep 6, 2024

0.0.17

Sep 6, 2024

0.0.16

Sep 6, 2024

0.0.15

Sep 6, 2024

0.0.14

Sep 6, 2024

0.0.12

Sep 6, 2024

0.0.11

Sep 5, 2024

0.0.10

Sep 5, 2024

0.0.9

Sep 5, 2024

0.0.8

Sep 5, 2024

0.0.7

Sep 3, 2024

0.0.6

Sep 3, 2024

0.0.5

Sep 3, 2024

0.0.4

Sep 2, 2024

0.0.3

Sep 2, 2024

0.0.1

Sep 1, 2024

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

transfusion_pytorch-0.17.0.tar.gz (30.1 kB view details)

Uploaded Mar 8, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

transfusion_pytorch-0.17.0-py3-none-any.whl (29.9 kB view details)

Uploaded Mar 8, 2026 Python 3

File details

Details for the file transfusion_pytorch-0.17.0.tar.gz.

File metadata

Download URL: transfusion_pytorch-0.17.0.tar.gz
Upload date: Mar 8, 2026
Size: 30.1 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: uv/0.8.13

File hashes

Hashes for transfusion_pytorch-0.17.0.tar.gz
Algorithm	Hash digest
SHA256	`aa6dc1ac01082000e476203c9daa185713cd7598198ffee0bb101a9c7e413011`
MD5	`c40c4134d1ad10f1828083aff429c25a`
BLAKE2b-256	`d5612f3b26a259462e3897fdaad0c4161bc2a39236fdddb99b965c0a32fe5c67`

See more details on using hashes here.

File details

Details for the file transfusion_pytorch-0.17.0-py3-none-any.whl.

File metadata

Download URL: transfusion_pytorch-0.17.0-py3-none-any.whl
Upload date: Mar 8, 2026
Size: 29.9 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: uv/0.8.13

File hashes

Hashes for transfusion_pytorch-0.17.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`48f1ab5ff9527b74f23643bc1412ccdb0aa024f32c367151bd50d263c3b5d2fa`
MD5	`0d705b427fa5a80b072b254c94cd52d9`
BLAKE2b-256	`803420eb9a0242c478b5ec6de9a25c868c2b8b5196d3082d53f824fea5f18bbf`

See more details on using hashes here.

transfusion-pytorch 0.17.0

Navigation

Verified details

Project links

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Transfusion - Pytorch

Appreciation

Install

Usage

Examples

Citations

Project details

Verified details

Project links

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes