AI Model Dynamic Offloader for ComfyUI

These details have been verified by PyPI

Project links

Repository

Owner

Comfy-Org

GitHub Statistics

Project description

AI Model Dynamic Offloader

This project is a pytorch VRAM allocator that implements on-demand offloading of model weights when the primary pytorch VRAM allocator comes under pressure.

Support:

Nvidia GPUs only
Pytorch 2.8+
Cuda 12.8+
Windows 11+ / Linux as per python ManyLinux support

How it works:

The pytorch application creates a Virtual Base Address Register (VBAR) for a model. Creating a VBAR doesn't cost any VRAM, only GPU virtual address space (which is pretty much free).
The pytorch application allocates tensors for model weights within the VBAR. These tensors are initially un-allocated and will segfault if touched.
The pytorch application faults in the tensors using the fault() API at the time the tensor is needed. This is where VRAM actually gets allocated.

If the `fault()` is successful (sufficient VRAM for this tensor):

If the fault() resultant signature is changed or unknown:
- The application uses tensor::_copy() to populate the weight data on the GPU.
- The application saves the returned signature against this weight for future comparison
The layer uses the weight tensor.
The application calls unpin() on the tensor to allow it to be freed under pressure later if needed.

If the `fault()` is unsuccessful (offloaded weight):

The application allocates a temporary regular GPU tensor.
Uses _copy to populate weight data on the GPU.
The layer uses the temporary as the weight.
Pytorch garbage collects the temp when the layer is finished.

see examples/example.py

Priorities:

The most recent VBARs are the highest priority and lower addresses in the VBAR take priority over higher addresses.
Applications should order their tensor allocations in the VBAR in load-priority order with the lowest addresses for the highest priority weights.
Calling fault() on a weight that is higher priority than other weights will cause those lower priority weights to get freed to make space.
Having a weight evicted sets that VBAR's watermark to that weight's level. Any weights in the same VBAR above the watermark automatically fail the fault() API. This avoids constantly faulting in all weights each model iteration while allowing the application to just blindly call fault() every layer and check the results. There is no need for the application to manage any VRAM quotas or watermarks.
Existing VBARs can be pushed to top priority with the prioritize() API. This allows use of an already loaded or partially model (e.g. using the same model twice in a complex workflow). Using prioritize resets the offload watermark of that model to no offloading, giving its weights priority over any other currently loaded models.

Backend:

VBAR allocation is done with cuMemAddressReserve(), faulting with cuMemCreate() and cuMemMap() and all frees done with appropriate converse APIs.
For consistency with VBAR memory management, main pytorch allocator plugin is also implemented with cuMemAddressReserve -> cuMemCreate -> cuMemMap. This also behaves a lot better on Windows systems with System Memory fallback.

Caveats:

There is no real way for this allocator to tell the difference between high usage and bad fragmentation in the pytorch caching allocator. As we always return success to the pytorch caching allocator it experiences no pressure while weights are being offloaded which means it can run in an extremely fragmented mode. The assumption is model weight access patterns are reasonably regular over blocks or iterations and it finds a good set of sizes to cache. What you should generally do though, is completely flush the pytorch caching allocator before each new model run, which avoids completely un-used reservations from taking priority over the next models weights.

Project details

These details have been verified by PyPI

Project links

Repository

Owner

Comfy-Org

GitHub Statistics

Release history Release notifications | RSS feed

0.2.12

Mar 14, 2026

0.2.11

Mar 13, 2026

0.2.10

Mar 11, 2026

0.2.9

Mar 8, 2026

0.2.8

Mar 5, 2026

This version

0.2.7

Mar 5, 2026

0.2.6

Mar 4, 2026

0.2.5

Mar 4, 2026

0.2.4

Mar 2, 2026

0.2.3

Mar 1, 2026

0.2.2

Feb 25, 2026

0.2.1

Feb 24, 2026

0.2.0

Feb 21, 2026

0.1.8

Feb 9, 2026

0.1.7

Feb 1, 2026

0.1.6

Jan 30, 2026

0.1.5

Jan 28, 2026

0.1.4

Jan 27, 2026

0.1.3

Jan 23, 2026

0.1.2

Jan 15, 2026

0.1.1

Jan 13, 2026

0.1.0

Jan 13, 2026

0.0.213

Apr 16, 2026

0.0.4

Jan 11, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

comfy_aimdo-0.2.7-py3-none-any.whl (18.6 kB view details)

Uploaded Mar 5, 2026 Python 3

comfy_aimdo-0.2.7-cp39-abi3-win_amd64.whl (121.9 kB view details)

Uploaded Mar 5, 2026 CPython 3.9+Windows x86-64

comfy_aimdo-0.2.7-cp39-abi3-manylinux1_x86_64.manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_5_x86_64.whl (80.7 kB view details)

Uploaded Mar 5, 2026 CPython 3.9+manylinux: glibc 2.17+ x86-64manylinux: glibc 2.5+ x86-64

File details

Details for the file comfy_aimdo-0.2.7-py3-none-any.whl.

File metadata

Download URL: comfy_aimdo-0.2.7-py3-none-any.whl
Upload date: Mar 5, 2026
Size: 18.6 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for comfy_aimdo-0.2.7-py3-none-any.whl
Algorithm	Hash digest
SHA256	`f95154f410ed01bc9bc76aae81bbdc907739b2b97755e740b87486628c709b95`
MD5	`ff074b127599e1df2c90d8a93dad1c03`
BLAKE2b-256	`cdb7aeb11a92468f97dd2ed1fb7ac25a54d6de9cd252dda0d1fb7d2f131cac4c`

See more details on using hashes here.

Provenance

The following attestation bundles were made for comfy_aimdo-0.2.7-py3-none-any.whl:

Publisher: build-wheels.yml on Comfy-Org/comfy-aimdo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: comfy_aimdo-0.2.7-py3-none-any.whl
- Subject digest: f95154f410ed01bc9bc76aae81bbdc907739b2b97755e740b87486628c709b95
- Sigstore transparency entry: 1042954585
- Sigstore integration time: Mar 5, 2026
Source repository:
- Permalink: Comfy-Org/comfy-aimdo@6c6a5e9a30308bb5a73bb3fdce14bf2497465f66
- Branch / Tag: refs/tags/v0.2.8
- Owner: https://github.com/Comfy-Org
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: build-wheels.yml@6c6a5e9a30308bb5a73bb3fdce14bf2497465f66
- Trigger Event: push

File details

Details for the file comfy_aimdo-0.2.7-cp39-abi3-win_amd64.whl.

File metadata

Download URL: comfy_aimdo-0.2.7-cp39-abi3-win_amd64.whl
Upload date: Mar 5, 2026
Size: 121.9 kB
Tags: CPython 3.9+, Windows x86-64
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for comfy_aimdo-0.2.7-cp39-abi3-win_amd64.whl
Algorithm	Hash digest
SHA256	`123c07ec7ff7f564b251962f1c8965b5f91ae6ac4f3c85ffb8b53af5605f3cff`
MD5	`f84302e88daaa6dfee6999927e4dcf1a`
BLAKE2b-256	`406f32f04f7334478babdc65e5c7a37100272fd6979d724dc3772816c0c892b7`

See more details on using hashes here.

Provenance

The following attestation bundles were made for comfy_aimdo-0.2.7-cp39-abi3-win_amd64.whl:

Publisher: build-wheels.yml on Comfy-Org/comfy-aimdo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: comfy_aimdo-0.2.7-cp39-abi3-win_amd64.whl
- Subject digest: 123c07ec7ff7f564b251962f1c8965b5f91ae6ac4f3c85ffb8b53af5605f3cff
- Sigstore transparency entry: 1042954588
- Sigstore integration time: Mar 5, 2026
Source repository:
- Permalink: Comfy-Org/comfy-aimdo@6c6a5e9a30308bb5a73bb3fdce14bf2497465f66
- Branch / Tag: refs/tags/v0.2.8
- Owner: https://github.com/Comfy-Org
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: build-wheels.yml@6c6a5e9a30308bb5a73bb3fdce14bf2497465f66
- Trigger Event: push

File details

Details for the file comfy_aimdo-0.2.7-cp39-abi3-manylinux1_x86_64.manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_5_x86_64.whl.

File metadata

Download URL: comfy_aimdo-0.2.7-cp39-abi3-manylinux1_x86_64.manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_5_x86_64.whl
Upload date: Mar 5, 2026
Size: 80.7 kB
Tags: CPython 3.9+, manylinux: glibc 2.17+ x86-64, manylinux: glibc 2.5+ x86-64
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for comfy_aimdo-0.2.7-cp39-abi3-manylinux1_x86_64.manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_5_x86_64.whl
Algorithm	Hash digest
SHA256	`5f514c82c8089a8a92ff1fb8d511f1ca494f84748bb66af0e6feda30e92fa72a`
MD5	`d54cd220390bf19c5f66f6b62c1ef658`
BLAKE2b-256	`f0604a2a3fb982db9498bd1c9c65f5e816574aee9dee3c6b86e730c9024fecb5`

See more details on using hashes here.

Provenance

The following attestation bundles were made for comfy_aimdo-0.2.7-cp39-abi3-manylinux1_x86_64.manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_5_x86_64.whl:

Publisher: build-wheels.yml on Comfy-Org/comfy-aimdo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: comfy_aimdo-0.2.7-cp39-abi3-manylinux1_x86_64.manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_5_x86_64.whl
- Subject digest: 5f514c82c8089a8a92ff1fb8d511f1ca494f84748bb66af0e6feda30e92fa72a
- Sigstore transparency entry: 1042954586
- Sigstore integration time: Mar 5, 2026
Source repository:
- Permalink: Comfy-Org/comfy-aimdo@6c6a5e9a30308bb5a73bb3fdce14bf2497465f66
- Branch / Tag: refs/tags/v0.2.8
- Owner: https://github.com/Comfy-Org
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: build-wheels.yml@6c6a5e9a30308bb5a73bb3fdce14bf2497465f66
- Trigger Event: push

comfy-aimdo 0.2.7

Navigation

Verified details

Project links

Owner

GitHub Statistics

Unverified details

Meta

Project description

AI Model Dynamic Offloader

Support:

How it works:

If the `fault()` is successful (sufficient VRAM for this tensor):

If the `fault()` is unsuccessful (offloaded weight):

Priorities:

Backend:

Caveats:

Project details

Verified details

Project links

Owner

GitHub Statistics

Unverified details

Meta

Release history Release notifications | RSS feed

Download files

Source Distributions

Built Distributions

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance

comfy-aimdo 0.2.7

Navigation

Verified details

Project links

Owner

GitHub Statistics

Unverified details

Meta

Project description

AI Model Dynamic Offloader

Support:

How it works:

If the fault() is successful (sufficient VRAM for this tensor):

If the fault() is unsuccessful (offloaded weight):

Priorities:

Backend:

Caveats:

Project details

Verified details

Project links

Owner

GitHub Statistics

Unverified details

Meta

Release history Release notifications | RSS feed

Download files

Source Distributions

Built Distributions

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance

If the `fault()` is successful (sufficient VRAM for this tensor):

If the `fault()` is unsuccessful (offloaded weight):