Skip to main content

AI Model Dynamic Offloader

This project is a pytorch VRAM allocator that implements on-demand offloading of model weights when the primary pytorch VRAM allocator comes under pressure.

Support:

  • Nvidia GPUs (CUDA) and AMD GPUs (ROCm/HIP)
  • PyTorch 2.8+
  • CUDA 12.8+ (Nvidia) / ROCm 7+ (AMD)
  • Windows 11+ / Linux as per python ManyLinux support

How it works:

  • The pytorch application creates a Virtual Base Address Register (VBAR) for a model. Creating a VBAR doesn't cost any VRAM, only GPU virtual address space (which is pretty much free).
  • The pytorch application allocates tensors for model weights within the VBAR. These tensors are initially un-allocated and will segfault if touched.
  • The pytorch application faults in the tensors using the fault() API at the time the tensor is needed. This is where VRAM actually gets allocated.
If the fault() is successful (sufficient VRAM for this tensor):
  1. If the fault() resultant signature is changed or unknown:
    • The application uses tensor::_copy() to populate the weight data on the GPU.
    • The application saves the returned signature against this weight for future comparison
  2. The layer uses the weight tensor.
  3. The application calls unpin() on the tensor to allow it to be freed under pressure later if needed.
If the fault() is unsuccessful (offloaded weight):
  1. The application allocates a temporary regular GPU tensor.
  2. Uses _copy to populate weight data on the GPU.
  3. The layer uses the temporary as the weight.
  4. Pytorch garbage collects the temp when the layer is finished.

see examples/example.py


Priorities:

  • The most recent VBARs are the highest priority and lower addresses in the VBAR take priority over higher addresses.
  • Applications should order their tensor allocations in the VBAR in load-priority order with the lowest addresses for the highest priority weights.
  • Calling fault() on a weight that is higher priority than other weights will cause those lower priority weights to get freed to make space.
  • Having a weight evicted sets that VBAR's watermark to that weight's level. Any weights in the same VBAR above the watermark automatically fail the fault() API. This avoids constantly faulting in all weights each model iteration while allowing the application to just blindly call fault() every layer and check the results. There is no need for the application to manage any VRAM quotas or watermarks.
  • Existing VBARs can be pushed to top priority with the prioritize() API. This allows use of an already loaded or partially model (e.g. using the same model twice in a complex workflow). Using prioritize resets the offload watermark of that model to no offloading, giving its weights priority over any other currently loaded models.

Backend:

  • VBAR allocation is done with cuMemAddressReserve(), faulting with cuMemCreate() and cuMemMap() and all frees done with appropriate converse APIs.
  • For consistency with VBAR memory management, main pytorch allocator plugin is also implemented with cuMemAddressReserve -> cuMemCreate -> cuMemMap. This also behaves a lot better on Windows systems with System Memory fallback.

On AMD, the equivalent HIP APIs (hipMemAddressReserve -> hipMemCreate -> hipMemMap, and their converse calls) are used throughout via the same flow.

Caveats:

  • There is no real way for this allocator to tell the difference between high usage and bad fragmentation in the pytorch caching allocator. As we always return success to the pytorch caching allocator it experiences no pressure while weights are being offloaded which means it can run in an extremely fragmented mode. The assumption is model weight access patterns are reasonably regular over blocks or iterations and it finds a good set of sizes to cache. What you should generally do though, is completely flush the pytorch caching allocator before each new model run, which avoids completely un-used reservations from taking priority over the next models weights.

Metadata

Release files for comfy-aimdo 0.5.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for comfy-aimdo 0.5.4
File
comfy_aimdo-0.5.4-py3-none-any.whl Python 3 none any Details
comfy_aimdo-0.5.4-cp39-abi3-win_arm64.whl CPython 3.9 abi3 Windows ARM64 Details
comfy_aimdo-0.5.4-cp39-abi3-win_amd64.whl CPython 3.9 abi3 Windows x86-64 Details
comfy_aimdo-0.5.4-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl CPython 3.9 abi3 Linux glibc 2.17+ x86-64 Details
comfy_aimdo-0.5.4-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl CPython 3.9 abi3 Linux glibc 2.17+ ARM64 Details

Total release size: 1.7 MB

Release files / comfy_aimdo-0.5.4-py3-none-any.whl

Download URL comfy_aimdo-0.5.4-py3-none-any.whl
Size 25.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2f2ad79cabf56742f727709d1f4f4190ddeeed10233bee853793efc5c7a46365
BLAKE2b-256 checksum
How to use checksums
b90ae6c2f358f54c0ce8007b96348812ee04244b3e36e03c614ec9a6d522804f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release files / comfy_aimdo-0.5.4-cp39-abi3-win_arm64.whl

Download URL comfy_aimdo-0.5.4-cp39-abi3-win_arm64.whl
Size 255.0 kB
Tags CPython 3.9 Windows ARM64 abi3
SHA-256 checksum
How to use checksums
5019a1e117d758c206450421586e6a09aad091565c1984e2e559972816a4a9d4
BLAKE2b-256 checksum
How to use checksums
6d7c08b4fd8ba8e0cfb79f3ca94042c3588f4bbd915d52d75c5dba9bc9ad57d1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release files / comfy_aimdo-0.5.4-cp39-abi3-win_amd64.whl

Download URL comfy_aimdo-0.5.4-cp39-abi3-win_amd64.whl
Size 280.9 kB
Tags CPython 3.9 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
e53c9f374d2319b1e2447c79c3189ac77676c75c9dc7f8f98be988196267948b
BLAKE2b-256 checksum
How to use checksums
3a3d4a74fb38aba0d2d6aba4c2a6ea04ce17275c3617f544eea59a1b537a08ec
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release files / comfy_aimdo-0.5.4-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl

Download URL comfy_aimdo-0.5.4-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl
Size 432.2 kB
Tags CPython 3.9 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
65051ce08cee4ed6d8cf68d2340ddb339300975d2781d128a694c534b03e4236
BLAKE2b-256 checksum
How to use checksums
1b7525fd2d0aecdb8e01a67a1dc93250e8031b9b1e71fc28237438faa240b626
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release files / comfy_aimdo-0.5.4-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl

Download URL comfy_aimdo-0.5.4-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl
Size 659.2 kB
Tags CPython 3.9 Linux glibc 2.17+ ARM64 abi3
SHA-256 checksum
How to use checksums
6c7b645d91baf2bf87cb48b0a80267fbd4689bbb6dfd84192a3d7de6e1101d90
BLAKE2b-256 checksum
How to use checksums
5ed6b62e59dbbd1dc8dd471964f1fa51ecc04143db1303678f14bada909c87ce
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.5

5 release files

This release

0.5.4 This release

5 release files

0.5.3

5 release files

0.5.2

5 release files

0.5.1

5 release files

0.5.0

5 release files

0.4.15

5 release files

0.4.14

5 release files

0.4.10

5 release files

0.4.9

5 release files

0.4.8

5 release files

0.4.7

5 release files

0.4.6

5 release files

0.4.5

5 release files

0.4.4

5 release files

0.4.3

5 release files

0.4.2

5 release files

0.4.1

5 release files

0.4.0

5 release files

0.3.0

5 release files

0.2.14

3 release files

0.2.12

3 release files

0.2.11

3 release files

0.2.10

3 release files

0.2.9

3 release files

0.2.8

3 release files

0.2.7

3 release files

0.2.6

3 release files

0.2.5

3 release files

0.2.4

3 release files

0.2.3

3 release files

0.2.2

3 release files

0.2.1

3 release files

0.2.0

3 release files

0.1.8

3 release files

0.1.7

3 release files

0.1.6

3 release files

0.1.5

3 release files

0.1.4

3 release files

0.1.3

3 release files

0.1.2

11 release files

0.1.1

11 release files

0.1.0

11 release files

0.0.4

11 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page