AI Model Dynamic Offloader
This project is a pytorch VRAM allocator that implements on-demand offloading of model weights when the primary pytorch VRAM allocator comes under pressure.
Support:
- Nvidia GPUs (CUDA) and AMD GPUs (ROCm/HIP)
- PyTorch 2.8+
- CUDA 12.8+ (Nvidia) / ROCm 7+ (AMD)
- Windows 11+ / Linux as per python ManyLinux support
How it works:
- The pytorch application creates a Virtual Base Address Register (VBAR) for a model. Creating a VBAR doesn't cost any VRAM, only GPU virtual address space (which is pretty much free).
- The pytorch application allocates tensors for model weights within the VBAR. These tensors are initially un-allocated and will segfault if touched.
- The pytorch application faults in the tensors using the
fault()API at the time the tensor is needed. This is where VRAM actually gets allocated.
If the fault() is successful (sufficient VRAM for this tensor):
- If the fault() resultant signature is changed or unknown:
- The application uses
tensor::_copy()to populate the weight data on the GPU. - The application saves the returned signature against this weight for future comparison
- The application uses
- The layer uses the weight tensor.
- The application calls
unpin()on the tensor to allow it to be freed under pressure later if needed.
If the fault() is unsuccessful (offloaded weight):
- The application allocates a temporary regular GPU tensor.
- Uses
_copyto populate weight data on the GPU. - The layer uses the temporary as the weight.
- Pytorch garbage collects the temp when the layer is finished.
see examples/example.py
Priorities:
- The most recent VBARs are the highest priority and lower addresses in the VBAR take priority over higher addresses.
- Applications should order their tensor allocations in the VBAR in load-priority order with the lowest addresses for the highest priority weights.
- Calling
fault()on a weight that is higher priority than other weights will cause those lower priority weights to get freed to make space. - Having a weight evicted sets that VBAR's watermark to that weight's level. Any weights in the same VBAR above the watermark automatically fail the
fault()API. This avoids constantly faulting in all weights each model iteration while allowing the application to just blindly callfault()every layer and check the results. There is no need for the application to manage any VRAM quotas or watermarks. - Existing VBARs can be pushed to top priority with the
prioritize()API. This allows use of an already loaded or partially model (e.g. using the same model twice in a complex workflow). Usingprioritizeresets the offload watermark of that model to no offloading, giving its weights priority over any other currently loaded models.
Backend:
- VBAR allocation is done with
cuMemAddressReserve(), faulting withcuMemCreate()andcuMemMap()and all frees done with appropriate converse APIs. - For consistency with VBAR memory management, main pytorch allocator plugin is also implemented with
cuMemAddressReserve->cuMemCreate->cuMemMap. This also behaves a lot better on Windows systems with System Memory fallback.
On AMD, the equivalent HIP APIs (hipMemAddressReserve -> hipMemCreate -> hipMemMap, and their converse calls) are used throughout via the same flow.
Caveats:
- There is no real way for this allocator to tell the difference between high usage and bad fragmentation in the pytorch caching allocator. As we always return success to the pytorch caching allocator it experiences no pressure while weights are being offloaded which means it can run in an extremely fragmented mode. The assumption is model weight access patterns are reasonably regular over blocks or iterations and it finds a good set of sizes to cache. What you should generally do though, is completely flush the pytorch caching allocator before each new model run, which avoids completely un-used reservations from taking priority over the next models weights.
Metadata
Release files for comfy-aimdo 0.5.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| comfy_aimdo-0.5.5-py3-none-any.whl | Python 3 | none | any | Details |
| comfy_aimdo-0.5.5-cp39-abi3-win_arm64.whl | CPython 3.9 | abi3 | Windows ARM64 | Details |
| comfy_aimdo-0.5.5-cp39-abi3-win_amd64.whl | CPython 3.9 | abi3 | Windows x86-64 | Details |
| comfy_aimdo-0.5.5-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl | CPython 3.9 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| comfy_aimdo-0.5.5-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl | CPython 3.9 | abi3 | Linux glibc 2.17+ ARM64 | Details |
Total release size: 1.7 MB
Release files / comfy_aimdo-0.5.5-py3-none-any.whl
| Download URL | comfy_aimdo-0.5.5-py3-none-any.whl |
|---|---|
| Size | 25.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
208b4379a080d6a78db49e6236751642b8e9e098ee270b035a1740ec21839fbd
|
|
BLAKE2b-256 checksum How to use checksums |
be523ae1892775f0138af1c0cc27341cf073e0890731f0c68848e8831738f9c2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / comfy_aimdo-0.5.5-cp39-abi3-win_arm64.whl
| Download URL | comfy_aimdo-0.5.5-cp39-abi3-win_arm64.whl |
|---|---|
| Size | 255.1 kB |
| Tags | CPython 3.9 Windows ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
5f1ae34b2e9ba6b5f810246602a1e9ee00e2cd5145745bd2197408ebd0b04d9a
|
|
BLAKE2b-256 checksum How to use checksums |
841776580317af4cd592fe01b52bdf6a5c1d68d76d23612153c4c6f12db2bf3f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / comfy_aimdo-0.5.5-cp39-abi3-win_amd64.whl
| Download URL | comfy_aimdo-0.5.5-cp39-abi3-win_amd64.whl |
|---|---|
| Size | 280.8 kB |
| Tags | CPython 3.9 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
a86906b230d067f7b7390afd7783742777d56df4ad52790c137d8a0236ab8524
|
|
BLAKE2b-256 checksum How to use checksums |
cac483483278fefcb56d76b21c3ceab92e63426ba4a9c49227685d8049524819
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / comfy_aimdo-0.5.5-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl
| Download URL | comfy_aimdo-0.5.5-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl |
|---|---|
| Size | 432.7 kB |
| Tags | CPython 3.9 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
7a4dc76831273a2f837f67cf49c9e507e778b19aff17e6ef8a34a6e9694fd86a
|
|
BLAKE2b-256 checksum How to use checksums |
5d18807dd84d80469c9620928429911b9ff04c699e8b47204423a8804ac3f09d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / comfy_aimdo-0.5.5-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl
| Download URL | comfy_aimdo-0.5.5-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl |
|---|---|
| Size | 659.4 kB |
| Tags | CPython 3.9 Linux glibc 2.17+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
22b66419f9f1887fd9fb49f0e5ff9a88b7a8c38f6b94729e5c2e8bf936ff2d70
|
|
BLAKE2b-256 checksum How to use checksums |
8c53751aae9e7635b3921647579fdca2d7f29c5bc3e8081743fef9dea1fdf36d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency log