Skip to main content

ComfyUI-OrbitQuant

ComfyUI custom nodes for inspecting OrbitQuant artifacts and attaching a quantized transformer component to an existing pipeline object.

The quantization implementation lives in the orbitquant Python package. This node pack only validates artifacts, reports metadata, and calls OrbitQuant's component-loading API.

Nodes

Node Purpose
OrbitQuant Inspect Artifact Validate an OrbitQuant artifact directory and return a text summary plus structured metadata.
OrbitQuant Pipeline Component Loader Attach any compatible universal or model-specific OrbitQuant component artifact to a pipeline attribute such as transformer.
OrbitQuant FLUX Loader Attach a FLUX or FLUX.2 transformer artifact and reject non-FLUX policies.
OrbitQuant Z-Image Loader Attach a Z-Image transformer artifact and reject other target policies.
OrbitQuant Wan Loader Attach a Wan transformer artifact and reject other target policies.
OrbitQuant Release Loader Validate any supported multicomponent release and select its allowlisted adapter from comfyui_orbitquant.json.
OrbitQuant Generate Video Run a supported video release with durable intermediate artifacts and a standard ComfyUI video preview.

The same nodes are exposed through the legacy NODE_CLASS_MAPPINGS interface and the modern ComfyUI V3 comfy_entrypoint interface when comfy_api is available.

Install

Install through ComfyUI-Manager, or clone this repository into ComfyUI's custom node directory:

cd ComfyUI/custom_nodes
git clone https://github.com/iamwavecut/ComfyUI-OrbitQuant.git

ComfyUI-Manager installs requirements.txt (the orbitquant package) and then runs install.py, which provisions the optimized native kernel package for the current runtime by downloading the matching prebuilt variant wheel from the OrbitQuant GitHub release. Provisioning is best effort: when no variant matches the runtime, packed runtime modes fall back to OrbitQuant's Triton or dequantized paths and the node pack keeps working.

For a manual clone, install the orbitquant package into the Python environment used by ComfyUI and provision the native kernels explicitly:

python -m pip install "orbitquant>=0.9.2,<0.10"
python -m orbitquant.cli.main kernels-install

For the default optimized runtime_mode="auto_fused" path on CUDA, install OrbitQuant with its kernel runtime extra. This provides the Triton fallback used when no native variant matches:

python -m pip install "orbitquant[hf,kernels]>=0.9.2,<0.10"

If you install this node pack from PyPI, the same kernel runtime dependencies are available through the node pack extra:

python -m pip install "comfyui-orbitquant[kernels]"

For a source checkout, install the package from the local OrbitQuant repository:

python -m pip install -e /path/to/OrbitQuant

For a source checkout with the kernel runtime dependencies:

python -m pip install -e "/path/to/OrbitQuant[kernels]"

Restart ComfyUI after installation.

Usage

Use an OrbitQuant artifact directory produced by the OrbitQuant package or downloaded from Hugging Face.

  1. Load or create the source Diffusers pipeline in your workflow.
  2. Add the matching OrbitQuant loader node.
  3. Set artifact_path to the local artifact directory.
  4. Connect the pipeline object into the loader node.
  5. Keep runtime_mode at auto_fused for optimized packed-weight inference.
  6. Use the returned pipeline object for the downstream generation nodes.

For model-specific loaders, the artifact target_policy is checked before the component is attached:

Loader Accepted target_policy
OrbitQuant FLUX Loader flux, flux2
OrbitQuant Z-Image Loader z_image
OrbitQuant Wan Loader wan

Use OrbitQuant Pipeline Component Loader for artifacts with target_policy="universal" or for future transformer components that do not have a specialized node. This loader validates the artifact schema without restricting the source architecture name.

Runtime Modes

runtime_mode defaults to auto_fused. On supported devices, OrbitQuant will use packed low-bit matmul kernels instead of materializing a full BF16/FP16 weight matrix. activation_kernel_backend defaults to auto; the triton_rocm and triton_xpu backends are experimental in OrbitQuant.

Use runtime_mode="dequant_bf16" only as an explicit compatibility or debug path when packed kernels are not installed in the ComfyUI Python environment.

MiniMax H3 W4A4 video

The generic release nodes consume the Diffusers-native multicomponent release instead of the older single-component artifact layout described below. Download the public model into a local directory using the same environment as ComfyUI:

hf download WaveCut/MiniMax-H3-OrbitQuant-W4A4 \
  --local-dir /models/MiniMax-H3-OrbitQuant-W4A4
python -m pip install "orbitquant[hf,kernels]>=0.9.2,<0.10"
python -m pip install \
  "diffusers @ git+https://github.com/huggingface/diffusers.git@abc5e9bf71fd38f53cd471bc3acaa84bc5ecbfdc" \
  "transformers>=5.13,<6" accelerate av soundfile

On the RunPod ComfyUI image, start ComfyUI with its global allocator and offload layers disabled. The OrbitQuant generator subprocess then owns the bounded memory policy instead of competing with ComfyUI's DynamicVRAM and async-offload hooks:

python main.py --listen 0.0.0.0 --port 8188 \
  --disable-cuda-malloc \
  --disable-dynamic-vram \
  --disable-async-offload

Load the public workflow or build the same graph with OrbitQuant Release Loader and OrbitQuant Generate Video. The public node types are model-agnostic; H3-specific component and execution rules live in the release config and an internal allowlisted adapter, so another model family does not require another pair of nodes.

  1. Set OrbitQuant Release Loader.model_path to the downloaded directory.
  2. Connect its release output to OrbitQuant Generate Video.
  3. For T2VA keep task=t2va. For Ref2VA choose ref2va and set reference_path to a local image.
  4. Use width=608, height=480, and steps=24 for the verified 480p recipe.
  5. Start with inference_profile=balanced; choose another profile only for its documented memory/latency tradeoff.

The release schedule uses 24 sigma points and 23 denoiser forwards. MiniMax H3 requires 5–15 seconds at 24 FPS; num_frames=124 is the shortest verified VAE packing sequence and is therefore the default smoke. The text encoder enters GPU memory layer-by-layer for conditioning and is then moved back to RAM before the selected transformer runs.

Inference profiles

All measurements below use W4A4 native packed kernels, no exact INT8 weight cache, 608×480, 124 frames, 24 sigma points, and source FP32 VAEs. Process peaks include CUDA allocations outside PyTorch's own accounting.

Profile Placement Hardware Task Process peak Denoise / generation
balanced (default) streamed leaf offload, 12 GiB allocator cap RTX PRO 6000 T2VA 6.36 GiB child; 6.90 GiB with idle ComfyUI 46.68 s denoise
speed resident transformer RTX PRO 6000 T2VA 21.14 GiB 46.84 s denoise; 51.10 s generation
minimum_vram low-CPU-memory streamed leaf offload, 8 GiB cap RTX 4090 T2VA 4.07 GiB 154.25 s denoise; 188.70 s generation
speed resident transformer_ref RTX PRO 6000 Ref2VA 24.06 GiB 118.48 s denoise; 155.42 s generation

balanced is the Pareto default on the tested PRO 6000: streamed transfers overlap compute closely enough to match resident denoising while cutting the child's physical peak by about 70%. minimum_vram is the verified absolute minimum endpoint. speed removes transformer transfers when VRAM is available. Native-auto Torch Flash SDPA was the fastest supported attention path on the tested CUDA 13 / SM120 stack; SageAttention2's available binary did not contain SM120 code, and cuDNN was slower. These unsupported branches are not part of the public recipe.

T2VA and Ref2VA both use sequential CUDA text conditioning. Ref2VA also encodes the reference through the untouched source FP32 visual VAE before loading the quantized transformer_ref; with a release runner from OrbitQuant 0.11 the runner keeps the VAEs on the GPU for that encode when the device and the profile's memory cap leave room, and streams them tiled otherwise. Visual decode uses tiled source FP32 VAE offload; the source FP32 audio VAE enters GPU only for its audio stage. Neither VAE is quantized.

The node saves generation logs, per-step checkpoints, and the latent bundle as soon as each exists. Only after denoising succeeds does it decode with the untouched source FP32 VAEs. The output is a standard ComfyUI VIDEO, so the core preview and downstream video nodes work without VideoHelperSuite. Decode also retains a high-quality CRF 1 yuv444p master next to the preview; the published HEVC example is derived from that master at CRF 10.

The proof bundle includes the official ComfyUI workflow-image export with embedded JSON, live /prompt results, CRF 1 and HEVC media, frame timelines, audio spectrum, exact revisions, and machine-readable measurements.

Artifact Requirements

The loader expects the standard OrbitQuant component artifact layout:

artifact/
  README.md
  SHA256SUMS
  model_index.json
  model.safetensors
  quantization_config.json
  orbitquant_manifest.json
  orbitquant_codebooks.safetensors
  orbitquant_rotations.safetensors
  prompts.json
  benchmark/summary.json

OrbitQuant Inspect Artifact validates required files, checksums, tensor shapes, source model metadata, bit settings, runtime mode, target policy, and module counts.

Python API

The node classes can also be called directly from Python when building a custom ComfyUI workflow wrapper.

Inspect an artifact:

from comfyui_orbitquant.nodes import OrbitQuantArtifactInspector

summary, info = OrbitQuantArtifactInspector().inspect(
    "/models/orbitquant/flux1-schnell-w4a4"
)
print(summary)
print(info["target_policy"])

Attach a FLUX-family transformer artifact to an existing pipeline object:

from comfyui_orbitquant.nodes import OrbitQuantFluxLoader

pipeline, info = OrbitQuantFluxLoader().load(
    pipeline,
    "/models/orbitquant/flux1-schnell-w4a4",
    strict=True,
    runtime_mode="auto_fused",
    activation_kernel_backend="auto",
)

The nodes delegate artifact parsing and component loading to OrbitQuant:

from orbitquant.artifacts import OrbitQuantManifest, validate_orbitquant_artifact
from orbitquant.pipeline import load_quantized_pipeline_component

No quantization math or artifact parsing logic is duplicated in this repository.

Metadata

Release files for comfyui-orbitquant 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for comfyui-orbitquant 0.5.0
File Size Uploaded
comfyui_orbitquant-0.5.0.tar.gz 31.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for comfyui-orbitquant 0.5.0
File Interpreter ABI Platform
comfyui_orbitquant-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 31.0 MB

Release files / comfyui_orbitquant-0.5.0.tar.gz

Download URL comfyui_orbitquant-0.5.0.tar.gz
Size 31.0 MB
Tags Source
SHA-256 checksum
How to use checksums
cb5b2d51193a44e0ed64d3d657b57d83986c35ab1d6fbae3a7d35f8872ed26e8
BLAKE2b-256 checksum
How to use checksums
8e3d5200cd22da6bdbb1eba1a94625d33c823fc94acd99401c6b96226872a933
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / comfyui_orbitquant-0.5.0-py3-none-any.whl

Download URL comfyui_orbitquant-0.5.0-py3-none-any.whl
Size 20.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6ef071efdf48d7943e746181203a9ce39a05be42bc98eca539131873e6efcef3
BLAKE2b-256 checksum
How to use checksums
00540c22a465779146386df2e57711ba1dd59b97ca77db8af36137e238306c5c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

0.6.0

2 release files

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page