Skip to main content

Comfy Kitchen

Fast kernel library for Diffusion inference with multiple compute backends.

Backend Capabilities Matrix

Function eager cuda triton hip
quantize_per_tensor_fp8
dequantize_per_tensor_fp8
stochastic_rounding_fp8
quantize_nvfp4
dequantize_nvfp4
scaled_mm_nvfp4
quantize_mxfp8
dequantize_mxfp8
scaled_mm_mxfp8
adaln
rms_adaln
na3d
na2d
sol_attn
int8_attention
apply_rope
apply_rope1
apply_rope_split_half
apply_rope_split_half1
rms_rope
rms_rope1
rms_rope_split_half
rms_rope_split_half1
quantize_int8_rowwise
quantize_int8_tensorwise
quantize_and_rotate_rowwise
quantize_int8_convrot_weight
dequantize_int8_simple_dtype
dequantize_int8_convrot_weight_dtype
int8_linear
gemv_awq_w4a16
quantize_svdquant_w4a4
scaled_mm_svdquant_w4a4
convrot_w4a4_linear
quantize_convrot_w4a4_weight
dequantize_convrot_w4a4_weight

Each of the eight rope entries also has an in-place form (apply_rope_, rms_rope_split_half1_, ...) with the same backend coverage as the row above.

HIP backend (AMD RDNA2 / RDNA3 / RDNA3.5 / RDNA4)

The hip backend implements the quantized paths with its own kernels: WMMA matrix-core GEMMs on RDNA3/RDNA3.5/RDNA4, and non-WMMA kernels (quantizers, INT8 dequantizers, RoPE and the fused RMSNorm+RoPE, AdaLN and RMS-AdaLN, the AWQ GEMV) that also run on RDNA2. It does not link or call hipBLAS/hipBLASLt; every matmul is compiled from the sources in comfy_kitchen/backends/hip/.

Both rope kernels address their inputs through the tensor's own strides, so a q/k pair permuted or sliced out of a packed qkv is read where it lies rather than copied contiguous first. The in-place entries (apply_rope_, rms_rope_ and their split-half and single-tensor siblings) rotate a strided view in place for the same reason: each thread owns one element pair and loads both before storing either.

int8_linear(input_act=...) folds the activation into the fused ConvRot quantizer's load, so an MLP's linear(act(proj(x))) never writes act's output to HBM just to read it straight back.

What a GPU gets depends on whether it has matrix cores:

Generation gfx targets Matrix cores What runs
RDNA4 gfx1200, gfx1201 WMMA + fp8 All HIP-supported kernels, fp8 native
RDNA3.5 gfx1150-gfx1153 WMMA, no fp8 All HIP-supported kernels; fp8 widened
RDNA3 gfx1100-gfx1103 WMMA, no fp8 All HIP-supported kernels; fp8 widened
RDNA2 gfx1030-gfx1036 none Non-WMMA kernels incl. AWQ GEMV; WMMA GEMMs decline

fp8, int8 and int4 share one byte-addressed tile kernel (gemm_wmma.h). RDNA3 and RDNA4 spread a WMMA operand across the wave differently and RDNA3 has no fp8 WMMA (it widens to bf16, which is exact), so each has its own set of Mma policies in mma.h; the tile kernel itself is shared.

RDNA2 has no matrix cores. It runs the kernels that do not need them (RoPE, AdaLN and RMS-AdaLN, the quantizers, stochastic rounding, the AWQ GEMV) and does not advertise the GEMMs, which fall through to triton/eager. In a process with a mix of GPUs the capability set is the intersection, since kernels launch on the tensor's own device.

A request outside a kernel's domain (swizzled operands, scaling other than tensor-wise, a K that is not a multiple of 16) falls back to torch or eager. NVFP4 and MXFP8 stay on eager everywhere: RDNA has neither fp4 WMMA nor microscaling hardware. Set COMFY_KITCHEN_DISABLE_HIP=1 to remove the backend from dispatch.

Building

On a ROCm-only host, the backend is selected automatically when CUDA's nvcc is absent. Both a system ROCm install and the pip rocm-sdk layout (which a ROCm PyTorch build already pulls in) are detected, so on Linux and Windows alike the usual build is:

pip install .

No environment variables, CC/CXX override or Visual Studio developer shell are needed: the ROCm clang builds C, C++ and HIP alike and locates the MSVC toolchain itself. CMake >= 3.26 and Ninja are required (Windows only ships a Visual Studio generator, which has no HIP language support). On Windows the Microsoft C++ build tools and Windows SDK must be installed, since clang links against them and CMake compiles a resource file with the SDK's rc.exe, which the build locates itself rather than expecting on PATH. Use the Visual Studio 2022 v143 toolset; newer MSVC toolsets are not yet reliable with ROCm.

When CUDA and ROCm toolchains are both installed, the source build defaults to CUDA only. This avoids compiling an unused multi-architecture HIP binary on an NVIDIA workstation. Request a combined build explicitly:

COMFY_KITCHEN_BUILD_HIP=1 pip install .

Architectures default to the validated GPUs the build machine can see, or to every target in the backend's architecture manifest when it can see none (a CI box), which is what the wheels carry. Detection reads the visible devices through PyTorch, so under PEP 517 build isolation (a plain pip install .) it sees nothing and falls back to the full target list; set COMFY_HIP_ARCHS, or pass --no-build-isolation, to build for the local GPU instead. Building for one target is much faster:

COMFY_HIP_ARCHS=gfx1201 pip install .
$env:COMFY_HIP_ARCHS = "gfx1201"; pip install .

PYTORCH_ROCM_ARCH and GPU_ARCHS are honoured too. When the build machine sees AMD GPUs but none is RDNA2/3/3.5/4 (CDNA has MFMA, not WMMA), the extension is skipped rather than built (seeing no GPU at all falls back to the full target list above instead); COMFY_KITCHEN_BUILD_HIP=1 requests HIP explicitly (and makes an unsupported visible AMD GPU a hard error), while COMFY_KITCHEN_BUILD_NO_HIP=1 suppresses the backend entirely.

Architecture overrides are exact and fail closed. A compiler-recognized target that is not in the manifest is rejected until its device and WMMA policies have been reviewed and added.

Both extensions are built against the Python limited API on 3.12+, so a wheel carrying CUDA and HIP side by side keeps its abi3 tag. At runtime only the extension matching PyTorch's CUDA or ROCm runtime is loaded.

Quantized Tensors

The library provides QuantizedTensor, a torch.Tensor subclass that transparently intercepts PyTorch operations and dispatches them to optimized quantized kernels when available.

Layout Format HW Requirement Description
TensorCoreFP8Layout FP8 E4M3 SM ≥ 8.9 (Ada) Per-tensor scaling, 1:1 element mapping
TensorCoreNVFP4Layout NVFP4 E2M1 SM ≥ 10.0 (Blackwell) Block quantization with 16-element blocks
TensorCoreMXFP8Layout MXFP8 E4M3 SM ≥ 10.0 (Blackwell) Block quantization with 32-element blocks, E8M0 scales
from comfy_kitchen.tensor import QuantizedTensor, TensorCoreFP8Layout, TensorCoreNVFP4Layout

# Quantize a tensor
x = torch.randn(128, 256, device="cuda", dtype=torch.bfloat16)
qt = QuantizedTensor.from_float(x, TensorCoreFP8Layout)

# Operations dispatch to optimized kernels automatically
output = torch.nn.functional.linear(qt, weight_qt)

# Dequantize back to float
dq = qt.dequantize()

Installation

From PyPI

# Install default (Linux/Windows/MacOS)
pip install comfy-kitchen

# Install with CUBLAS for NVFP4 (+Blackwell)
pip install comfy-kitchen[cublas]

Package Variants

  • CUDA wheels: Linux x86_64 and Windows x64
  • Pure Python wheel: Any platform, eager and triton backends only

Wheels are built for Python 3.10, 3.11, and 3.12+ (using Stable ABI for 3.12+).

From Source

# Standard installation with CUDA support
pip install .

# Development installation
pip install -e ".[dev]"

# For faster rebuilds during development (skip build isolation)
pip install -e . --no-build-isolation -v

Build Options

These options require using setup.py directly (not pip install):

Option Command Description Default
--no-cuda python setup.py bdist_wheel --no-cuda Disable CUDA; without --hip, build a CPU-only wheel Enabled (build with CUDA)
--hip python setup.py bdist_wheel --hip Add HIP explicitly (including to a CUDA build) Auto only when CUDA is unavailable
--no-hip python setup.py bdist_wheel --no-hip Disable HIP Disabled
--hip-archs=... python setup.py build_ext --hip-archs="gfx1200;gfx1201" HIP architectures to build for Visible supported AMD GPUs, otherwise all supported targets
--cuda-archs=... python setup.py build_ext --cuda-archs="80;89" CUDA architectures to build for 75-virtual;80;89;90a;100f;120f (Linux), 75-virtual;80;89;120f (Windows)
--debug-build python setup.py build_ext --debug-build Build in debug mode with symbols Disabled (Release)
--lineinfo python setup.py build_ext --lineinfo Enable NVCC line info for profiling Disabled
# Build CPU-only wheel (pure Python, no CUDA required)
python setup.py bdist_wheel --no-cuda

# Build with custom CUDA architectures
python setup.py build_ext --cuda-archs="80;89" bdist_wheel

# Debug build with line info for profiling
python setup.py build_ext --debug-build --lineinfo bdist_wheel

Requirements

  • Python: ≥3.10
  • PyTorch: ≥2.7.0
  • CUDA Runtime (for CUDA wheels): ≥13.0
    • Pre-built wheels require NVIDIA Driver r580+
    • Building from source requires CUDA Toolkit ≥12.8 and CUDA_HOME environment variable
  • nanobind: ≥2.0.0 (for building from source)
  • CMake: ≥3.26 (for building from source; the abi3 modules need FindPython's Development.SABIModule)

Quick Start

import comfy_kitchen as ck
import torch

# Automatic backend selection (hip -> cuda -> triton -> eager)
x = torch.randn(100, 100, device="cuda")
scale = torch.tensor([1.0], device="cuda")
result = ck.quantize_per_tensor_fp8(x, scale)

# Check which backends are available
print(ck.list_backends())

# Force a specific backend
result = ck.quantize_per_tensor_fp8(x, scale, backend="eager")

# Temporarily use a different backend
with ck.use_backend("triton"):
    result = ck.quantize_per_tensor_fp8(x, scale)

Backend System

The library supports multiple backends:

  • eager: Pure PyTorch implementation
  • cuda: Custom CUDA C kernels (CUDA only)
  • hip: Custom HIP kernels (WMMA GEMMs on RDNA3/3.5/4; non-WMMA kernels also on RDNA2)
  • triton: Triton JIT-compiled kernels

Automatic Backend Selection

When you call a function, the registry selects the best backend by checking constraints in priority order (hipcudatritoneager):

# Backend is selected automatically based on input constraints
result = ck.quantize_per_tensor_fp8(x, scale)

# On CPU tensors → falls back to eager (only backend supporting CPU)
# On CUDA tensors → uses cuda or triton (higher priority)

Constraint System

Each backend declares constraints for its functions:

Constraint Description
Device Which device types are supported
Dtype Allowed input/output dtypes per parameter
Shape Shape requirements (e.g., 2D tensors, dimensions divisible by 16)
Compute Capability Minimum GPU architecture (e.g., SM 8.0 for FP8, SM 10.0 for NVFP4)

The registry validates inputs against these constraints before calling the backend—no try/except fallback patterns. If no backend can handle the inputs, a NoCapableBackendError is raised with details.

# Debug logging to see backend selection
import logging
logging.getLogger("comfy_kitchen.dispatch").setLevel(logging.DEBUG)

Testing

Run the test suite with pytest:

# Run all tests
pytest

# Run specific test file
pytest tests/test_backends.py

# Run with verbose output
pytest -v

# Run specific test
pytest tests/test_backends.py::TestBackendSystem::test_list_backends

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

comfy_kitchen-0.2.32-py3-none-any.whl (198.5 kB view details)

Uploaded Python 3

comfy_kitchen-0.2.32-cp312-abi3-win_arm64.whl (9.2 MB view details)

Uploaded CPython 3.12+Windows ARM64

comfy_kitchen-0.2.32-cp312-abi3-win_amd64.whl (48.8 MB view details)

Uploaded CPython 3.12+Windows x86-64

comfy_kitchen-0.2.32-cp312-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (64.9 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

comfy_kitchen-0.2.32-cp312-abi3-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl (48.6 MB view details)

Uploaded CPython 3.12+manylinux: glibc 2.26+ ARM64manylinux: glibc 2.28+ ARM64

comfy_kitchen-0.2.32-cp311-cp311-win_amd64.whl (48.8 MB view details)

Uploaded CPython 3.11Windows x86-64

comfy_kitchen-0.2.32-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (64.9 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

comfy_kitchen-0.2.32-cp311-cp311-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl (48.6 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.26+ ARM64manylinux: glibc 2.28+ ARM64

comfy_kitchen-0.2.32-cp310-cp310-win_amd64.whl (48.8 MB view details)

Uploaded CPython 3.10Windows x86-64

comfy_kitchen-0.2.32-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (64.9 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

comfy_kitchen-0.2.32-cp310-cp310-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl (48.6 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.26+ ARM64manylinux: glibc 2.28+ ARM64

File details

Details for the file comfy_kitchen-0.2.32-py3-none-any.whl.

File metadata

  • Download URL: comfy_kitchen-0.2.32-py3-none-any.whl
  • Upload date:
  • Size: 198.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for comfy_kitchen-0.2.32-py3-none-any.whl
Algorithm Hash digest
SHA256 6a5fba5224abbb7c9d8248bb7fe607bfab26ee623d311fcfae70066f1c7cfd9b
MD5 3d08cb56649f6d586bb341edcc361db9
BLAKE2b-256 a343ceed9307bf92bccdc420703c3800ed46eafcafbfd764cbd93726f43db2b6

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp312-abi3-win_arm64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp312-abi3-win_arm64.whl
Algorithm Hash digest
SHA256 dd7d79ac93b14e8f7b2de9edf05c76443f84b8aeb8d651909c7e3a1d3e1a5454
MD5 9d1227ee89c603b8da36aa5eb05790bc
BLAKE2b-256 b44faecfe8ad470bf94851ac679b259db74e92bd340d9d8a6b3308475e6f735e

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp312-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp312-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 85cc8b5f81956b9f5fc8fc03b3b8e193af8f0d1f145286368448c6477e4bbe42
MD5 1a5488b3def7ef61e1c9797a71925bf0
BLAKE2b-256 9b20205c11c2d391c4599aa59a340f328b5df2ec24db09bfa8f34b3a03490eb8

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp312-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp312-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 e60d87e2f63703dccf8ce76ce0b9118042306388b39a9605c0748e56c0bb85d4
MD5 1770bd5d92ddc8e34871f78e2032448a
BLAKE2b-256 80f983cd0892c30bde91665b267ce8ce5b9f43cf48185cac579abd9d5c831827

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp312-abi3-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp312-abi3-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 4704aef5fffda62dad04f25f5cfb38746c4b0388de7643a44f5877bb28b32aaf
MD5 a917d6e04101909f6c38cfd080781e46
BLAKE2b-256 8d6c3ce80e717855d453601933ba8caca90d4a950567b39d89b345f42cdc917d

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp311-cp311-win_amd64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp311-cp311-win_amd64.whl
Algorithm Hash digest
SHA256 26cb8161b793efd182a1948ea2dc4834211980f61018e2ff89510fb93d158911
MD5 1f555f741446f56b033b8c4464400970
BLAKE2b-256 84a81cc2d17848a1af8a0a37a5ef64e4aafba9d61647d5c3c60d987cd74ab075

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 ed33b1400d8f537ef072af3335aaf73626247ef5c1fa3f33108d38eb37ef3dbd
MD5 7cd2041f4152d8900dcd05ce266f9a43
BLAKE2b-256 0848f0c85a71522b5d0f761e9e5a79899926d842ff48e92f123ebf5f85a3533d

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp311-cp311-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp311-cp311-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 2cc079b77e30bb93135d5baa841a73252eb06e36dea340898dea02df849901dc
MD5 53529baa16bf4b95e78eedb9b20b5f1e
BLAKE2b-256 5894d9eea27be405e5839fa5bbae57eb2efd4ca66948597c6fae639e7d78696d

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp310-cp310-win_amd64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp310-cp310-win_amd64.whl
Algorithm Hash digest
SHA256 a0490fed1cc89a53ccc96b2ee00eefb994d1968f3fd41220ab40ac3f0478d764
MD5 1c0d7edb89ee66dfb534efc8f7f04103
BLAKE2b-256 7ecb5f228727a728d8fa5a47d9a9b85f9541e3295c2caa09a5e49e51a0a726ff

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 88b2cfc0160f45e26c85c3e8e0b9ba3ca5abbfe2dac871e33338c566701fcbf7
MD5 4b503208c140a902ca35d17f332a6994
BLAKE2b-256 e683bde9d2dc076515268cf4444e1c8d6bb9156e85745b57db2c2ec8b5b0ffc4

See more details on using hashes here.

File details

Details for the file comfy_kitchen-0.2.32-cp310-cp310-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for comfy_kitchen-0.2.32-cp310-cp310-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 a30b8507c3366735b69e3023b63fe0c44027fa01553214754eb8273e070429a4
MD5 ac907fe44fcc0150ee3642cffb9f7c09
BLAKE2b-256 41fdb51325ca73e320a5f112aa3c3f7aa1dfe76f7b1f7475afa420fe2d3648d9

See more details on using hashes here.

Release history Release notifications | RSS feed

0.2.33

11 files

This release

0.2.32 This release

11 files

0.2.31

11 files

0.2.30

11 files

0.2.28

11 files

0.2.27

11 files

0.2.26

11 files

0.2.25

11 files

0.2.24

11 files

0.2.23

11 files

0.2.22

10 files

0.2.21

10 files

0.2.20

10 files

0.2.19

10 files

0.2.18

10 files

0.2.17

10 files

0.2.16

10 files

0.2.15

10 files

0.2.14

10 files

0.2.13

10 files

0.2.12

10 files

0.2.11

10 files

0.2.10

10 files

0.2.9

10 files

0.2.8

10 files

0.2.7

7 files

0.2.6

7 files

0.2.5

7 files

0.2.4

7 files

0.2.3

7 files

0.2.2

7 files

0.2.1

7 files

0.2.0

7 files

0.1.6

7 files

0.1.5

7 files

0.1.4

7 files

0.1.3

7 files

0.1.2

7 files

0.1.1

3 files

0.1.0

3 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page