Skip to main content

Humming

Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

Key Features

  • High Flexibility
    • Supports inference for any weight type under 8-bit across FP16 / BF16 / FP8 / FP4 / INT8 / INT4 activations (provided the activation's dynamic range covers the weight type).
    • Supports various quantization strategies.
    • Supports various scale types (BF16, FP16, E4M3, E5M2, and UE8M0).
    • Supports both Dense GEMM and MoE GEMM.
  • High Compatibility: supports all NVIDIA GPUs from SM75+ (Turing architecture) and beyond.
  • High Performance
    • Delivers State-of-the-Art (SOTA) throughput and efficiency across a wide range of computational scenarios.
  • Ultra-Lightweight
    • Minimal dependencies: Requires only PyTorch and NVCC.
    • Compact footprint: The package size is only 100+KB.

Support Matrix

Activation Type Supported Devices Supported Weight Types
FP16 (e5m10) SM75+ • Symmetric INT1-8
• INT1-8 with dynamic zero point
• Arbitrary signed FP (kBits ≤ 8, kExp ≤ 5)
BF16 (e8m7) SM80+ • Symmetric INT1-8
• INT1-8 with dynamic zero point
• Arbitrary signed FP (kBits ≤ 8)
FP8 (e4m3) SM89+ • Symmetric INT1-5
• INT1-4 with dynamic zero point
• Arbitrary signed FP (kExp ≤ 4, kMan ≤ 3)
FP8 (e5m2) SM89+ • Symmetric INT1-4
• INT1-3 with dynamic zero point
• Arbitrary signed FP (kExp ≤ 5, kMan ≤ 2)
FP4 (e2m1) SM120+ • Symmetric INT1-3
• INT1-2 with dynamic zero point
• Arbitrary signed FP (kExp ≤ 2, kMan ≤ 1)
INT8 SM75+ • Symmetric INT1-8
• INT1-7 with dynamic zero point
INT4 SM80+ • Symmetric INT1-4
• INT1-3 with dynamic zero point

Getting Started

Installation

pip install git+https://github.com/inclusionAI/humming.git

Usage Example

import torch
from humming.layer import HummingLayer

layer = HummingLayer(
    shape_n=8192,
    shape_k=8192,
    weight_config={"dtype": "int6"},
    torch_dtype=torch.float16,
).cuda()

weight = torch.randn((8192, 8192), dtype=torch.float16, device="cuda:0")
inputs = torch.randn((128, 8192), dtype=torch.float16, device="cuda:0")

# Load unquantized weight and quantize to layer quantization format
layer.load_from_unquantized(weight)
# Transform weight to humming format and prepare default kernels
layer.transform()

# Run quantized GEMM (tuning_config is optional, auto-selected by default)
output = layer(inputs)

print("Quantized GEMM Output:")
print(output)
print("\nReference Output:")
print(inputs.matmul(weight.T))

Acknowledgement

This project is highly inspired by

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

humming_kernels-0.1.13.tar.gz (280.6 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

humming_kernels-0.1.13-py3-none-manylinux_2_28_x86_64.whl (342.3 kB view details)

Uploaded Python 3manylinux: glibc 2.28+ x86-64

humming_kernels-0.1.13-py3-none-manylinux_2_28_aarch64.whl (337.8 kB view details)

Uploaded Python 3manylinux: glibc 2.28+ ARM64

File details

Details for the file humming_kernels-0.1.13.tar.gz.

File metadata

  • Download URL: humming_kernels-0.1.13.tar.gz
  • Upload date:
  • Size: 280.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for humming_kernels-0.1.13.tar.gz
Algorithm Hash digest
SHA256 03261275b6821add9d23375f6c2616c8a8e0d370a11371a3bb6a1afb138bc815
MD5 ab001fc9b007f21ebfc771bcdf9392ef
BLAKE2b-256 ef14b1e93b4ced758d786c6b580e314dcd49201a982fc12e63473f329b507bdd

See more details on using hashes here.

Provenance

The following attestation bundles were made for humming_kernels-0.1.13.tar.gz:

Publisher: publish.yml on inclusionAI/humming

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file humming_kernels-0.1.13-py3-none-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for humming_kernels-0.1.13-py3-none-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 2ab59b2f9d19ee437c8dacee5a3e6f1a014e226dfd13c9d7c7d61d48dd33aacc
MD5 31ed38cf2c8745532120e055c5c4144c
BLAKE2b-256 fe47756faf6a9e100e511e7e8fc11e48ea7d8fd5d6a9c0962f69eca7b7d4dfce

See more details on using hashes here.

Provenance

The following attestation bundles were made for humming_kernels-0.1.13-py3-none-manylinux_2_28_x86_64.whl:

Publisher: publish.yml on inclusionAI/humming

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file humming_kernels-0.1.13-py3-none-manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for humming_kernels-0.1.13-py3-none-manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 bbba0cf993a6afa8845d1ee4e0cda135559c66cf6ce399ac101b15d031497d8b
MD5 8c90d45e5ff49ef02dab8be77bd20ce7
BLAKE2b-256 436f7ed467a6d430e557ea381044d456adcdf53ba42d19faae991a97cd75c6b4

See more details on using hashes here.

Provenance

The following attestation bundles were made for humming_kernels-0.1.13-py3-none-manylinux_2_28_aarch64.whl:

Publisher: publish.yml on inclusionAI/humming

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.1.13 This release

3 files

0.1.12

3 files

0.1.11

2 files

0.1.10

2 files

0.1.9

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page