Skip to main content

Uni-Quant

Small library to quantize/dequantize TensorFlow models using PyTorch CUDA kernels.

Requirements

  • Python: 3.13.13 (haven't tested on any other)
  • CUDA Toolkit: >=12.8
  • Python Dependencies: All required packages are listed in requirements.txt

Installing Dependencies

pip install -r requirements.txt

Installation from pip

pip install uni-quant-cuda

Usage

Importing Functions

from uniquant import quantize, dequantize, dequantize_save

Main Functions

quantize(model_path, quant_directory="", quant_name="", pack_size=32, quant_size=4, overwrite=False)

Quantizes a TensorFlow or XGBoost model.

Arguments:

  • model_path (str): Path to the model to quantize (with extension)
  • quant_directory (str): Directory path to save the quantized model
  • quant_name (str): Filename for the quantized model
  • pack_size (int): Number of weights in one quantization batch (must be divisible by 2)
  • quant_size (int): Number of bits per weight (available: 4 or 8)
  • overwrite (bool): Whether to overwrite existing file

dequantize(quant_path, literal=False, balanced=True)

Dequantizes a model and returns it.

Arguments:

  • quant_path (str): Path to the .uniq file to dequantize
  • literal (bool): Whether weights should be unscaled
  • balanced (bool): Whether weights should be balanced around 0

dequantize_save(quant_path, model_directory="", model_name="", overwrite=False)

Dequantizes a model, saves it, and returns it.

Arguments:

  • quant_path (str): Path to the .uniq file to dequantize
  • model_directory (str): Directory path to save the dequantized model
  • model_name (str): Filename for the dequantized model
  • overwrite (bool): Whether to overwrite existing file

Notes

  • This package compiles CUDA kernels at runtime using torch.utils.cpp_extension.load_inline.
  • Installing and using the CUDA compilation requires a compatible CUDA toolkit on the target machine (tested with >=12.8).

Release files for uni-quant-cuda 0.2.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for uni-quant-cuda 0.2.8
File Size Uploaded
uni_quant_cuda-0.2.8.tar.gz 17.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for uni-quant-cuda 0.2.8
File Interpreter ABI Platform
uni_quant_cuda-0.2.8-py3-none-any.whl Python 3 none any Details

Total release size: 30.5 kB

Release files / uni_quant_cuda-0.2.8.tar.gz

Download URL uni_quant_cuda-0.2.8.tar.gz
Size 17.2 kB
Tags Source
SHA-256 checksum
How to use checksums
46dd5f06928c41d120797bb264180ccd7817bc6b4d4a47ce85465f0a7195fef5
BLAKE2b-256 checksum
How to use checksums
4f4227f775c28fd82e634ab1703f9d9b32f54a19a726cedc30c601679d5dfcba
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.13

Release files / uni_quant_cuda-0.2.8-py3-none-any.whl

Download URL uni_quant_cuda-0.2.8-py3-none-any.whl
Size 13.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fe1ebd81e4a1ec2cf3f7fd27c5bd0896e07342f469a705787f51fcb567fbc147
BLAKE2b-256 checksum
How to use checksums
47fc7f6d16dca50f72172e49a501e65222c83d207fc485d77d700d59a38d1b04
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.13

Release history Release notifications | RSS feed

This release

0.2.8 This release

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page