Skip to main content

LeanQuant

Overview

This package provides efficient inference kernels for running non-uniformly quantized LeanQuant models on CUDA-enabled GPUs. LeanQuant is a scalable and accurate quantization algorithm that compresses large language models by 4-8x while maintaining competitive performance.

Installation

Ensure your GPU supports CUDA 11 or CUDA 12. You can check your CUDA version with the command nvidia-smi | grep CUDA.

To install:

# For CUDA 11.x
pip install leanquant[cuda11]

# For CUDA 12.x
pip install leanquant[cuda12]

Models

Quantized LeanQuant models are available for download on our HuggingFace page: huggingface.co/LeanQuant

Technical Details

LeanQuant introduces an innovative loss-error-aware grid approach to quantization that significantly outperforms traditional methods. Our technique:

  • Achieves superior compression ratios: Reduces model size by 4-8x without sacrificing capability
  • Preserves model intelligence: Maintains performance comparable to full-precision models across challenging benchmarks
  • Optimizes GPU execution: Features custom CUDA kernels specifically designed for non-uniform quantization format

The algorithm strategically allocates quantization precision based on parameter sensitivity, ensuring computational resources are focused where they matter most.

For a comprehensive explanation of our methodology and benchmark results, please refer to our research paper.

Citation

If you find LeanQuant useful in your research or applications, please consider citing our work:

@inproceedings{
    zhang2025leanquant,
    title={LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid},
    author={Tianyi Zhang and Anshumali Shrivastava},
    booktitle={The Thirteenth International Conference on Learning Representations},
    year={2025},
    url={https://openreview.net/forum?id=ISqx8giekS}
}

Metadata

Release files for leanquant 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for leanquant 0.1.1
File Size Uploaded
leanquant-0.1.1.tar.gz 5.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for leanquant 0.1.1
File Interpreter ABI Platform
leanquant-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 10.9 kB

Release files / leanquant-0.1.1.tar.gz

Download URL leanquant-0.1.1.tar.gz
Size 5.3 kB
Tags Source
SHA-256 checksum
How to use checksums
5b269964faa392adffdd07f07528debecbe43ba27d9b64f0ae149a6450068c35
BLAKE2b-256 checksum
How to use checksums
e70eee91cad8fff5cdc9c543f17eba09c80a374ebcdf5836662dfbeb737d4f57
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.16

Release files / leanquant-0.1.1-py3-none-any.whl

Download URL leanquant-0.1.1-py3-none-any.whl
Size 5.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ba123232c54910ae7985cbf3d6109133292d42fca7b72710535454760a961b61
BLAKE2b-256 checksum
How to use checksums
83212839ba4a349c27a971ecae978986949a0180f06c1ceec806df012f4e3e13
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.16

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page