Skip to main content

Neural Compressor

An open-source Python library supporting popular model compression techniques for ONNX

python version license


Neural Compressor aims to provide popular model compression techniques inherited from Intel Neural Compressor yet focused on ONNX model quantization such as SmoothQuant, weight-only quantization through ONNX Runtime. In particular, the tool provides the key features, typical examples, and open collaborations as below:

Installation

Install from source

git clone https://github.com/onnx/neural-compressor.git
cd neural-compressor
pip install -r requirements.txt
pip install .

Note: Further installation methods can be found under Installation Guide.

Getting Started

Setting up the environment:

pip install onnx-neural-compressor "onnxruntime>=1.17.0" onnx

After successfully installing these packages, try your first quantization program.

Notes: please install from source before the formal pypi release.

Weight-Only Quantization (LLMs)

Following example code demonstrates Weight-Only Quantization on LLMs, device will be selected for efficiency automatically when multiple devices are available.

Run the example:

from onnx_neural_compressor.quantization import matmul_nbits_quantizer

algo_config = matmul_nbits_quantizer.RTNWeightOnlyQuantConfig()
quant = matmul_nbits_quantizer.MatMulNBitsQuantizer(
    model,
    n_bits=4,
    block_size=32,
    is_symmetric=True,
    algo_config=algo_config,
)
quant.process()
best_model = quant.model

Static Quantization

from onnx_neural_compressor.quantization import quantize, config
from onnx_neural_compressor import data_reader


class DataReader(data_reader.CalibrationDataReader):
    def __init__(self):
        self.encoded_list = []
        # append data into self.encoded_list

        self.iter_next = iter(self.encoded_list)

    def get_next(self):
        return next(self.iter_next, None)

    def rewind(self):
        self.iter_next = iter(self.encoded_list)


data_reader = DataReader()
qconfig = config.StaticQuantConfig(calibration_data_reader=data_reader)
quantize(model, output_model_path, qconfig)

Documentation

Overview
Architecture Workflow Examples
Feature
Quantization SmoothQuant
Weight-Only Quantization (INT8/INT4) Layer-Wise Quantization

Additional Content

Communication

  • GitHub Issues: mainly for bug reports, new feature requests, question asking, etc.
  • Email: welcome to raise any interesting research ideas on model compression techniques by email for collaborations.

Metadata

Release files for onnx-neural-compressor 1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for onnx-neural-compressor 1.0
File Size Uploaded
onnx_neural_compressor-1.0.tar.gz 95.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for onnx-neural-compressor 1.0
File Interpreter ABI Platform
onnx_neural_compressor-1.0-py3-none-any.whl Python 3 none any Details

Total release size: 238.3 kB

Release files / onnx_neural_compressor-1.0.tar.gz

Download URL onnx_neural_compressor-1.0.tar.gz
Size 95.3 kB
Tags Source
SHA-256 checksum
How to use checksums
7d04a517a36c1bb0e976b014dbf51bea7b6b747136a409ea0959351d4a8acce1
BLAKE2b-256 checksum
How to use checksums
696a25cdb4307e361d54ca7c824e35fa325925134d5df91c50538fad846fe774
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.3.0 pkginfo/1.5.0.1 requests/2.23.0 setuptools/46.4.0.post20200518 requests-toolbelt/0.9.1 tqdm/4.46.0 CPython/3.8.3

Release files / onnx_neural_compressor-1.0-py3-none-any.whl

Download URL onnx_neural_compressor-1.0-py3-none-any.whl
Size 143.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6896dd9084e75812ce6d28bc794e8f5d7e2f8a394c23c81d5232624aa1db8722
BLAKE2b-256 checksum
How to use checksums
9ffb750b57c3174bccb6b77518f5fa1f0cf308e38cbdf24209f62be31d8eceff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.3.0 pkginfo/1.5.0.1 requests/2.23.0 setuptools/46.4.0.post20200518 requests-toolbelt/0.9.1 tqdm/4.46.0 CPython/3.8.3

Release history Release notifications | RSS feed

This release

1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page