This release is a pre-release and may not be stable for production use.
AI Edge Quantizer
AI Edge Quantizer (AEQ) is a flexible, high-performance post-training quantization (PTQ) toolkit designed for LiteRT (formerly TensorFlow Lite) and LiteRT-LM. It enables developers to optimize resource-intensive models (vision models, LLMs, and GenAI pipelines) for edge deployment on mobile CPUs, GPUs, and NPUs.
Key Features
-
Selective Quantization: Target specific operations or layer subgraphs using regex scopes (e.g., FullyConnected Ops in FeedForward layers and leaving all other Ops as float).
-
Mixed-Precision Quantization: Mix precision schemes across layers (e.g., INT4 weights for FullyConnected in FeedForward layers but INT8 weights in Attention layers).
-
Advanced Quantization Algorithms:
- Blockwise Quantization (Block sizes: 32, 64, 128, 256)
- Hadamard Transformations to suppress outlier activations and preserve accuracy for INT4/INT2 schemes
- GPTQ & OCTAV optimization algorithms
-
Full Integer Quantization (Static Range, SRQ): INT8/INT16 activations with INT8/INT4 weights, required for mobile hardware.
-
Integrated Numerical Validation: Built-in tensor-level distortion analysis (MSE, SNR, Cosine Similarity, KL Divergence) compatible with Model Explorer.
Build Status
| Build Type | Status |
|---|---|
| Unit Tests (Linux) | |
| Nightly Release | |
| Nightly Colab |
Installation
Requirements and Dependencies
- Python versions: 3.10, 3.11, 3.12, 3.13
- Operating system: Linux, MacOS
- LiteRT:
ai-edge-litert-nightly
Install
Nightly PyPi package:
pip install ai-edge-quantizer-nightly
Quick Start
The quantizer requires two inputs:
- An unquantized source LiteRT (FP32 data type in the FlatBuffer
format with
.tfliteextension) / LiteRT-LM (with.litertlmextension) model - A quantization recipe (details below)
and outputs a quantized LiteRT/LiteRT-LM model that's ready for deployment on edge devices.
Command Line (aeq)
Quantize a model directly from your terminal:
# Quantize standard .tflite model
aeq --model_file="path/to/input.tflite" \
--recipe=dynamic_wi8_afp32 \
--output_dir="/path/to/output"
# Quantize a .litertlm LLM container
aeq --model_file="path/to/gemma.litertlm" \
--recipe=gemma4_mixed48 \
--output_dir="/path/to/output"
Python API
from ai_edge_quantizer import quantizer, recipe
# 1. Initialize quantizer
qt = quantizer.Quantizer("path/to/model.tflite")
# 2. Load a ready-to-use recipe (e.g., dynamic int8 weights with float32
# activations).
qt.load_quantization_recipe(recipe.dynamic_wi8_afp32())
# Quantize and export.
qt.quantize().export_model("path/to/quantized_model.tflite")
Please see the getting started colab for the simplest quick start guide on those steps, and the selective quantization colab for more details on advanced features.
Hardware & Recipe Decision Guide
Generally, we recommend dynamic quantization for CPU/GPU deployment and static quantization for NPU deployment:
| Target Hardware | Recommended Recipe | Precision | Activation Calibration? |
|---|---|---|---|
| CPU/GPU | dynamic_wi8_afp32 |
Int8 Weight / FP32 Act | No |
| NPU | static_wi8_ai8 / static_wi8_ai16 |
Int8 Weight / Int8 or Int16 Act | Yes (Requires calibration data, supported only via Python API) |
Quantization Concepts & Methods
LiteRT Model
Please refer to the LiteRT documentation
for ways to generate LiteRT models from Jax, PyTorch and TensorFlow. The input
source model should be an FP32 (unquantized) model in the FlatBuffer format with
.tflite extension.
LiteRT-LM Model
Please refer to the LiteRT-LM documentation for details.
Quantization Recipe
A quantization recipe encodes all information on how a model is to be quantized, such as number of bits, data type, symmetry, scope name, etc.
Essentially, a quantization recipe is defined as a collection of commands of the following type:
“Apply Quantization Algorithm X on Operator Y under Scope Z with ConfigN”.
For example:
"Uniformly quantize the FullyConnected op under scope 'dense1/' with INT8 symmetric with Dynamic Quantization".
All the unspecified ops will be kept as FP32 (unquantized). The scope of an operator in TFLite is defined as the output tensor name of the op, which preserves the hierarchical model information from the source model (e.g., scope in TF). The best way to obtain scope name is by visualizing the model with Model Explorer.
Quantization Methods
Currently, there are three ways to quantize an operator:
-
dynamic quantization (recommended): weights are quantized while activations remain in a float format and are not processed by AI Edge Quantizer (AEQ). The runtime kernel handles the on-the-fly quantization of these activations, as identified by
compute_precision=integerandexplicit_dequantize=False.- Pros: reduced model size and memory usage. Latency improvement due to integer computation. No sample data requirement (calibration).
- Cons: on-the-fly quantization of activation tensors can affect model quality. Not supported in all hardware (e.g., some GPU and NPU).
-
weight only quantization: only model weights are quantized, not activations. The actual operation (op) computation remains in float. The quantized weight is explicitly dequantized before being fed into the op, by inserting a dequantize op between the quantized weight and the consuming op. To enable this,
compute_precisionwill be set tofloatandexplicit_dequantizetoTrue.- Pros: reduced model size and memory usage. No sample data requirement (calibration). Usually has the best model quality.
- Cons: no latency benefit (may be worse) due to float computation with explicit dequantization.
-
static quantization: both weights and activations are quantized. This requires a calibration phase to estimate quantization parameters of runtime tensors (activations).
- Pros: reduced model size, memory usage, and latency.
- Cons: requires sample data for calibration. Imposing static quantization parameters (derived from calibration) on runtime tensors can compromise quality.
We include commonly used recipes in recipe.py. This is demonstrated in the getting started colab example. Advanced users can build their own recipe through the quantizer API.
Quantization Workflow
Quantizing a model with AI Edge Quantizer follows a structured 7-step lifecycle:
- Load Model
- Load / Configure Recipe
- [Static recipes only] Calibrate
- Quantize & Export
- Validate Accuracy
- Visualize
- Deploy
Note on CLI Usage: The aeq command-line tool executes
Steps 1, 2, and4 in a single command for recipes that do not require
calibration (Dynamic Range, Weight-Only):
aeq --model_file="path/to/model.tflite" \
--recipe=dynamic_wi8_afp32 \
--output_dir="/path/to/output"
For Full-Integer Static Quantization requiring representative calibration datasets (Step 3), use the Python API.
Detailed examples on Steps 1-4 with Python API can be found in quantize_toy_model.py.
Step 1: Initialize Quantizer with Source Model
Load an unquantized FP32 .tflite model or a .litertlm generative model
bundle:
from ai_edge_quantizer import quantizer
qt = quantizer.Quantizer("path/to/model.tflite")
Step 2: Choose & Load a Quantization Recipe
Load a ready-to-use recipe (e.g., static int8 quantization) or custom recipe:
from ai_edge_quantizer import recipe
qt.load_quantization_recipe(recipe.static_wi8_ai8())
Supported Operators and Recipes
Please refer to the Operator Coverage section for more details on supported operators and configurations for each recipe.
Advanced Recipes & Customization
There are many ways the user can configure and customize the quantization recipe beyond using a template in recipe.py. For example, the user can configure the recipe to achieve these features:
- Selective quantization (exclude selected ops from being quantized)
- Flexible mixed scheme quantization (mixture of different precision, compute precision, scope, op, config, etc)
- 4-bit weight quantization
- Advanced algorithms (e.g., Hadamard Rotation, OCTAV)
The selective quantization colab shows some of these more advanced features.
For specifics of the recipe schema, please refer to the OpQuantizationRecipe
in recipe_manager.py.
For advanced usage involving mixed quantization, the following API may be useful:
- Use
Quantizer:load_quantization_recipe()in quantizer.py to load a custom recipe. - Use
Quantizer:update_quantization_recipe()in quantizer.py to extend or override specific parts of the recipe.
Step 3: Calibrate with Representative Data (Static Quantization Only)
Static range quantization (e.g., static_wi8_ai8, static_wi8_ai16) quantizes
both weights and activations into integers, requiring a calibration phase with
representative sample data to calculate quantization statistics values (QSVs).
Calibration data is structured as a dictionary mapping signature keys (e.g.,
'serving_default') to lists of input sample dictionaries:
from ai_edge_quantizer import quantizer, recipe
import numpy as np
qt = quantizer.Quantizer("path/to/model.tflite")
qt.load_quantization_recipe(recipe.static_wi8_ai8())
if qt.need_calibration:
# Provide representative calibration data matching the model signature inputs.
calibration_data = {
"serving_default": [
{
"input_tensor_name": np.random.uniform(
-1.0, 1.0, size=(1, 28, 28, 1)
).astype(np.float32)
}
for _ in range(256)
]
}
calibration_result = qt.calibrate(calibration_data)
qt.quantize(calibration_result=calibration_result).export_model(
"/path/to/output/quantized_model.tflite"
)
Step 4: Quantize & Export the Model
Execute the quantization engine and export the resulting model:
qt.quantize().export_model("/path/to/output/quantized_model.tflite")
Step 5: Validate Numerical Accuracy
Quantizing a model inherently introduces numerical noise. After calling
qt.quantize(), you can verify the mathematical distortion between the float
baseline and the quantized model using the built-in validate() method, which
returns a single ComparisonResult object mapping nodes to their error metric
values. You can print them or automatically save them to Model Explorer JSON
files:
# 1. Default validation (evaluates MSE metric by default)
comparison_results = qt.validate(test_data=sample_data)
print(
"Per-layer metrics:",
comparison_results.get_all_tensor_results(),
)
# 2. Multi-metric validation (save all metrics and validation json data directly)
comparison_results = qt.validate(
test_data=sample_data,
error_metrics=[
quantizer.ValidationErrorMetric.MSE,
quantizer.ValidationErrorMetric.SNR,
],
save_folder='/tmp/',
)
all_results = comparison_results.get_all_tensor_results()
for tensor_name, metrics in all_results.items():
print(
f"Tensor: {tensor_name} "
f"- MSE: {metrics.get(quantizer.ValidationErrorMetric.MSE.value, 0.0):.6f} "
f"- SNR: {metrics.get(quantizer.ValidationErrorMetric.SNR.value, 0.0):.6f}"
)
Step 6: Visualize Models with Model Explorer
The best way to obtain exact operator scope names and visually compare tensor shapes and quantization scales between baseline float and quantized graphs is using Model Explorer.
To visualize two exported .tflite models side-by-side in your terminal, run:
model_explorer --models \
"/path/to/baseline_float.tflite,/path/to/quantized_model.tflite"
Step 7: Deploy on Edge Hardware
Please refer to the LiteRT deployment documentation for ways to deploy a quantized LiteRT model.
Operator Coverage
Allowed Configurations for Available recipes
| Config | DYNAMIC_WI8_AFP32 | DYNAMIC_WI4_AFP32 | DYNAMIC_WI4_AFP32_BLOCKWISE | DYNAMIC_WI2_AFP32_BLOCKWISE | STATIC_WI8_AI8 | STATIC_WI8_AI16 | STATIC_WI4_AI8 | STATIC_WI4_AI16 | WEIGHTONLY_WI8_AFP32 | WEIGHTONLY_WI4_AFP32 | |
| activation | num_bits | None | None | None | None | 8 | 16 | 8 | 16 | None | None |
| symmetric | None | None | None | None | [TRUE, FALSE] | TRUE | [TRUE, FALSE] | TRUE | None | None | |
| granularity | None | None | None | None | TENSORWISE | TENSORWISE | TENSORWISE | TENSORWISE | None | None | |
| dtype | None | None | None | None | INT | INT | INT | INT | None | None | |
| weight | num_bits | 8 | 4 | 4 | 2 | 8 | 8 | 4 | 4 | 8 | 4 |
| symmetric | TRUE | TRUE | TRUE | TRUE | TRUE | TRUE | TRUE | TRUE | [TRUE, FALSE] | [TRUE, FALSE] | |
| granularity | [CHANNELWISE, TENSORWISE] | [CHANNELWISE, TENSORWISE] | [BLOCKWISE_32, BLOCKWISE_64, BLOCKWISE_128, BLOCKWISE_256] | [BLOCKWISE_32, BLOCKWISE_64, BLOCKWISE_128, BLOCKWISE_256] | [CHANNELWISE, TENSORWISE] | [CHANNELWISE, TENSORWISE] | [CHANNELWISE, TENSORWISE] | [CHANNELWISE, TENSORWISE] | [CHANNELWISE, TENSORWISE] | [CHANNELWISE, TENSORWISE] | |
| dtype | INT | INT | INT | INT | INT | INT | INT | INT | INT | INT | |
| explicit_dequantize | FALSE | FALSE | FALSE | FALSE | FALSE | FALSE | FALSE | FALSE | TRUE | TRUE | |
| compute_precision | INTEGER | INTEGER | INTEGER | INTEGER | INTEGER | INTEGER | INTEGER | INTEGER | FLOAT | FLOAT |
Quantization Support for Operators with Weights
| Config | DYNAMIC_WI8_AFP32 | DYNAMIC_WI4_AFP32 | DYNAMIC_WI4_AFP32_BLOCKWISE | DYNAMIC_WI2_AFP32_BLOCKWISE | STATIC_WI8_AI8 | STATIC_WI8_AI16 | STATIC_WI4_AI8 | STATIC_WI4_AI16 | WEIGHTONLY_WI8_AFP32 | WEIGHTONLY_WI4_AFP32 |
| BATCH_MATMUL | ✓ |
✓ |
✓ |
✓ |
✓ |
|||||
| CONV_2D | ✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
||
| CONV_2D_TRANSPOSE | ✓ |
✓ |
✓ |
✓ |
||||||
| DEPTHWISE_CONV_2D | ✓ |
✓ |
✓ |
✓ |
||||||
| EMBEDDING_LOOKUP | ✓ |
✓ |
✓ |
✓ |
✓ |
|||||
| FULLY_CONNECTED | ✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
Quantization Support for Activations-Only Operators
| Config | STATIC_WI8_AI8 | STATIC_WI8_AI16 |
| ADD | ✓ |
✓ |
| AVERAGE_POOL_2D | ✓ |
✓ |
| BROADCAST_TO | ✓ |
✓ |
| CONCATENATION | ✓ |
✓ |
| DIV | ✓ |
✓ |
| DYNAMIC_UPDATE_SLICE | ✓ |
✓ |
| EQUAL | ✓ |
✓ |
| GATHER | ✓ |
✓ |
| GATHER_ND | ✓ |
✓ |
| GELU | ✓ |
✓ |
| HARD_SWISH | ✓ |
|
| LOGISTIC | ✓ |
✓ |
| MAX_POOL_2D | ✓ |
✓ |
| MAXIMUM | ✓ |
✓ |
| MEAN | ✓ |
✓ |
| MIRROR_PAD | ✓ |
✓ |
| MUL | ✓ |
✓ |
| NOT_EQUAL | ✓ |
✓ |
| PACK | ✓ |
✓ |
| PAD | ✓ |
✓ |
| PADV2 | ✓ |
✓ |
| REDUCE_MIN | ✓ |
✓ |
| RELU | ✓ |
✓ |
| RESHAPE | ✓ |
✓ |
| RESIZE_BILINEAR | ✓ |
✓ |
| RESIZE_NEAREST_NEIGHBOR | ✓ |
✓ |
| RSQRT | ✓ |
✓ |
| SELECT | ✓ |
✓ |
| SELECT_V2 | ✓ |
✓ |
| SLICE | ✓ |
✓ |
| SOFTMAX | ✓ |
✓ |
| SPACE_TO_DEPTH | ✓ |
|
| SPLIT | ✓ |
✓ |
| SQRT | ✓ |
✓ |
| SQUARED_DIFFERENCE | ✓ |
|
| STRIDED_SLICE | ✓ |
✓ |
| SUB | ✓ |
✓ |
| SUM | ✓ |
✓ |
| TANH | ✓ |
✓ |
| TRANSPOSE | ✓ |
✓ |
| UNPACK | ✓ |
✓ |
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ai_edge_quantizer_nightly-0.10.0.dev20260821.tar.gz.
File metadata
- Download URL: ai_edge_quantizer_nightly-0.10.0.dev20260821.tar.gz
- Upload date:
- Size: 266.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
615eb4873d8376d62f0211b43bcdac7684543f5a40ae3183fd18102fe8de5f91
|
|
| MD5 |
6c6a62faf6bd375a72b8bad3443fbd81
|
|
| BLAKE2b-256 |
7bc25069c3974e90da0e959650f0053ef2327fc314efc117889f68d2c2953a81
|
File details
Details for the file ai_edge_quantizer_nightly-0.10.0.dev20260821-py3-none-any.whl.
File metadata
- Download URL: ai_edge_quantizer_nightly-0.10.0.dev20260821-py3-none-any.whl
- Upload date:
- Size: 481.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1e0d5c4f95b8a7199d0159d06805496d59ee87fb8ff6902d1d7e293e7bdaa85c
|
|
| MD5 |
b246d28c611a92ae97f6c0e18e4778f8
|
|
| BLAKE2b-256 |
be15a35128bf8217efd18997bf48b0d1bc8ab0fff6ae464250c35aabf5aba18d
|