PQuantML
PQuantML is an end-to-end library for training compressed machine learning models, developed at CERN as part of the Next Generation Triggers project.
It supports:
- Joint pruning + quantization
- Layer-wise precision configuration
- Flexible training pipelines
- PyTorch and TensorFlow backends
- Knowledge distillation
- HGQ library integration
- Integration with hardware-friendly toolchains (e.g., hls4ml)
PQuantML enables efficient deployment of compact neural networks on resource-constrained hardware such as FPGAs and embedded accelerators.
Installation
Install the base package via pip:
pip install pquant-ml
Install with a specific backend:
pip install "pquant-ml[tensorflow]" # TensorFlow backend
pip install "pquant-ml[torch]" # PyTorch backend
Supported layers
| Layer | Description |
|---|---|
PQConv*D |
Convolutional layers |
PQAvgPool*D |
Average pooling layers |
PQBatchNorm*D |
Batch normalization layers |
PQDense |
Linear (fully connected) layer |
PQActivation |
Activation layers: ReLU, Tanh, Leaky ReLU, GELU, Hard Tanh, or a user-provided activation function (Torch only) |
MultiHeadAttention |
Multi-head attention layer |
LayerNorm |
Layer normalization layer (currently Torch only) |
Training
Different pruning methods involve different training stages, such as pre-training and fine-tuning. PQuantML provides a generic training function: you supply your own training and validation functions along with the number of epochs, and PQuant handles the training loop while automatically triggering the appropriate stages for the chosen pruning method.
Quantization
PQuantML supports two quantization modes, each with several granularity options.
Fixed-point quantization (for weights):
- per-weight
- per-channel
- per-tensor
HGQ (High Granularity Quantization):
- per-weight
- per-tensor
Example
Example notebooks are available in the examples/ directory. It shows how to:
- Create a Torch model and data loaders.
- Create the training and validation functions.
- Load a default configuration for a pruning method.
- Train and compress the model by passing the configuration, model, and training/validation functions to PQuant's training function.
- Build a custom quantization and pruning configuration for a given model (e.g. disabling pruning for some layers, or using different quantization bit-widths per layer).
- Use the direct-layer and layer-replacement approaches.
- Use the HPO platform.
Documentation
Full documentation is available at pquantml.readthedocs.io.
Citation
The framework is described in PQuantML: A Tool for End-to-End Hardware-aware Model Compression (arXiv:2603.26595).
If you use PQuantML in your work, please cite:
@article{niemi2026pquantml,
title = {PQuantML: A Tool for End-to-End Hardware-aware Model Compression},
author = {Niemi, Roope and Petrovych, Anastasiia and Das, Arghya and
Lupi, Enrico and Sun, Chang and Danopoulos, Dimitrios and
Helbing, Marlon Joshua and Liu, Mia and Kagan, Michael and
Loncar, Vladimir and Pierini, Maurizio},
journal = {arXiv preprint arXiv:2603.26595},
year = {2026}
}
Authors
- Roope Niemi (CERN)
- Anastasiia Petrovych (CERN)
- Arghya Das (Purdue University)
- Enrico Lupi (CERN)
- Chang Sun (Caltech)
- Dimitrios Danopoulos (CERN)
- Marlon Joshua Helbing
- Mia Liu (Purdue University)
- Michael Kagan (SLAC National Accelerator Laboratory)
- Vladimir Loncar (CERN)
- Maurizio Pierini (CERN)
Metadata
Release files for pquant-ml 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pquant_ml-0.1.0.tar.gz | 1.8 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pquant_ml-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.9 MB
Release files / pquant_ml-0.1.0.tar.gz
| Download URL | pquant_ml-0.1.0.tar.gz |
|---|---|
| Size | 1.8 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1fe2c787750b32f5752797b1ce2ef4b84165c25a1c93654d47c0e5091be632b7
|
|
BLAKE2b-256 checksum How to use checksums |
2ca7bc58ebd99e6592c2079f80f2fbec2290dee0304b64936131a27870625061
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 31, 2026.
Transparency logRelease files / pquant_ml-0.1.0-py3-none-any.whl
| Download URL | pquant_ml-0.1.0-py3-none-any.whl |
|---|---|
| Size | 169.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2ee28b52ee71175512ef69ca239f6dfbd45318d11d4ab7d12fbb14a2f4213146
|
|
BLAKE2b-256 checksum How to use checksums |
91325ac3bc237c9547bab9e5dd4a0d61547cbe7cd58e1c59d37c7786924d1a84
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 31, 2026.
Transparency log