FMS Model Optimizer
Introduction
FMS Model Optimizer is a framework for developing reduced precision neural network models. Quantization techniques, such as quantization-aware-training (QAT), post-training quantization (PTQ), and several other optimization techniques on popular deep learning workloads are supported.
Highlights
- Python API to enable model quantization: With the addition of a few lines of codes, module-level and/or function-level operations replacement will be performed.
- Robust: Verified for INT 8/4-bit quantization on important vision/speech/NLP/object detection/LLMs.
- Flexible: Options to analyze the network using PyTorch Dynamo, apply best practices, such as clip_val initialization, layer-level precision setting, optimizer param group setting, etc. during quantization.
- State-of-the-art INT and FP quantization techniques for weights and activations, such as SmoothQuant, SAWB+ and PACT+.
- Supports key compute-intensive operations like Conv2d, Linear, LSTM, MM and BMM
Supported Models
| GPTQ | FP8 | PTQ | QAT | |
|---|---|---|---|---|
| Granite | ✅ | ✅ | ✅ | 🔲 |
| Llama | ✅ | ✅ | ✅ | 🔲 |
| Mixtral | ✅ | ✅ | ✅ | 🔲 |
| BERT/Roberta | ✅ | ✅ | ✅ | ✅ |
Note: Direct QAT on LLMs is not recommended
Getting Started
Requirements
- 🐧 Linux system with Nvidia GPU (V100/A100/H100)
- Python 3.10 to Python 3.12
- CUDA >=12
Optional packages based on optimization functionality required:
- GPTQ is a popular compression method for LLMs:
- If you want to experiment with INT8 deployment in QAT and PTQ examples:
- FP8 is a reduced precision format like INT8:
- Nvidia A100 family or higher
- llm-compressor
- To enable compute graph plotting function (mostly for troubleshooting purpose):
Installation
We recommend using a Python virtual environment with Python 3.9+. Here is how to setup a virtual environment using Python venv:
python3 -m venv fms_mo_venv
source fms_mo_venv/bin/activate
There are 2 ways to install the FMS Model Optimizer as follows:
From Release
To install from release (PyPi package):
python3 -m venv fms_mo_venv
source fms_mo_venv/bin/activate
pip install fms-model-optimizer
From Source
To install from source(GitHub Repository):
python3 -m venv fms_mo_venv
source fms_mo_venv/bin/activate
git clone https://github.com/foundation-model-stack/fms-model-optimizer
cd fms-model-optimizer
pip install -e .
Optional Dependencies
The following optional dependencies are available:
fp8:llmcompressorandtorchaopackages for fp8 quantization and inferencefp8-infer:torchaopackage for fp8 inferencegptq:GPTQModelpackage for W4A16 quantizationmx:microxcalingpackage for MX quantizationopt: Shortcut forfp8,gptq, andmxinstallsaiu:ibm-fmspackage for AIU model deploymenttorchvision:torchpackage for image recognition training and inferencetriton:tritonpackage for matrix multiplication kernelsexamples: Dependencies needed for examplesvisualize: Dependencies for visualizing models and performance datatest: Dependencies needed for unit testingdev: Dependencies needed for development
To install an optional dependency, modify the pip install commands above with a list of these names enclosed in brackets. The example below installs llm-compressor and torchvision with FMS Model Optimizer:
pip install fms-model-optimizer[fp8,torchvision]
pip install -e .[fp8,torchvision]
If you have already installed FMS Model Optimizer, then only the optional packages will be installed.
Try It Out!
To help you get up and running as quickly as possible with the FMS Model Optimizer framework, check out the following resources which demonstrate how to use the framework with different quantization techniques:
- Jupyter notebook tutorials (It is recommended to begin here):
- Quantization tutorial:
- Visualizes a random Gaussian tensor step-by-step along the quantization process
- Build a quantizer and quantized convolution module based on this process
- Quantization tutorial:
- Python script examples
Docs
Dive into the design document to get a better understanding of the framework motivation and concepts.
Contributing
Check out our contributing guide to learn how to contribute.
Release files for fms-model-optimizer 0.8.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fms_model_optimizer-0.8.6.tar.gz | 5.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fms_model_optimizer-0.8.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.6 MB
Release files / fms_model_optimizer-0.8.6.tar.gz
| Download URL | fms_model_optimizer-0.8.6.tar.gz |
|---|---|
| Size | 5.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f939b3c1ce955fe0be93f76628199005d1303a5608d5cb468e3dd09a51d66a8d
|
|
BLAKE2b-256 checksum How to use checksums |
16203eae9ea4eaf363c8e46d5e8aa863c58bfda78a13b1bf4aded43b2c080d83
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / fms_model_optimizer-0.8.6-py3-none-any.whl
| Download URL | fms_model_optimizer-0.8.6-py3-none-any.whl |
|---|---|
| Size | 364.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
eeeb016b4cf8530ce2686809904a9aa805d03707b1a3896b5ef8d1489650ef4f
|
|
BLAKE2b-256 checksum How to use checksums |
fe64ec4b9c8fa49dc05dc3761239e535acea2ae6463dc691e2c1669bb7bd530d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log