Skip to main content

FMS Model Optimizer

Lint Tests Build Minimum Python Version Release License

Introduction

FMS Model Optimizer is a framework for developing reduced precision neural network models. Quantization techniques, such as quantization-aware-training (QAT), post-training quantization (PTQ), and several other optimization techniques on popular deep learning workloads are supported.

Highlights

  • Python API to enable model quantization: With the addition of a few lines of codes, module-level and/or function-level operations replacement will be performed.
  • Robust: Verified for INT 8/4-bit quantization on important vision/speech/NLP/object detection/LLMs.
  • Flexible: Options to analyze the network using PyTorch Dynamo, apply best practices, such as clip_val initialization, layer-level precision setting, optimizer param group setting, etc. during quantization.
  • State-of-the-art INT and FP quantization techniques for weights and activations, such as SmoothQuant, SAWB+ and PACT+.
  • Supports key compute-intensive operations like Conv2d, Linear, LSTM, MM and BMM

Supported Models

GPTQ FP8 PTQ QAT
Granite ✅ ✅ ✅ 🔲
Llama ✅ ✅ ✅ 🔲
Mixtral ✅ ✅ ✅ 🔲
BERT/Roberta ✅ ✅ ✅ ✅

Note: Direct QAT on LLMs is not recommended

Getting Started

Requirements

  1. 🐧 Linux system with Nvidia GPU (V100/A100/H100)
  2. Python 3.10 to Python 3.12
  3. CUDA >=12

Optional packages based on optimization functionality required:

  • GPTQ is a popular compression method for LLMs:
  • If you want to experiment with INT8 deployment in QAT and PTQ examples:
    • Nvidia GPU with compute capability > 8.0 (A100 family or higher)
    • Option 1:
      • Ninja
      • Clone the CUTLASS repository
      • PyTorch 2.3.1 (as newer version will cause issue for the custom CUDA kernel used in these examples)
    • Option 2:
      • use triton kernel included. But this kernel is currently not faster than FP16.
  • FP8 is a reduced precision format like INT8:
  • To enable compute graph plotting function (mostly for troubleshooting purpose):

Installation

We recommend using a Python virtual environment with Python 3.9+. Here is how to setup a virtual environment using Python venv:

python3 -m venv fms_mo_venv
source fms_mo_venv/bin/activate

There are 2 ways to install the FMS Model Optimizer as follows:

From Release

To install from release (PyPi package):

python3 -m venv fms_mo_venv
source fms_mo_venv/bin/activate
pip install fms-model-optimizer

From Source

To install from source(GitHub Repository):

python3 -m venv fms_mo_venv
source fms_mo_venv/bin/activate
git clone https://github.com/foundation-model-stack/fms-model-optimizer
cd fms-model-optimizer
pip install -e .

Optional Dependencies

The following optional dependencies are available:

  • fp8: llmcompressor and torchao packages for fp8 quantization and inference
  • fp8-infer: torchao package for fp8 inference
  • gptq: GPTQModel package for W4A16 quantization
  • mx: microxcaling package for MX quantization
  • opt: Shortcut for fp8, gptq, and mx installs
  • aiu: ibm-fms package for AIU model deployment
  • torchvision: torch package for image recognition training and inference
  • triton: triton package for matrix multiplication kernels
  • examples: Dependencies needed for examples
  • visualize: Dependencies for visualizing models and performance data
  • test: Dependencies needed for unit testing
  • dev: Dependencies needed for development

To install an optional dependency, modify the pip install commands above with a list of these names enclosed in brackets. The example below installs llm-compressor and torchvision with FMS Model Optimizer:

pip install fms-model-optimizer[fp8,torchvision]

pip install -e .[fp8,torchvision]

If you have already installed FMS Model Optimizer, then only the optional packages will be installed.

Try It Out!

To help you get up and running as quickly as possible with the FMS Model Optimizer framework, check out the following resources which demonstrate how to use the framework with different quantization techniques:

  • Jupyter notebook tutorials (It is recommended to begin here):
    • Quantization tutorial:
      • Visualizes a random Gaussian tensor step-by-step along the quantization process
      • Build a quantizer and quantized convolution module based on this process
  • Python script examples

Docs

Dive into the design document to get a better understanding of the framework motivation and concepts.

Contributing

Check out our contributing guide to learn how to contribute.

Release files for fms-model-optimizer 0.8.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fms-model-optimizer 0.8.6
File Size Uploaded
fms_model_optimizer-0.8.6.tar.gz 5.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for fms-model-optimizer 0.8.6
File Interpreter ABI Platform
fms_model_optimizer-0.8.6-py3-none-any.whl Python 3 none any Details

Total release size: 5.6 MB

Release files / fms_model_optimizer-0.8.6.tar.gz

Download URL fms_model_optimizer-0.8.6.tar.gz
Size 5.2 MB
Tags Source
SHA-256 checksum
How to use checksums
f939b3c1ce955fe0be93f76628199005d1303a5608d5cb468e3dd09a51d66a8d
BLAKE2b-256 checksum
How to use checksums
16203eae9ea4eaf363c8e46d5e8aa863c58bfda78a13b1bf4aded43b2c080d83
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / fms_model_optimizer-0.8.6-py3-none-any.whl

Download URL fms_model_optimizer-0.8.6-py3-none-any.whl
Size 364.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
eeeb016b4cf8530ce2686809904a9aa805d03707b1a3896b5ef8d1489650ef4f
BLAKE2b-256 checksum
How to use checksums
fe64ec4b9c8fa49dc05dc3761239e535acea2ae6463dc691e2c1669bb7bd530d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.8.6 This release

2 release files

0.8.5

2 release files

0.8.4

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page