Skip to main content

NeMo Export-Deploy

codecov CI/CD Python 3.10+ GitHub Stars

📖 Documentation 🔧 Installation 🚀 Quick start 🤝 Contributing

The Export-Deploy library ("NeMo Export-Deploy") provides tools and APIs for exporting and deploying NeMo and Hugging Face models to production environments. It supports various deployment paths including TensorRT and vLLM deployment through NVIDIA Triton Inference Server and Ray Serve.

image

📣 News

  • [03/12/2026] Deprecating Python 3.10 support: We're officially dropping Python 3.10 support with the upcoming 0.4.0 release. Downstream applications must raise their lower boundary to 3.12 to stay compatible with Export-Deploy.

🚀 Key Features

  • Support for Large Language Models (LLMs) and Multimodal Models (MMs)
  • Export Megatron-Bridge and Hugging Face models to optimized inference formats including vLLM
  • Deploy Megatron-Bridge, Megatron-LM and Hugging Face models using Ray Serve or NVIDIA Triton Inference Server
  • Multi-GPU and distributed inference capabilities
  • Multi-instance deployment options

Feature Support Matrix

Model Export Capabilities

Model / Checkpoint vLLM ONNX TensorRT
Megatron Bridge bf16 N/A N/A
Hugging Face bf16 N/A N/A
NIM Embedding N/A bf16, fp8, int8 (PTQ) bf16, fp8, int8 (PTQ)
NIM Reranking N/A bf16, fp8, int8 (PTQ) bf16, fp8, int8 (PTQ)

The support matrix above outlines the export capabilities for each model or checkpoint, including the supported precision options across various inference-optimized libraries. The export module enables exporting models that have been quantized using post-training quantization (PTQ) with the TensorRT Model Optimizer library, as shown above. Models trained with low precision or quantization-aware training are also supported, as indicated in the table.

The inference-optimized libraries listed above also support on-the-fly quantization during model export, with configurable parameters available in the export APIs. However, please note that the precision options shown in the table above indicate support for exporting models that have already been quantized, rather than the ability to quantize models during export.

Please note that not all large language models (LLMs) and multimodal models (MMs) are currently supported. For the most complete and up-to-date information, please refer to the LLM documentation and MM documentation.

Model Deployment Capabilities

Model / Checkpoint RayServe PyTriton
Megatron Bridge Single/Multi-Node Multi-GPU Single-Node Multi-GPU
Megatron-LM Limited Limited
Hugging Face Single-Node Multi-GPU,
Multi-instance
Single-Node Multi-GPU
vLLM N/A Single-Node Multi-GPU

The support matrix above outlines the available deployment options for each model or checkpoint, highlighting multi-node and multi-GPU support where applicable. For comprehensive details, please refer to the documentation.

Refer to the table below for an overview of optimized inference and deployment support for NeMo Framework and Hugging Face models with Triton Inference Server.

Model / Checkpoint vLLM + Triton Inference Server Direct Triton Inference Server
Hugging Face ☑ ☑

🔧 Install

For quick exploration of NeMo Export-Deploy, we recommend installing our pip package:

pip install nemo-export-deploy

This installation comes without extra dependencies like TransformerEngine or vLLM. The installation serves for navigating around and for exploring the project.

For a feature-complete install, please refer to the following sections.

Use NeMo-FW Container

Best experience, highest performance and full feature support is guaranteed by the NeMo Framework container. Please fetch the most recent $TAG and run the following command to start a container:

docker run --rm -it -w /workdir -v $(pwd):/workdir \
  --entrypoint bash \
  --gpus all \
  nvcr.io/nvidia/nemo:${TAG}

Build with Dockerfile

For containerized development, use our Dockerfile for building your own container. There are two flavors: INFERENCE_FRAMEWORK=inframework and INFERENCE_FRAMEWORK=vllm:

docker build \
    -f docker/Dockerfile.pytorch \
    -t nemo-export-deploy \
    --build-arg INFERENCE_FRAMEWORK=$INFERENCE_FRAMEWORK \
    .

Start your container:

docker run --rm -it -w /workdir -v $(pwd):/workdir \
  --entrypoint bash \
  --gpus all \
  nemo-export-deploy

Install from Source

For complete feature coverage, we recommend to install TransformerEngine and additionally vLLM.

Recommended Requirements

  • Python 3.12
  • PyTorch 2.7
  • CUDA 12.9
  • Ubuntu 24.04

Install TransformerEngine + InFramework

For highly optimized TransformerEngine path with PyTriton backend, please make sure to install the following prerequisites first:

pip install torch==2.7.0 setuptools pybind11 wheel_stub  # Required for TE

Now proceed with the main installation:

git clone https://github.com/NVIDIA-NeMo/Export-Deploy
cd Export-Deploy/
pip install --no-build-isolation .

Install TransformerEngine + vLLM

For highly optimized TransformerEngine path with vLLM backend, please make sure to install the following prerequisites first:

pip install torch==2.7.0 setuptools pybind11 wheel_stub  # Required for TE

Now proceed with the main installation:

pip install --no-build-isolation .[vllm]

Install TransformerEngine + TRT-ONNX

For highly optimized TransformerEngine path with TRT-ONNX backend, please make sure to install the following prerequisites first:

pip install torch==2.7.0 setuptools pybind11 wheel_stub  # Required for TE

Now proceed with the main installation:

pip install --no-build-isolation .[trt-onnx]

🤝 Contributing

We welcome contributions to NeMo Export-Deploy! Please see our Contributing Guidelines for more information on how to get involved.

License

NeMo Export-Deploy is licensed under the Apache License 2.0.

Release files for NeMo-Export-Deploy 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for NeMo-Export-Deploy 0.7.0
File Size Uploaded
nemo_export_deploy-0.7.0.tar.gz 102.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for NeMo-Export-Deploy 0.7.0
File Interpreter ABI Platform
nemo_export_deploy-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 237.7 kB

Release files / nemo_export_deploy-0.7.0.tar.gz

Download URL nemo_export_deploy-0.7.0.tar.gz
Size 102.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f8ba9620f3ea7fb2e399b18db075645d96408251532f42a4c107f500fde31588
BLAKE2b-256 checksum
How to use checksums
9942b5218740eb400e4720003d7a6fd5f5f141e1784d533b819e2093cfd1459f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.12.3

Release files / nemo_export_deploy-0.7.0-py3-none-any.whl

Download URL nemo_export_deploy-0.7.0-py3-none-any.whl
Size 135.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a946d20343f54410a171b4be163da987283874bd7085111abdec417de28ffc8a
BLAKE2b-256 checksum
How to use checksums
9c5333446f422ce07193b18039dcd7466b1f336c493b1e4413aa816ee7e5d300
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.12.3
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page