TensorRT Edge-LLM
High-Performance Large Language Model Inference Framework for NVIDIA Edge Platforms
Overview | Support Matrix | Quick Start | Performance | Documentation | Roadmap
Latest News
- [2026/09] TensorRT Edge-LLM completes the MLPerf Edge Agentic benchmark 6.4x faster on Jetson AGX Thor.
- [2026/09] Release 0.11.0 adds published Python wheels, Muse-Glimmer, experimental pi0.5, Hunyuan dense models, guided decoding, opt-in in-flight batching, native provider Jinja chat templates, and GPU-free CuTe DSL artifact generation.
- [2026/09] Release 0.10.1 adds experimental TP=2 inference on Dual NVIDIA DGX Spark and redesigns the experimental OpenAI-compatible server for faster cold launches and lower memory usage.
- [2026/08] TensorRT Edge-LLM 0.10.0 adds Day-0 support for Qwen3.8-27B.
- [2026/08] Release 0.10.0 adds support for NVIDIA Nemotron-3.5 Lightning with MTP and DFlash, Cosmos3-Edge, DiffusionGemma, Nemotron-3.5-ASR, and DSpark speculative decoding, alongside an experimental direct TensorRT engine builder without ONNX export, multi-turn KV-cache reuse, and video input for the experimental OpenAI-compatible server.
- [2026/07] Support for the full Gemma 4 family (E2B / E4B / 12B / 26B-A4B / 31B — multimodal text + image + audio, with MTP), Qwen3-Omni and Nemotron-3 NVFP4, and DFlash speculative decoding (with DDTree for Qwen3 / Qwen3.5) landed across releases 0.9.0 and 0.9.1.
Overview
TensorRT Edge-LLM is NVIDIA's C++ inference runtime for text, vision, audio, speech, and action models on NVIDIA Jetson, NVIDIA DRIVE, and NVIDIA DGX Spark. The supported frontend exports Hugging Face checkpoints to ONNX for C++ engine building; an experimental direct frontend builds engines from checkpoints without ONNX. Both paths use the same C++ deployment runtimes.
Getting Started
Check the Official Support Matrix, then follow the Quick Start Guide. Checkpoint IDs are listed in Supported Models.
For a supported target, install a published Python wheel
without compiling Edge-LLM. Use tensorrt-edgellm[server] for the high-level
Python API and HTTP serving; see the extras guide
for export/tools dependencies and the minimal base workflow.
Documentation
Introduction
- Overview - What is TensorRT Edge-LLM and key features
- Official Support Matrix - Platform, JetPack, DriveOS, CUDA, TensorRT, and TensorRT Edge-LLM compatibility
- Supported Models - Complete model compatibility matrix
- Checkpoint Exporter - Recommended ONNX export pipeline
- Experimental Direct Engine Builder - Build all model components directly from a checkpoint
User Guide
- Installation - Set up quantization,
tensorrt_edgellm, and the C++ runtime - Quick Start Guide - Run your first inference in ~15 minutes
- Examples - End-to-end workflows
- Quantization - Create quantized checkpoints for
tensorrt_edgellm - Experimental High-Level Python API and Server - vLLM-style API and OpenAI-compatible server
- Input Format Guide - Request format and specifications
- Chat Template Format - Chat template configuration
Developer Guide
Software Design
- Quantization Package Design - Quantization package architecture
- Engine Builder - Building TensorRT engines
- C++ Runtime Overview - Runtime system architecture
Advanced Topics
- Customization Guide - Customizing TensorRT Edge-LLM for your needs
- TensorRT Plugins - Custom plugin development
- Tests - Comprehensive test suite for contributors
Performance
See the Performance Benchmarks page for released benchmark results covering LLM and VLM prefill, generation throughput, memory usage, and EAGLE speculative decoding speedups.
Use Cases
🚗 Automotive
- In-vehicle AI assistants
- Voice-controlled interfaces
- Scene understanding
- Driver assistance systems
🤖 Robotics
- Natural language interaction
- Task planning and reasoning
- Visual question answering
- Human-robot collaboration
🏭 Industrial IoT
- Equipment monitoring with NLP
- Automated inspection
- Predictive maintenance
- Voice-controlled machinery
📱 Edge Devices
- On-device chatbots
- Offline language processing
- Privacy-preserving AI
- Low-latency inference
Featured Websites
- TensorRT Edge-LLM Jetson AI Lab tutorial
- Maximizing Memory Efficiency to Run Bigger Models on NVIDIA Jetson
- Build Next-Gen Physical AI with Edge-First LLMs for Autonomous Vehicles and Robotics
- Accelerate AI Inference for Edge and Robotics with NVIDIA Jetson T4000 and NVIDIA JetPack 7.1
- Accelerating LLM and VLM Inference for Automotive and Robotics with NVIDIA TensorRT Edge-LLM
Follow our GitHub repository for the latest updates, releases, and announcements.
Support
- Documentation: Full Documentation
- Quick Start: Quick Start Guide
- Roadmap: Developer Roadmap
- Issues: GitHub Issues
- Discussions: GitHub Discussions
- Forums: NVIDIA Developer Forums
License
Contributing
We welcome contributions! Please see our Contributing Guidelines for details.
Release files for tensorrt-edgellm 0.11.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| tensorrt_edgellm-0.11.0-cp312-cp312-manylinux_2_39_aarch64.whl | CPython 3.12 | CPython 3.12 | Linux glibc 2.39+ ARM64 | Details |
| tensorrt_edgellm-0.11.0-cp312-cp312-manylinux_2_35_x86_64.whl | CPython 3.12 | CPython 3.12 | Linux glibc 2.35+ x86-64 | Details |
| tensorrt_edgellm-0.11.0-cp311-cp311-manylinux_2_39_aarch64.whl | CPython 3.11 | CPython 3.11 | Linux glibc 2.39+ ARM64 | Details |
| tensorrt_edgellm-0.11.0-cp311-cp311-manylinux_2_35_x86_64.whl | CPython 3.11 | CPython 3.11 | Linux glibc 2.35+ x86-64 | Details |
| tensorrt_edgellm-0.11.0-cp310-cp310-manylinux_2_39_aarch64.whl | CPython 3.10 | CPython 3.10 | Linux glibc 2.39+ ARM64 | Details |
| tensorrt_edgellm-0.11.0-cp310-cp310-manylinux_2_35_x86_64.whl | CPython 3.10 | CPython 3.10 | Linux glibc 2.35+ x86-64 | Details |
Total release size: 1.9 GB
Release files / tensorrt_edgellm-0.11.0-cp312-cp312-manylinux_2_39_aarch64.whl
| Download URL | tensorrt_edgellm-0.11.0-cp312-cp312-manylinux_2_39_aarch64.whl |
|---|---|
| Size | 447.1 MB |
| Tags | CPython 3.12 Linux glibc 2.39+ ARM64 |
|
SHA-256 checksum How to use checksums |
9e6722a5a250d0ba18665911a908aecfc9684722b9d56d944efe31fe6c889d9b
|
|
BLAKE2b-256 checksum How to use checksums |
c20732efeb227b4b4c952b051ab9b990b21146756296b5ee65450de827cc5072
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / tensorrt_edgellm-0.11.0-cp312-cp312-manylinux_2_35_x86_64.whl
| Download URL | tensorrt_edgellm-0.11.0-cp312-cp312-manylinux_2_35_x86_64.whl |
|---|---|
| Size | 200.6 MB |
| Tags | CPython 3.12 Linux glibc 2.35+ x86-64 |
|
SHA-256 checksum How to use checksums |
f3ea7a0f96c52a12d80b84d42f6f9b1902d2a0018f0d3156370a62a84583801c
|
|
BLAKE2b-256 checksum How to use checksums |
af5a8ca441d785b91b4168f42b0aa43647250364b23e48c6a72d694ba1b905f3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / tensorrt_edgellm-0.11.0-cp311-cp311-manylinux_2_39_aarch64.whl
| Download URL | tensorrt_edgellm-0.11.0-cp311-cp311-manylinux_2_39_aarch64.whl |
|---|---|
| Size | 447.1 MB |
| Tags | CPython 3.11 Linux glibc 2.39+ ARM64 |
|
SHA-256 checksum How to use checksums |
21a66a57de5264bcf24e3e699e3e7b660e9541409d978cb5e76fe92e890b9c33
|
|
BLAKE2b-256 checksum How to use checksums |
c01f2b46c643045322cf06b1f895b1111fc4f5eeb3fc38c025762d35773311ea
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / tensorrt_edgellm-0.11.0-cp311-cp311-manylinux_2_35_x86_64.whl
| Download URL | tensorrt_edgellm-0.11.0-cp311-cp311-manylinux_2_35_x86_64.whl |
|---|---|
| Size | 200.6 MB |
| Tags | CPython 3.11 Linux glibc 2.35+ x86-64 |
|
SHA-256 checksum How to use checksums |
2cef657ced3a30b6d03d5c4b1ad511a707f77b64d7f59d5e1dd117de8ceb8f50
|
|
BLAKE2b-256 checksum How to use checksums |
00e3702017cbcd0907b137ae1bb87ef77a49e02347ad818832f5f91ec6e81eb3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / tensorrt_edgellm-0.11.0-cp310-cp310-manylinux_2_39_aarch64.whl
| Download URL | tensorrt_edgellm-0.11.0-cp310-cp310-manylinux_2_39_aarch64.whl |
|---|---|
| Size | 447.1 MB |
| Tags | CPython 3.10 Linux glibc 2.39+ ARM64 |
|
SHA-256 checksum How to use checksums |
9bb4383ef013a2108d2c077c8758a1a1ecbb2995ac9450ae02d5611dcecba611
|
|
BLAKE2b-256 checksum How to use checksums |
908b2a8e32f3fd7bf25446eb777ffd0bc8f8583b699a0de2eb3b5bc599cc45be
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / tensorrt_edgellm-0.11.0-cp310-cp310-manylinux_2_35_x86_64.whl
| Download URL | tensorrt_edgellm-0.11.0-cp310-cp310-manylinux_2_35_x86_64.whl |
|---|---|
| Size | 200.6 MB |
| Tags | CPython 3.10 Linux glibc 2.35+ x86-64 |
|
SHA-256 checksum How to use checksums |
73c407eb625ab4382dd0b380bd51a6d08d4a3d3b18ec34154187dfc42b3b98fe
|
|
BLAKE2b-256 checksum How to use checksums |
aca119927da3d9fb30f84deb83b3e59237a2642cbd93b5b38ea04f18759963c6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|