Skip to main content

Latest News 🎉

2026
2025
  • [2025/09] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the performance of vLLM on H800 for openai gpt-oss models!
  • [2025/06] Comprehensive inference optimization for FP8 MoE Models
  • [2025/06] DeepSeek PD Disaggregation deployment is now supported through integration with DLSlime and Mooncake. Huge thanks to both teams!
  • [2025/04] Enhance DeepSeek inference performance by integrating deepseek-ai techniques: FlashMLA, DeepGemm, DeepEP, MicroBatch and eplb
  • [2025/01] Support DeepSeek V3 and R1
2024
  • [2024/11] Support Mono-InternVL with PyTorch engine
  • [2024/10] PyTorchEngine supports graph mode on ascend platform, doubling the inference speed
  • [2024/09] LMDeploy PyTorchEngine adds support for Huawei Ascend. See supported models here
  • [2024/09] LMDeploy PyTorchEngine achieves 1.3x faster on Llama3-8B inference by introducing CUDA graph
  • [2024/08] LMDeploy is integrated into modelscope/swift as the default accelerator for VLMs inference
  • [2024/07] Support Llama3.1 8B, 70B and its TOOLS CALLING
  • [2024/07] Support InternVL2 full-series models, InternLM-XComposer2.5 and function call of InternLM2.5
  • [2024/06] PyTorch engine support DeepSeek-V2 and several VLMs, such as CogVLM2, Mini-InternVL, LlaVA-Next
  • [2024/05] Balance vision model when deploying VLMs with multiple GPUs
  • [2024/05] Support 4-bit weight-only quantization and inference on VLMs, such as InternVL v1.5, LLaVa, InternLMXComposer2
  • [2024/04] Support Llama3 and more VLMs, such as InternVL v1.1, v1.2, MiniGemini, InternLMXComposer2.
  • [2024/04] TurboMind adds online int8/int4 KV cache quantization and inference for all supported devices. Refer here for detailed guide
  • [2024/04] TurboMind latest upgrade boosts GQA, rocketing the internlm2-20b model inference to 16+ RPS, about 1.8x faster than vLLM.
  • [2024/04] Support Qwen1.5-MOE and dbrx.
  • [2024/03] Support DeepSeek-VL offline inference pipeline and serving.
  • [2024/03] Support VLM offline inference pipeline and serving.
  • [2024/02] Support Qwen 1.5, Gemma, Mistral, Mixtral, Deepseek-MOE and so on.
  • [2024/01] OpenAOE seamless integration with LMDeploy Serving Service.
  • [2024/01] Support for multi-model, multi-machine, multi-card inference services. For usage instructions, please refer to here
  • [2024/01] Support PyTorch inference engine, developed entirely in Python, helping to lower the barriers for developers and enable rapid experimentation with new features and technologies.
2023
  • [2023/12] Turbomind supports multimodal input.
  • [2023/11] Turbomind supports loading hf model directly. Click here for details.
  • [2023/11] TurboMind major upgrades, including: Paged Attention, faster attention kernels without sequence length limitation, 2x faster KV8 kernels, Split-K decoding (Flash Decoding), and W4A16 inference for sm_75
  • [2023/09] TurboMind supports Qwen-14B
  • [2023/09] TurboMind supports InternLM-20B
  • [2023/09] TurboMind supports all features of Code Llama: code completion, infilling, chat / instruct, and python specialist. Click here for deployment guide
  • [2023/09] TurboMind supports Baichuan2-7B
  • [2023/08] TurboMind supports flash-attention2.
  • [2023/08] TurboMind supports Qwen-7B, dynamic NTK-RoPE scaling and dynamic logN scaling
  • [2023/08] TurboMind supports Windows (tp=1)
  • [2023/08] TurboMind supports 4-bit inference, 2.4x faster than FP16, the fastest open-source implementation. Check this guide for detailed info
  • [2023/08] LMDeploy has launched on the HuggingFace Hub, providing ready-to-use 4-bit models.
  • [2023/08] LMDeploy supports 4-bit quantization using the AWQ algorithm.
  • [2023/07] TurboMind supports Llama-2 70B with GQA.
  • [2023/07] TurboMind supports Llama-2 7B/13B.
  • [2023/07] TurboMind supports tensor-parallel inference of InternLM.

Introduction

LMDeploy is a toolkit for compressing, deploying, and serving LLM, developed by the MMRazor and MMDeploy teams. It has the following core features:

  • Efficient Inference: LMDeploy delivers up to 1.8x higher request throughput than vLLM, by introducing key features like persistent batch(a.k.a. continuous batching), blocked KV cache, dynamic split&fuse, tensor parallelism, high-performance CUDA kernels and so on.

  • Effective Quantization: LMDeploy supports weight-only and k/v quantization, and the 4-bit inference performance is 2.4x higher than FP16. The quantization quality has been confirmed via OpenCompass evaluation.

  • Effortless Distribution Server: Leveraging the request distribution service, LMDeploy facilitates an easy and efficient deployment of multi-model services across multiple machines and cards.

  • Excellent Compatibility: LMDeploy supports KV Cache Quant, AWQ and Automatic Prefix Caching to be used simultaneously.

Performance

v0 1 0-benchmark

Supported Models

LLMs VLMs
  • Llama (7B - 65B)
  • Llama2 (7B - 70B)
  • Llama3 (8B, 70B)
  • Llama3.1 (8B, 70B)
  • Llama3.2 (1B, 3B)
  • InternLM2 (7B - 20B)
  • InternLM3 (8B)
  • InternLM2.5 (7B)
  • Qwen1.5 (0.5B - 110B)
  • Qwen1.5 - MoE (0.5B - 72B)
  • Qwen2 (0.5B - 72B)
  • Qwen2-MoE (57BA14B)
  • Qwen2.5 (0.5B - 32B)
  • Qwen3, Qwen3-MoE
  • Qwen3-Next(80B)
  • Code Llama (7B - 34B)
  • ChatGLM2 (6B)
  • GLM-4 (9B)
  • GLM-4-0414 (9B, 32B)
  • CodeGeeX4 (9B)
  • YI (6B-34B)
  • Mistral (7B)
  • DeepSeek-MoE (16B)
  • DeepSeek-V2 (16B, 236B)
  • DeepSeek-V2.5 (236B)
  • DeepSeek-V3 (685B)
  • DeepSeek-V3.2 (685B)
  • DeepSeek-V4 (284B, 1.6T)
  • Hy3 (295B-A21B)
  • Mixtral (8x7B, 8x22B)
  • Gemma (2B - 7B)
  • Phi-3-mini (3.8B)
  • Phi-3.5-mini (3.8B)
  • Phi-3.5-MoE (16x3.8B)
  • Phi-4-mini (3.8B)
  • MiniCPM3 (4B)
  • SDAR (1.7B-30B)
  • gpt-oss (20B, 120B)
  • GLM-4.7-Flash (30B)
  • GLM-5 (754B)
  • GLM-5.2 (754B)
  • LLaVA(1.5,1.6) (7B-34B)
  • Qwen2-VL (2B, 7B, 72B)
  • Qwen2.5-VL (3B, 7B, 72B)
  • Qwen3-VL (2B - 235B)
  • Qwen3.5 (0.8B - 397B)
  • Qwen3-Omni (30B-A3B)
  • Kimi-K2.6 (1T-A32B)
  • DeepSeek-VL (7B)
  • DeepSeek-VL2 (3B, 16B, 27B)
  • InternVL-Chat (v1.1-v1.5)
  • InternVL2 (1B-76B)
  • InternVL2.5(MPO) (1B-78B)
  • InternVL3 (1B-78B)
  • InternVL3.5 (1B-241BA28B)
  • Intern-S1 (241B)
  • Intern-S1-mini (8.3B)
  • Intern-S1-Pro (1TB)
  • Intern-S2-Preview (35B-A3B, 397B)
  • Intern-S2-Mobius (35B)
  • ChemVLM (8B-26B)
  • CogVLM-Chat (17B)
  • CogVLM2-Chat (19B)
  • MiniCPM-Llama3-V-2_5
  • MiniCPM-V-2_6
  • Phi-3-vision (4.2B)
  • Phi-3.5-vision (4.2B)
  • GLM-4V (9B)
  • GLM-4.1V-Thinking (9B)
  • Molmo (7B-D,72B)
  • Gemma3 (1B - 27B)
  • Llama4 (Scout, Maverick)

LMDeploy has developed two inference engines - TurboMind and PyTorch, each with a different focus. The former strives for ultimate optimization of inference performance, while the latter, developed purely in Python, aims to decrease the barriers for developers.

They differ in the types of supported models and the inference data type. Please refer to this table for each engine's capability and choose the proper one that best fits your actual needs.

Quick Start Open In Colab

Installation

It is recommended to install lmdeploy using pip in a conda environment (python 3.10 - 3.13):

conda create -n lmdeploy python=3.12 -y
conda activate lmdeploy
pip install lmdeploy

Starting from v0.13.0, the default prebuilt wheels published on PyPI are built against CUDA 12.8, so pip install lmdeploy is sufficient for typical setups including GeForce RTX 50 series.

Offline Batch Inference

import lmdeploy
with lmdeploy.pipeline("internlm/internlm3-8b-instruct") as pipe:
    response = pipe(["Hi, pls intro yourself", "Shanghai is"])
    print(response)

[!NOTE] By default, LMDeploy downloads model from HuggingFace. If you would like to use models from ModelScope, please install ModelScope by pip install modelscope and set the environment variable:

export LMDEPLOY_USE_MODELSCOPE=True

If you would like to use models from openMind Hub, please install openMind Hub by pip install openmind_hub and set the environment variable:

export LMDEPLOY_USE_OPENMIND_HUB=True

For more information about inference pipeline, please refer to here.

Tutorials

Please review getting_started section for the basic usage of LMDeploy.

For detailed user guides and advanced guides, please refer to our tutorials:

Third-party projects

  • Deploying LLMs offline on the NVIDIA Jetson platform by LMDeploy: LMDeploy-Jetson

  • Example project for deploying LLMs using LMDeploy and BentoML: BentoLMDeploy

Contributing

We appreciate all contributions to LMDeploy. Please refer to CONTRIBUTING.md for the contributing guideline.

Acknowledgement

Citation

@misc{2023lmdeploy,
    title={LMDeploy: A Toolkit for Compressing, Deploying, and Serving LLM},
    author={LMDeploy Contributors},
    howpublished = {\url{https://github.com/InternLM/lmdeploy}},
    year={2023}
}
@article{zhang2025lmdeploy,
  title={LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind},
  author={Zhang, Li and Jiang, Youhe and He, Guoliang and Chen, Xin and Lv, Han and Yao, Qian and Ma, Ningsheng and Fu, Fangcheng and Chen, Kai},
  journal={arXiv preprint arXiv:2508.15601},
  year={2025}
}

License

This project is released under the Apache 2.0 license.

Release files for lmdeploy 0.18.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for lmdeploy 0.18.0
File
lmdeploy-0.18.0-cp314-cp314-win_amd64.whl CPython 3.14 CPython 3.14 Windows x86-64 Details
lmdeploy-0.18.0-cp314-cp314-manylinux_2_28_x86_64.whl CPython 3.14 CPython 3.14 Linux glibc 2.28+ x86-64 Details
lmdeploy-0.18.0-cp313-cp313-win_amd64.whl CPython 3.13 CPython 3.13 Windows x86-64 Details
lmdeploy-0.18.0-cp313-cp313-manylinux_2_28_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.28+ x86-64 Details
lmdeploy-0.18.0-cp312-cp312-win_amd64.whl CPython 3.12 CPython 3.12 Windows x86-64 Details
lmdeploy-0.18.0-cp312-cp312-manylinux_2_28_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.28+ x86-64 Details
lmdeploy-0.18.0-cp311-cp311-win_amd64.whl CPython 3.11 CPython 3.11 Windows x86-64 Details
lmdeploy-0.18.0-cp311-cp311-manylinux_2_28_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.28+ x86-64 Details
lmdeploy-0.18.0-cp310-cp310-win_amd64.whl CPython 3.10 CPython 3.10 Windows x86-64 Details
lmdeploy-0.18.0-cp310-cp310-manylinux_2_28_x86_64.whl CPython 3.10 CPython 3.10 Linux glibc 2.28+ x86-64 Details

Total release size: 827.9 MB

Release files / lmdeploy-0.18.0-cp314-cp314-win_amd64.whl

Download URL lmdeploy-0.18.0-cp314-cp314-win_amd64.whl
Size 59.2 MB
Tags CPython 3.14 Windows x86-64
SHA-256 checksum
How to use checksums
34e2b6c772a68aaf45481e86015e5f341f0bc61c857f2138f211c76f87cdb15d
BLAKE2b-256 checksum
How to use checksums
3a3b5aa6d73c6c4b6b0dc2a26fedafd6fdf3d41cd6bc1655730db75891a11d54
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp314-cp314-manylinux_2_28_x86_64.whl

Download URL lmdeploy-0.18.0-cp314-cp314-manylinux_2_28_x86_64.whl
Size 106.7 MB
Tags CPython 3.14 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
629a136c7ca532e1251c3b6609b2c8b4c8b6d748042a7d5f4214cb4d4a581ec3
BLAKE2b-256 checksum
How to use checksums
0e0dc3ec6984ad7c982ed36ab3214e75c094aff06838734fc5ab22da3b8e65b5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp313-cp313-win_amd64.whl

Download URL lmdeploy-0.18.0-cp313-cp313-win_amd64.whl
Size 58.8 MB
Tags CPython 3.13 Windows x86-64
SHA-256 checksum
How to use checksums
a2de45d2d292871795d50b22d9cb4185983768783abde58ff376e2072da51893
BLAKE2b-256 checksum
How to use checksums
2b61f223a174cc2db4b9eb3eee6f947597366890df58e4404f7f83f60f71fa0c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp313-cp313-manylinux_2_28_x86_64.whl

Download URL lmdeploy-0.18.0-cp313-cp313-manylinux_2_28_x86_64.whl
Size 106.7 MB
Tags CPython 3.13 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
e5f727a2b61e88aa10a9902592d6cd08e13fb41d695dda4c007b6cc8f7f568cb
BLAKE2b-256 checksum
How to use checksums
af994bdea6d1933d85e85f6c1191370dc6064b04fc732a03eabca3845ecabf3c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp312-cp312-win_amd64.whl

Download URL lmdeploy-0.18.0-cp312-cp312-win_amd64.whl
Size 58.8 MB
Tags CPython 3.12 Windows x86-64
SHA-256 checksum
How to use checksums
26c974dcedf547d0e868998002bda691575932ae77c65e0cb9b027946610603e
BLAKE2b-256 checksum
How to use checksums
77d0e2642c425be45809bb2059e77ae7aefdb503292cd9bd04d73b510552d320
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp312-cp312-manylinux_2_28_x86_64.whl

Download URL lmdeploy-0.18.0-cp312-cp312-manylinux_2_28_x86_64.whl
Size 106.7 MB
Tags CPython 3.12 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
fcef7cf0d210d766e02e2e6df025378d2f8ae36080400b2818eb9a34ef3392dc
BLAKE2b-256 checksum
How to use checksums
198db880e02991ce03947c1efa9504c540d409b06d5d7ea4f95b1d8f8ae85c99
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp311-cp311-win_amd64.whl

Download URL lmdeploy-0.18.0-cp311-cp311-win_amd64.whl
Size 58.8 MB
Tags CPython 3.11 Windows x86-64
SHA-256 checksum
How to use checksums
31d21f6dd1c7b914fb622971f9a8fd01f6856c1dabab650140f8ad1182c376ee
BLAKE2b-256 checksum
How to use checksums
c88d70aecf94d6d062bfe52f79d18ac4d5427be4c1614d398e409b114d4470ee
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp311-cp311-manylinux_2_28_x86_64.whl

Download URL lmdeploy-0.18.0-cp311-cp311-manylinux_2_28_x86_64.whl
Size 106.7 MB
Tags CPython 3.11 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
e6790a25dcdb0345c160567fdf2ac604fec319389523a8b806834c9d92c8ffe1
BLAKE2b-256 checksum
How to use checksums
f636bb00cf18ebe16237b17065adeb6720886f054e95975443d5112fec312dd8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp310-cp310-win_amd64.whl

Download URL lmdeploy-0.18.0-cp310-cp310-win_amd64.whl
Size 58.8 MB
Tags CPython 3.10 Windows x86-64
SHA-256 checksum
How to use checksums
42398e1f3b7139b164997b954d0739cd9fb8b79001ce6540c0a593abf5e02467
BLAKE2b-256 checksum
How to use checksums
59bcddd505e20069c5549a14ee7eeea17f2e1336ef06de12cb7e2319b12cdf29
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release files / lmdeploy-0.18.0-cp310-cp310-manylinux_2_28_x86_64.whl

Download URL lmdeploy-0.18.0-cp310-cp310-manylinux_2_28_x86_64.whl
Size 106.7 MB
Tags CPython 3.10 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
529a131a98b2fafeb4f9b3b10b0bad2b5d6cd48931adc9d93da15fd1671803bb
BLAKE2b-256 checksum
How to use checksums
27367f7218c3cbef553eb6cff4785f3a2f5e80ddc658727cf0cb61255841e8e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.21

Release history Release notifications | RSS feed

This release

0.18.0 This release

10 release files

0.14.0

8 release files

0.13.0

8 release files

0.12.3

8 release files

0.12.2

4 release files

0.12.1

8 release files

0.11.1

8 release files

0.9.2

10 release files

0.9.0

10 release files

0.7.3

10 release files

0.7.2

10 release files

0.7.1

10 release files

0.7.0

10 release files

0.6.5

10 release files

0.6.3

10 release files

0.6.2

10 release files

0.6.1

10 release files

0.6.0

10 release files

0.5.2

10 release files

0.5.1

10 release files

0.4.2

10 release files

0.4.1

8 release files

0.4.0

8 release files

0.3.0

8 release files

0.2.6

8 release files

0.2.5

8 release files

0.2.4

8 release files

0.2.3

8 release files

0.2.2

8 release files

0.2.1

8 release files

0.2.0

8 release files

0.1.0

8 release files

0.0.13

8 release files

0.0.12

8 release files

0.0.11

8 release files

0.0.10

8 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page