Skip to main content

Latest News 🎉

2026
  • [2026/04] PyPI has expanded the storage quota for LMDeploy and wheel uploads have resumed. v0.12.3 is now available on PyPI, so you can install it directly via pip install lmdeploy.
  • [2026/02] Support Qwen3.5
  • [2026/02] Support vllm-project/llm-compressor 4bit symmetric/asymmetric quantization. Refer here for detailed guide
2025
  • [2025/09] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the performmance of vLLM on H800 for openai gpt-oss models!
  • [2025/06] Comprehensive inference optimization for FP8 MoE Models
  • [2025/06] DeepSeek PD Disaggregation deployment is now supported through integration with DLSlime and Mooncake. Huge thanks to both teams!
  • [2025/04] Enhance DeepSeek inference performance by integration deepseek-ai techniques: FlashMLA, DeepGemm, DeepEP, MicroBatch and eplb
  • [2025/01] Support DeepSeek V3 and R1
2024
  • [2024/11] Support Mono-InternVL with PyTorch engine
  • [2024/10] PyTorchEngine supports graph mode on ascend platform, doubling the inference speed
  • [2024/09] LMDeploy PyTorchEngine adds support for Huawei Ascend. See supported models here
  • [2024/09] LMDeploy PyTorchEngine achieves 1.3x faster on Llama3-8B inference by introducing CUDA graph
  • [2024/08] LMDeploy is integrated into modelscope/swift as the default accelerator for VLMs inference
  • [2024/07] Support Llama3.1 8B, 70B and its TOOLS CALLING
  • [2024/07] Support InternVL2 full-series models, InternLM-XComposer2.5 and function call of InternLM2.5
  • [2024/06] PyTorch engine support DeepSeek-V2 and several VLMs, such as CogVLM2, Mini-InternVL, LlaVA-Next
  • [2024/05] Balance vision model when deploying VLMs with multiple GPUs
  • [2024/05] Support 4-bits weight-only quantization and inference on VLMs, such as InternVL v1.5, LLaVa, InternLMXComposer2
  • [2024/04] Support Llama3 and more VLMs, such as InternVL v1.1, v1.2, MiniGemini, InternLMXComposer2.
  • [2024/04] TurboMind adds online int8/int4 KV cache quantization and inference for all supported devices. Refer here for detailed guide
  • [2024/04] TurboMind latest upgrade boosts GQA, rocketing the internlm2-20b model inference to 16+ RPS, about 1.8x faster than vLLM.
  • [2024/04] Support Qwen1.5-MOE and dbrx.
  • [2024/03] Support DeepSeek-VL offline inference pipeline and serving.
  • [2024/03] Support VLM offline inference pipeline and serving.
  • [2024/02] Support Qwen 1.5, Gemma, Mistral, Mixtral, Deepseek-MOE and so on.
  • [2024/01] OpenAOE seamless integration with LMDeploy Serving Service.
  • [2024/01] Support for multi-model, multi-machine, multi-card inference services. For usage instructions, please refer to here
  • [2024/01] Support PyTorch inference engine, developed entirely in Python, helping to lower the barriers for developers and enable rapid experimentation with new features and technologies.
2023
  • [2023/12] Turbomind supports multimodal input.
  • [2023/11] Turbomind supports loading hf model directly. Click here for details.
  • [2023/11] TurboMind major upgrades, including: Paged Attention, faster attention kernels without sequence length limitation, 2x faster KV8 kernels, Split-K decoding (Flash Decoding), and W4A16 inference for sm_75
  • [2023/09] TurboMind supports Qwen-14B
  • [2023/09] TurboMind supports InternLM-20B
  • [2023/09] TurboMind supports all features of Code Llama: code completion, infilling, chat / instruct, and python specialist. Click here for deployment guide
  • [2023/09] TurboMind supports Baichuan2-7B
  • [2023/08] TurboMind supports flash-attention2.
  • [2023/08] TurboMind supports Qwen-7B, dynamic NTK-RoPE scaling and dynamic logN scaling
  • [2023/08] TurboMind supports Windows (tp=1)
  • [2023/08] TurboMind supports 4-bit inference, 2.4x faster than FP16, the fastest open-source implementation. Check this guide for detailed info
  • [2023/08] LMDeploy has launched on the HuggingFace Hub, providing ready-to-use 4-bit models.
  • [2023/08] LMDeploy supports 4-bit quantization using the AWQ algorithm.
  • [2023/07] TurboMind supports Llama-2 70B with GQA.
  • [2023/07] TurboMind supports Llama-2 7B/13B.
  • [2023/07] TurboMind supports tensor-parallel inference of InternLM.

Introduction

LMDeploy is a toolkit for compressing, deploying, and serving LLM, developed by the MMRazor and MMDeploy teams. It has the following core features:

  • Efficient Inference: LMDeploy delivers up to 1.8x higher request throughput than vLLM, by introducing key features like persistent batch(a.k.a. continuous batching), blocked KV cache, dynamic split&fuse, tensor parallelism, high-performance CUDA kernels and so on.

  • Effective Quantization: LMDeploy supports weight-only and k/v quantization, and the 4-bit inference performance is 2.4x higher than FP16. The quantization quality has been confirmed via OpenCompass evaluation.

  • Effortless Distribution Server: Leveraging the request distribution service, LMDeploy facilitates an easy and efficient deployment of multi-model services across multiple machines and cards.

  • Excellent Compatibility: LMDeploy supports KV Cache Quant, AWQ and Automatic Prefix Caching to be used simultaneously.

Performance

v0 1 0-benchmark

Supported Models

LLMs VLMs
  • Llama (7B - 65B)
  • Llama2 (7B - 70B)
  • Llama3 (8B, 70B)
  • Llama3.1 (8B, 70B)
  • Llama3.2 (1B, 3B)
  • InternLM2 (7B - 20B)
  • InternLM3 (8B)
  • InternLM2.5 (7B)
  • Qwen1.5 (0.5B - 110B)
  • Qwen1.5 - MoE (0.5B - 72B)
  • Qwen2 (0.5B - 72B)
  • Qwen2-MoE (57BA14B)
  • Qwen2.5 (0.5B - 32B)
  • Qwen3, Qwen3-MoE
  • Qwen3-Next(80B)
  • Code Llama (7B - 34B)
  • ChatGLM2 (6B)
  • GLM-4 (9B)
  • GLM-4-0414 (9B, 32B)
  • CodeGeeX4 (9B)
  • YI (6B-34B)
  • Mistral (7B)
  • DeepSeek-MoE (16B)
  • DeepSeek-V2 (16B, 236B)
  • DeepSeek-V2.5 (236B)
  • DeepSeek-V3 (685B)
  • DeepSeek-V3.2 (685B)
  • DeepSeek-V4 (284B, 1.6T)
  • Hy3 (295B-A21B)
  • Mixtral (8x7B, 8x22B)
  • Gemma (2B - 7B)
  • Phi-3-mini (3.8B)
  • Phi-3.5-mini (3.8B)
  • Phi-3.5-MoE (16x3.8B)
  • Phi-4-mini (3.8B)
  • MiniCPM3 (4B)
  • SDAR (1.7B-30B)
  • gpt-oss (20B, 120B)
  • GLM-4.7-Flash (30B)
  • GLM-5 (754B)
  • GLM-5.2 (754B)
  • LLaVA(1.5,1.6) (7B-34B)
  • Qwen2-VL (2B, 7B, 72B)
  • Qwen2.5-VL (3B, 7B, 72B)
  • Qwen3-VL (2B - 235B)
  • Qwen3.5 (0.8B - 397B)
  • Qwen3-Omni (30B-A3B)
  • DeepSeek-VL (7B)
  • DeepSeek-VL2 (3B, 16B, 27B)
  • InternVL-Chat (v1.1-v1.5)
  • InternVL2 (1B-76B)
  • InternVL2.5(MPO) (1B-78B)
  • InternVL3 (1B-78B)
  • InternVL3.5 (1B-241BA28B)
  • Intern-S1 (241B)
  • Intern-S1-mini (8.3B)
  • Intern-S1-Pro (1TB)
  • Intern-S2-Preview (35B-A3B, 397B)
  • Intern-S2-Mobius (35B)
  • ChemVLM (8B-26B)
  • CogVLM-Chat (17B)
  • CogVLM2-Chat (19B)
  • MiniCPM-Llama3-V-2_5
  • MiniCPM-V-2_6
  • Phi-3-vision (4.2B)
  • Phi-3.5-vision (4.2B)
  • GLM-4V (9B)
  • GLM-4.1V-Thinking (9B)
  • Molmo (7B-D,72B)
  • Gemma3 (1B - 27B)
  • Llama4 (Scout, Maverick)

LMDeploy has developed two inference engines - TurboMind and PyTorch, each with a different focus. The former strives for ultimate optimization of inference performance, while the latter, developed purely in Python, aims to decrease the barriers for developers.

They differ in the types of supported models and the inference data type. Please refer to this table for each engine's capability and choose the proper one that best fits your actual needs.

Quick Start Open In Colab

Installation

It is recommended installing lmdeploy using pip in a conda environment (python 3.10 - 3.13):

conda create -n lmdeploy python=3.12 -y
conda activate lmdeploy
pip install lmdeploy

Starting from v0.13.0, the default prebuilt wheels published on PyPI are built against CUDA 12.8, so pip install lmdeploy is sufficient for typical setups including GeForce RTX 50 series.

Offline Batch Inference

import lmdeploy
with lmdeploy.pipeline("internlm/internlm3-8b-instruct") as pipe:
    response = pipe(["Hi, pls intro yourself", "Shanghai is"])
    print(response)

[!NOTE] By default, LMDeploy downloads model from HuggingFace. If you would like to use models from ModelScope, please install ModelScope by pip install modelscope and set the environment variable:

export LMDEPLOY_USE_MODELSCOPE=True

If you would like to use models from openMind Hub, please install openMind Hub by pip install openmind_hub and set the environment variable:

export LMDEPLOY_USE_OPENMIND_HUB=True

For more information about inference pipeline, please refer to here.

Tutorials

Please review getting_started section for the basic usage of LMDeploy.

For detailed user guides and advanced guides, please refer to our tutorials:

Third-party projects

  • Deploying LLMs offline on the NVIDIA Jetson platform by LMDeploy: LMDeploy-Jetson

  • Example project for deploying LLMs using LMDeploy and BentoML: BentoLMDeploy

Contributing

We appreciate all contributions to LMDeploy. Please refer to CONTRIBUTING.md for the contributing guideline.

Acknowledgement

Citation

@misc{2023lmdeploy,
    title={LMDeploy: A Toolkit for Compressing, Deploying, and Serving LLM},
    author={LMDeploy Contributors},
    howpublished = {\url{https://github.com/InternLM/lmdeploy}},
    year={2023}
}
@article{zhang2025efficient,
  title={Efficient Mixed-Precision Large Language Model Inference with TurboMind},
  author={Zhang, Li and Jiang, Youhe and He, Guoliang and Chen, Xin and Lv, Han and Yao, Qian and Fu, Fangcheng and Chen, Kai},
  journal={arXiv preprint arXiv:2508.15601},
  year={2025}
}

License

This project is released under the Apache 2.0 license.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

lmdeploy-0.16.0-cp314-cp314-win_amd64.whl (60.1 MB view details)

Uploaded CPython 3.14Windows x86-64

lmdeploy-0.16.0-cp314-cp314-manylinux_2_28_x86_64.whl (106.0 MB view details)

Uploaded CPython 3.14manylinux: glibc 2.28+ x86-64

lmdeploy-0.16.0-cp313-cp313-win_amd64.whl (59.8 MB view details)

Uploaded CPython 3.13Windows x86-64

lmdeploy-0.16.0-cp313-cp313-manylinux_2_28_x86_64.whl (106.0 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.28+ x86-64

lmdeploy-0.16.0-cp312-cp312-win_amd64.whl (59.8 MB view details)

Uploaded CPython 3.12Windows x86-64

lmdeploy-0.16.0-cp312-cp312-manylinux_2_28_x86_64.whl (106.0 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.28+ x86-64

lmdeploy-0.16.0-cp311-cp311-win_amd64.whl (59.8 MB view details)

Uploaded CPython 3.11Windows x86-64

lmdeploy-0.16.0-cp311-cp311-manylinux_2_28_x86_64.whl (106.0 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.28+ x86-64

lmdeploy-0.16.0-cp310-cp310-win_amd64.whl (59.8 MB view details)

Uploaded CPython 3.10Windows x86-64

lmdeploy-0.16.0-cp310-cp310-manylinux_2_28_x86_64.whl (106.0 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.28+ x86-64

File details

Details for the file lmdeploy-0.16.0-cp314-cp314-win_amd64.whl.

File metadata

  • Download URL: lmdeploy-0.16.0-cp314-cp314-win_amd64.whl
  • Upload date:
  • Size: 60.1 MB
  • Tags: CPython 3.14, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.20

File hashes

Hashes for lmdeploy-0.16.0-cp314-cp314-win_amd64.whl
Algorithm Hash digest
SHA256 4a2ee31623862eae7a31349cf91b14da577a338a4979c12dd9fa7602117375c3
MD5 12033ae652483f0aebad31ff43f531cb
BLAKE2b-256 aa9318385f763ed6b69bf499fec69fae224ccbb91095ba94e38ed6fdf1f4868b

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp314-cp314-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for lmdeploy-0.16.0-cp314-cp314-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 d8384d13d9cc832e4d21f0cf3d1922362b7f5efab7fe32d155bc77f9170249a4
MD5 062c11a7596c701d58361cb7a61f67c4
BLAKE2b-256 d73e7622eb8ff4c472dbb788e166cf579f1dacff5e61db23b72f4f4d9584fce1

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp313-cp313-win_amd64.whl.

File metadata

  • Download URL: lmdeploy-0.16.0-cp313-cp313-win_amd64.whl
  • Upload date:
  • Size: 59.8 MB
  • Tags: CPython 3.13, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.20

File hashes

Hashes for lmdeploy-0.16.0-cp313-cp313-win_amd64.whl
Algorithm Hash digest
SHA256 ad6e3705cf05c673b32976ca0aca9c00b27f90b61f2c57200dcc469fd851713d
MD5 243d78dd7150965755d7d9c725255a07
BLAKE2b-256 5f4462db137e6e1b049e10b734d8915679f2c0926262ca6aaad60c643423eb13

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp313-cp313-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for lmdeploy-0.16.0-cp313-cp313-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 e9016976d64ba679462167f5ab7df11475f764293daf2c7e1500bcbaaeae217a
MD5 b5aa3d567cb7f0033c698a3bbe522b2a
BLAKE2b-256 fb1759c1bae7789f8b2bb9e7b81b4a1027cafda1dc80a4bb9322d2a71d51c278

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp312-cp312-win_amd64.whl.

File metadata

  • Download URL: lmdeploy-0.16.0-cp312-cp312-win_amd64.whl
  • Upload date:
  • Size: 59.8 MB
  • Tags: CPython 3.12, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.20

File hashes

Hashes for lmdeploy-0.16.0-cp312-cp312-win_amd64.whl
Algorithm Hash digest
SHA256 df890d755e9afbfdd2f6e384b0033278110a6c3c80b67fd4a1a35530d631b539
MD5 9cb602cffb22bd770cb40314cecdce4b
BLAKE2b-256 471a77010cb436df5a9f56838ea5c320b8419b79206f10f7178ca06ceef53924

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp312-cp312-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for lmdeploy-0.16.0-cp312-cp312-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 3bdee3d6696fc28f8fb0dd7f638e429ef95255ba47e7ee33685a3b03105ea7e8
MD5 d6679ff48484ee970d8cca94ee248e79
BLAKE2b-256 9627c391dbf1cf9f043515fdc3c88a00f0ef21479c20caafff0627fd21481eb6

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp311-cp311-win_amd64.whl.

File metadata

  • Download URL: lmdeploy-0.16.0-cp311-cp311-win_amd64.whl
  • Upload date:
  • Size: 59.8 MB
  • Tags: CPython 3.11, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.20

File hashes

Hashes for lmdeploy-0.16.0-cp311-cp311-win_amd64.whl
Algorithm Hash digest
SHA256 1f4b3bd288ba920b976a0cc81124413e4da21c9547b95405ff4a0afe473d3db3
MD5 a5e008dd60c1c11fd54821ea626354cc
BLAKE2b-256 d044f3aab08f31657b8eef8537554f037902c23cf906cce93567f57e50946657

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp311-cp311-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for lmdeploy-0.16.0-cp311-cp311-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 0b5b1c7994a2bb77d0817376820f815d67433103510d87d09937d3be152d4c89
MD5 9846adcf35131973cd594356e9ef71db
BLAKE2b-256 97ffc0c1a8de366b6891cfdc5f12451fa59026f1dd7bba7510d3d1ea3b064ec6

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp310-cp310-win_amd64.whl.

File metadata

  • Download URL: lmdeploy-0.16.0-cp310-cp310-win_amd64.whl
  • Upload date:
  • Size: 59.8 MB
  • Tags: CPython 3.10, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.20

File hashes

Hashes for lmdeploy-0.16.0-cp310-cp310-win_amd64.whl
Algorithm Hash digest
SHA256 06336a61b05f5b06bc3fb2018e25d3614323005d4c6e6797c89888988865e6f7
MD5 f9b687ab4dd167bc23d3c3940984c619
BLAKE2b-256 360426185de2c9dde65c8e6dad12e8277bb67408b9f220f617452f06b447f485

See more details on using hashes here.

File details

Details for the file lmdeploy-0.16.0-cp310-cp310-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for lmdeploy-0.16.0-cp310-cp310-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 5ee7a42a7e62f74e189f22d503a3bf42d616bf8382cf49582576e70aa75d0551
MD5 288674b98cfff9060312837556a8f0ec
BLAKE2b-256 0630bc3de8f8dbe1b656073653419484a8e091efb858c8ad000df009e17d9149

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page