zh-nj/lmdeploy-v100

This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 2026. It does not account for compatibility with other architectures and has only been tested on an 8-card V100 32GB setup.

21

stars

6

commits

Python

primary language

Mar 18, 2026

updated

README

Since 2025Q4, running new open-source models on V100 for inference has become largely impractical — especially prefill performance is painful to watch. As an owner of 8× V100 GPUs, I finally had enough and spent the Spring Festival holiday and spare time afterwards developing this project to solve the pain points that had been bothering me. What a relief!

This project is developed specifically for V100, based on lmdeploy 0.12.1. It adds support for mainstream open-source models from 2025Q4 to 2026Q1 that I personally wanted to use. No effort has been made to ensure compatibility with other architectures — only verified on 8× V100 32GB. Due to the massive amount of work involved, merging back into the upstream project would be very difficult, so it is released independently first, for fellow V100 owners to enjoy. PRs are welcome.
Supported models:
Qwen3.5-122B-A10B-AWQ-4bit, Qwen3.5-35B-A3B-AWQ, Qwen3.5-27B-AWQ-4bit, Qwen3.5-9B-AWQ-4bit, Qwen3.5-4B-AWQ-4bit, Qwen3.5-2B, Qwen3.5-0.8B (multimodal with video support, MTP)
MiniMax-M2.5-AWQ, MiniMax-M2.1-AWQ
Step-3.5-Flash-int4-mixed-AutoRound (MTP supported), Step-3.5-Flash-AWQ-4bit (this model does not include MTP layers)
Qwen3-Coder-Next-AWQ-4bit, Qwen3-Next-80B-A3B-Thinking-AWQ-4bit
New features:
Anthropic API
AutoRound & compressed-tensors quantization support
GGUF support (in progress)
V100 low-level MMA_884 kernel refactoring (in progress)


Install from Source

# 1. Clone the repository
git clone https://github.com/zh-nj/lmdeploy-v100.git
cd lmdeploy-v100

# 2. Create a conda environment (Python 3.10 recommended)
conda create -n lmdeploy python=3.10 -y
conda activate lmdeploy

# 3. Install PyTorch (CUDA 11.8 example, adjust for your CUDA version)
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118

# 4. Install in development mode (automatically compiles C++/CUDA)
pip install -e .

# For all optional dependencies:
# pip install -e ".[all]"

Build requires CMake 3.11+, Ninja, and GCC with C++17 support. V100 users should ensure CUDA_DEVICE_ORDER=PCI_BUS_ID is set.


Model Inference Launch Examples

All examples use the TurboMind engine (lmdeploy serve api_server). Replace model paths with your actual paths.

Qwen3.5 Series

# Qwen3.5-122B-A10B-AWQ (MoE, TP=8)
lmdeploy serve api_server /path/to/Qwen3.5-122B-A10B-AWQ \
    --tp 8 --dtype float16 --session-len 8192 \
    --cache-max-entry-count 0.7 \
    --tool-call-parser qwen3_5 --reasoning-parser deepseek-r1

# Qwen3.5-35B-A3B-AWQ (MoE, TP=4)
lmdeploy serve api_server /path/to/Qwen3.5-35B-A3B-AWQ \
    --tp 4 --dtype float16 --cache-max-entry-count 0.7 \
    --tool-call-parser qwen3_5 --reasoning-parser deepseek-r1

# Qwen3.5-27B-AWQ (Dense, TP=4)
lmdeploy serve api_server /path/to/Qwen3.5-27B-AWQ \
    --tp 4 --dtype float16 --cache-max-entry-count 0.7 \
    --tool-call-parser qwen3_5 --reasoning-parser deepseek-r1

# Qwen3.5-9B-AWQ / 4B-AWQ (TP=1 or TP=2)
lmdeploy serve api_server /path/to/Qwen3.5-9B-AWQ \
    --tp 1 --dtype float16 --cache-max-entry-count 0.8

# Qwen3.5-2B / 0.8B (small models, TP=1)
lmdeploy serve api_server /path/to/Qwen3.5-2B \
    --tp 1 --dtype float16

Qwen3.5 MTP Speculative Decoding (faster decode)

# Add --num-draft-tokens 1 to any command above
lmdeploy serve api_server /path/to/Qwen3.5-27B-AWQ \
    --tp 4 --dtype float16 --num-draft-tokens 1 \
    --cache-max-entry-count 0.4

Qwen3.5 Multimodal (image + video)

Multimodal inputs are passed via the OpenAI-compatible API using image_url / video_url fields. No extra launch parameters needed.

MiniMax Series

# MiniMax-M2.5-AWQ (TP=8)
lmdeploy serve api_server /path/to/MiniMax-M2.5-AWQ \
    --tp 8 --dtype float16 --cache-max-entry-count 0.7 \
    --tool-call-parser minimax-m2 --reasoning-parser minimax-m2

# MiniMax-M2.1-AWQ (TP=8)
lmdeploy serve api_server /path/to/MiniMax-M2.1-AWQ \
    --tp 8 --dtype float16 --cache-max-entry-count 0.7

Step-3.5-Flash

# Step-3.5-Flash AutoRound int4 mixed quantization (TP=8, MTP supported)
lmdeploy serve api_server /path/to/Step-3.5-Flash-int4-mixed-AutoRound \
    --tp 8 --dtype float16 --session-len 8192 \
    --num-draft-tokens 1 \
    --tool-call-parser step3p5 --reasoning-parser step3p5

# Step-3.5-Flash AWQ (TP=8, no MTP layers)
lmdeploy serve api_server /path/to/Step-3.5-Flash-AWQ \
    --tp 8 --dtype float16 --session-len 8192 \
    --tool-call-parser step3p5 --reasoning-parser step3p5

Qwen3-Coder-Next / Qwen3-Next

# Qwen3-Coder-Next-AWQ (MoE, TP=4)
lmdeploy serve api_server /path/to/Qwen3-Coder-Next-AWQ-4bit \
    --tp 4 --dtype float16 --cache-max-entry-count 0.4 \
    --num-draft-tokens 1

# Qwen3-Next-80B-A3B-Thinking-AWQ (MoE, TP=4)
lmdeploy serve api_server /path/to/Qwen3-Next-80B-A3B-Thinking-AWQ-4bit \
    --tp 4 --dtype float16 --cache-max-entry-count 0.7

Latest News 🎉

2026
2025
  • [2025/09] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the performmance of vLLM on H800 for openai gpt-oss models!
  • [2025/06] Comprehensive inference optimization for FP8 MoE Models
  • [2025/06] DeepSeek PD Disaggregation deployment is now supported through integration with DLSlime and Mooncake. Huge thanks to both teams!
  • [2025/04] Enhance DeepSeek inference performance by integration deepseek-ai techniques: FlashMLA, DeepGemm, DeepEP, MicroBatch and eplb
  • [2025/01] Support DeepSeek V3 and R1
2024
  • [2024/11] Support Mono-InternVL with PyTorch engine
  • [2024/10] PyTorchEngine supports graph mode on ascend platform, doubling the inference speed
  • [2024/09] LMDeploy PyTorchEngine adds support for Huawei Ascend. See supported models here
  • [2024/09] LMDeploy PyTorchEngine achieves 1.3x faster on Llama3-8B inference by introducing CUDA graph
  • [2024/08] LMDeploy is integrated into modelscope/swift as the default accelerator for VLMs inference
  • [2024/07] Support Llama3.1 8B, 70B and its TOOLS CALLING
  • [2024/07] Support InternVL2 full-series models, InternLM-XComposer2.5 and function call of InternLM2.5
  • [2024/06] PyTorch engine support DeepSeek-V2 and several VLMs, such as CogVLM2, Mini-InternVL, LlaVA-Next
  • [2024/05] Balance vision model when deploying VLMs with multiple GPUs
  • [2024/05] Support 4-bits weight-only quantization and inference on VLMs, such as InternVL v1.5, LLaVa, InternLMXComposer2
  • [2024/04] Support Llama3 and more VLMs, such as InternVL v1.1, v1.2, MiniGemini, InternLMXComposer2.
  • [2024/04] TurboMind adds online int8/int4 KV cache quantization and inference for all supported devices. Refer here for detailed guide
  • [2024/04] TurboMind latest upgrade boosts GQA, rocketing the internlm2-20b model inference to 16+ RPS, about 1.8x faster than vLLM.
  • [2024/04] Support Qwen1.5-MOE and dbrx.
  • [2024/03] Support DeepSeek-VL offline inference pipeline and serving.
  • [2024/03] Support VLM offline inference pipeline and serving.
  • [2024/02] Support Qwen 1.5, Gemma, Mistral, Mixtral, Deepseek-MOE and so on.
  • [2024/01] OpenAOE seamless integration with LMDeploy Serving Service.
  • [2024/01] Support for multi-model, multi-machine, multi-card inference services. For usage instructions, please refer to here
  • [2024/01] Support PyTorch inference engine, developed entirely in Python, helping to lower the barriers for developers and enable rapid experimentation with new features and technologies.
2023
  • [2023/12] Turbomind supports multimodal input.
  • [2023/11] Turbomind supports loading hf model directly. Click here for details.
  • [2023/11] TurboMind major upgrades, including: Paged Attention, faster attention kernels without sequence length limitation, 2x faster KV8 kernels, Split-K decoding (Flash Decoding), and W4A16 inference for sm_75
  • [2023/09] TurboMind supports Qwen-14B
  • [2023/09] TurboMind supports InternLM-20B
  • [2023/09] TurboMind supports all features of Code Llama: code completion, infilling, chat / instruct, and python specialist. Click here for deployment guide
  • [2023/09] TurboMind supports Baichuan2-7B
  • [2023/08] TurboMind supports flash-attention2.
  • [2023/08] TurboMind supports Qwen-7B, dynamic NTK-RoPE scaling and dynamic logN scaling
  • [2023/08] TurboMind supports Windows (tp=1)
  • [2023/08] TurboMind supports 4-bit inference, 2.4x faster than FP16, the fastest open-source implementation. Check this guide for detailed info
  • [2023/08] LMDeploy has launched on the HuggingFace Hub, providing ready-to-use 4-bit models.
  • [2023/08] LMDeploy supports 4-bit quantization using the AWQ algorithm.
  • [2023/07] TurboMind supports Llama-2 70B with GQA.
  • [2023/07] TurboMind supports Llama-2 7B/13B.
  • [2023/07] TurboMind supports tensor-parallel inference of InternLM.

Introduction

LMDeploy is a toolkit for compressing, deploying, and serving LLM, developed by the MMRazor and MMDeploy teams. It has the following core features:

  • Efficient Inference: LMDeploy delivers up to 1.8x higher request throughput than vLLM, by introducing key features like persistent batch(a.k.a. continuous batching), blocked KV cache, dynamic split&fuse, tensor parallelism, high-performance CUDA kernels and so on.

  • Effective Quantization: LMDeploy supports weight-only and k/v quantization, and the 4-bit inference performance is 2.4x higher than FP16. The quantization quality has been confirmed via OpenCompass evaluation.

  • Effortless Distribution Server: Leveraging the request distribution service, LMDeploy facilitates an easy and efficient deployment of multi-model services across multiple machines and cards.

  • Excellent Compatibility: LMDeploy supports KV Cache Quant, AWQ and Automatic Prefix Caching to be used simultaneously.

Performance

v0 1 0-benchmark

Supported Models

LLMs VLMs
  • Llama (7B - 65B)
  • Llama2 (7B - 70B)
  • Llama3 (8B, 70B)
  • Llama3.1 (8B, 70B)
  • Llama3.2 (1B, 3B)
  • InternLM (7B - 20B)
  • InternLM2 (7B - 20B)
  • InternLM3 (8B)
  • InternLM2.5 (7B)
  • Qwen (1.8B - 72B)
  • Qwen1.5 (0.5B - 110B)
  • Qwen1.5 - MoE (0.5B - 72B)
  • Qwen2 (0.5B - 72B)
  • Qwen2-MoE (57BA14B)
  • Qwen2.5 (0.5B - 32B)
  • Qwen3, Qwen3-MoE
  • Qwen3-Next(80B)
  • Baichuan (7B)
  • Baichuan2 (7B-13B)
  • Code Llama (7B - 34B)
  • ChatGLM2 (6B)
  • GLM-4 (9B)
  • GLM-4-0414 (9B, 32B)
  • CodeGeeX4 (9B)
  • YI (6B-34B)
  • Mistral (7B)
  • DeepSeek-MoE (16B)
  • DeepSeek-V2 (16B, 236B)
  • DeepSeek-V2.5 (236B)
  • DeepSeek-V3 (685B)
  • DeepSeek-V3.2 (685B)
  • Mixtral (8x7B, 8x22B)
  • Gemma (2B - 7B)
  • StarCoder2 (3B - 15B)
  • Phi-3-mini (3.8B)
  • Phi-3.5-mini (3.8B)
  • Phi-3.5-MoE (16x3.8B)
  • Phi-4-mini (3.8B)
  • MiniCPM3 (4B)
  • SDAR (1.7B-30B)
  • gpt-oss (20B, 120B)
  • GLM-4.7-Flash (30B)
  • LLaVA(1.5,1.6) (7B-34B)
  • InternLM-XComposer2 (7B, 4khd-7B)
  • InternLM-XComposer2.5 (7B)
  • Qwen-VL (7B)
  • Qwen2-VL (2B, 7B, 72B)
  • Qwen2.5-VL (3B, 7B, 72B)
  • Qwen3-VL (2B - 235B)
  • DeepSeek-VL (7B)
  • DeepSeek-VL2 (3B, 16B, 27B)
  • InternVL-Chat (v1.1-v1.5)
  • InternVL2 (1B-76B)
  • InternVL2.5(MPO) (1B-78B)
  • InternVL3 (1B-78B)
  • InternVL3.5 (1B-241BA28B)
  • Intern-S1 (241B)
  • Intern-S1-mini (8.3B)
  • Intern-S1-Pro (1TB)
  • Mono-InternVL (2B)
  • ChemVLM (8B-26B)
  • CogVLM-Chat (17B)
  • CogVLM2-Chat (19B)
  • MiniCPM-Llama3-V-2_5
  • MiniCPM-V-2_6
  • Phi-3-vision (4.2B)
  • Phi-3.5-vision (4.2B)
  • GLM-4V (9B)
  • GLM-4.1V-Thinking (9B)
  • Llama3.2-vision (11B, 90B)
  • Molmo (7B-D,72B)
  • Gemma3 (1B - 27B)
  • Llama4 (Scout, Maverick)

LMDeploy has developed two inference engines - TurboMind and PyTorch, each with a different focus. The former strives for ultimate optimization of inference performance, while the latter, developed purely in Python, aims to decrease the barriers for developers.

They differ in the types of supported models and the inference data type. Please refer to this table for each engine's capability and choose the proper one that best fits your actual needs.

Quick Start Open In Colab

Installation

It is recommended installing lmdeploy using pip in a conda environment (python 3.10 - 3.13):

conda create -n lmdeploy python=3.10 -y
conda activate lmdeploy
pip install lmdeploy

The default prebuilt package is compiled on CUDA 12 since v0.3.0.

For the GeForce RTX 50 series, please install the LMDeploy prebuilt package complied with CUDA 12.8

export LMDEPLOY_VERSION=0.12.0
export PYTHON_VERSION=310
pip install https://github.com/InternLM/lmdeploy/releases/download/v${LMDEPLOY_VERSION}/lmdeploy-${LMDEPLOY_VERSION}+cu128-cp${PYTHON_VERSION}-cp${PYTHON_VERSION}-manylinux2014_x86_64.whl --extra-index-url https://download.pytorch.org/whl/cu128

For more information on installing on CUDA 11+ platform, or for instructions on building from source, please refer to the installation guide.

Offline Batch Inference

import lmdeploy
with lmdeploy.pipeline("internlm/internlm3-8b-instruct") as pipe:
    response = pipe(["Hi, pls intro yourself", "Shanghai is"])
    print(response)

[!NOTE] By default, LMDeploy downloads model from HuggingFace. If you would like to use models from ModelScope, please install ModelScope by pip install modelscope and set the environment variable:

export LMDEPLOY_USE_MODELSCOPE=True

If you would like to use models from openMind Hub, please install openMind Hub by pip install openmind_hub and set the environment variable:

export LMDEPLOY_USE_OPENMIND_HUB=True

For more information about inference pipeline, please refer to here.

Tutorials

Please review getting_started section for the basic usage of LMDeploy.

For detailed user guides and advanced guides, please refer to our tutorials:

Third-party projects

  • Deploying LLMs offline on the NVIDIA Jetson platform by LMDeploy: LMDeploy-Jetson

  • Example project for deploying LLMs using LMDeploy and BentoML: BentoLMDeploy

Contributing

We appreciate all contributions to LMDeploy. Please refer to CONTRIBUTING.md for the contributing guideline.

Acknowledgement

Citation

@misc{2023lmdeploy,
    title={LMDeploy: A Toolkit for Compressing, Deploying, and Serving LLM},
    author={LMDeploy Contributors},
    howpublished = {\url{https://github.com/InternLM/lmdeploy}},
    year={2023}
}
@article{zhang2025efficient,
  title={Efficient Mixed-Precision Large Language Model Inference with TurboMind},
  author={Zhang, Li and Jiang, Youhe and He, Guoliang and Chen, Xin and Lv, Han and Yao, Qian and Fu, Fangcheng and Chen, Kai},
  journal={arXiv preprint arXiv:2508.15601},
  year={2025}
}

License

This project is released under the Apache 2.0 license.

Contributors

zh-nj

6 commits

zh-nj/lmdeploy-v100

This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 2026. It does not account for compatibility with other architectures and has only been tested on an 8-card V100 32GB setup.

21

stars

6

commits

Python

primary language

Mar 18, 2026

updated

README

Since 2025Q4, running new open-source models on V100 for inference has become largely impractical — especially prefill performance is painful to watch. As an owner of 8× V100 GPUs, I finally had enough and spent the Spring Festival holiday and spare time afterwards developing this project to solve the pain points that had been bothering me. What a relief!

This project is developed specifically for V100, based on lmdeploy 0.12.1. It adds support for mainstream open-source models from 2025Q4 to 2026Q1 that I personally wanted to use. No effort has been made to ensure compatibility with other architectures — only verified on 8× V100 32GB. Due to the massive amount of work involved, merging back into the upstream project would be very difficult, so it is released independently first, for fellow V100 owners to enjoy. PRs are welcome.
Supported models:
Qwen3.5-122B-A10B-AWQ-4bit, Qwen3.5-35B-A3B-AWQ, Qwen3.5-27B-AWQ-4bit, Qwen3.5-9B-AWQ-4bit, Qwen3.5-4B-AWQ-4bit, Qwen3.5-2B, Qwen3.5-0.8B (multimodal with video support, MTP)
MiniMax-M2.5-AWQ, MiniMax-M2.1-AWQ
Step-3.5-Flash-int4-mixed-AutoRound (MTP supported), Step-3.5-Flash-AWQ-4bit (this model does not include MTP layers)
Qwen3-Coder-Next-AWQ-4bit, Qwen3-Next-80B-A3B-Thinking-AWQ-4bit
New features:
Anthropic API
AutoRound & compressed-tensors quantization support
GGUF support (in progress)
V100 low-level MMA_884 kernel refactoring (in progress)


Install from Source

# 1. Clone the repository
git clone https://github.com/zh-nj/lmdeploy-v100.git
cd lmdeploy-v100

# 2. Create a conda environment (Python 3.10 recommended)
conda create -n lmdeploy python=3.10 -y
conda activate lmdeploy

# 3. Install PyTorch (CUDA 11.8 example, adjust for your CUDA version)
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118

# 4. Install in development mode (automatically compiles C++/CUDA)
pip install -e .

# For all optional dependencies:
# pip install -e ".[all]"

Build requires CMake 3.11+, Ninja, and GCC with C++17 support. V100 users should ensure CUDA_DEVICE_ORDER=PCI_BUS_ID is set.


Model Inference Launch Examples

All examples use the TurboMind engine (lmdeploy serve api_server). Replace model paths with your actual paths.

Qwen3.5 Series

# Qwen3.5-122B-A10B-AWQ (MoE, TP=8)
lmdeploy serve api_server /path/to/Qwen3.5-122B-A10B-AWQ \
    --tp 8 --dtype float16 --session-len 8192 \
    --cache-max-entry-count 0.7 \
    --tool-call-parser qwen3_5 --reasoning-parser deepseek-r1

# Qwen3.5-35B-A3B-AWQ (MoE, TP=4)
lmdeploy serve api_server /path/to/Qwen3.5-35B-A3B-AWQ \
    --tp 4 --dtype float16 --cache-max-entry-count 0.7 \
    --tool-call-parser qwen3_5 --reasoning-parser deepseek-r1

# Qwen3.5-27B-AWQ (Dense, TP=4)
lmdeploy serve api_server /path/to/Qwen3.5-27B-AWQ \
    --tp 4 --dtype float16 --cache-max-entry-count 0.7 \
    --tool-call-parser qwen3_5 --reasoning-parser deepseek-r1

# Qwen3.5-9B-AWQ / 4B-AWQ (TP=1 or TP=2)
lmdeploy serve api_server /path/to/Qwen3.5-9B-AWQ \
    --tp 1 --dtype float16 --cache-max-entry-count 0.8

# Qwen3.5-2B / 0.8B (small models, TP=1)
lmdeploy serve api_server /path/to/Qwen3.5-2B \
    --tp 1 --dtype float16

Qwen3.5 MTP Speculative Decoding (faster decode)

# Add --num-draft-tokens 1 to any command above
lmdeploy serve api_server /path/to/Qwen3.5-27B-AWQ \
    --tp 4 --dtype float16 --num-draft-tokens 1 \
    --cache-max-entry-count 0.4

Qwen3.5 Multimodal (image + video)

Multimodal inputs are passed via the OpenAI-compatible API using image_url / video_url fields. No extra launch parameters needed.

MiniMax Series

# MiniMax-M2.5-AWQ (TP=8)
lmdeploy serve api_server /path/to/MiniMax-M2.5-AWQ \
    --tp 8 --dtype float16 --cache-max-entry-count 0.7 \
    --tool-call-parser minimax-m2 --reasoning-parser minimax-m2

# MiniMax-M2.1-AWQ (TP=8)
lmdeploy serve api_server /path/to/MiniMax-M2.1-AWQ \
    --tp 8 --dtype float16 --cache-max-entry-count 0.7

Step-3.5-Flash

# Step-3.5-Flash AutoRound int4 mixed quantization (TP=8, MTP supported)
lmdeploy serve api_server /path/to/Step-3.5-Flash-int4-mixed-AutoRound \
    --tp 8 --dtype float16 --session-len 8192 \
    --num-draft-tokens 1 \
    --tool-call-parser step3p5 --reasoning-parser step3p5

# Step-3.5-Flash AWQ (TP=8, no MTP layers)
lmdeploy serve api_server /path/to/Step-3.5-Flash-AWQ \
    --tp 8 --dtype float16 --session-len 8192 \
    --tool-call-parser step3p5 --reasoning-parser step3p5

Qwen3-Coder-Next / Qwen3-Next

# Qwen3-Coder-Next-AWQ (MoE, TP=4)
lmdeploy serve api_server /path/to/Qwen3-Coder-Next-AWQ-4bit \
    --tp 4 --dtype float16 --cache-max-entry-count 0.4 \
    --num-draft-tokens 1

# Qwen3-Next-80B-A3B-Thinking-AWQ (MoE, TP=4)
lmdeploy serve api_server /path/to/Qwen3-Next-80B-A3B-Thinking-AWQ-4bit \
    --tp 4 --dtype float16 --cache-max-entry-count 0.7

Latest News 🎉

2026
2025
  • [2025/09] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the performmance of vLLM on H800 for openai gpt-oss models!
  • [2025/06] Comprehensive inference optimization for FP8 MoE Models
  • [2025/06] DeepSeek PD Disaggregation deployment is now supported through integration with DLSlime and Mooncake. Huge thanks to both teams!
  • [2025/04] Enhance DeepSeek inference performance by integration deepseek-ai techniques: FlashMLA, DeepGemm, DeepEP, MicroBatch and eplb
  • [2025/01] Support DeepSeek V3 and R1
2024
  • [2024/11] Support Mono-InternVL with PyTorch engine
  • [2024/10] PyTorchEngine supports graph mode on ascend platform, doubling the inference speed
  • [2024/09] LMDeploy PyTorchEngine adds support for Huawei Ascend. See supported models here
  • [2024/09] LMDeploy PyTorchEngine achieves 1.3x faster on Llama3-8B inference by introducing CUDA graph
  • [2024/08] LMDeploy is integrated into modelscope/swift as the default accelerator for VLMs inference
  • [2024/07] Support Llama3.1 8B, 70B and its TOOLS CALLING
  • [2024/07] Support InternVL2 full-series models, InternLM-XComposer2.5 and function call of InternLM2.5
  • [2024/06] PyTorch engine support DeepSeek-V2 and several VLMs, such as CogVLM2, Mini-InternVL, LlaVA-Next
  • [2024/05] Balance vision model when deploying VLMs with multiple GPUs
  • [2024/05] Support 4-bits weight-only quantization and inference on VLMs, such as InternVL v1.5, LLaVa, InternLMXComposer2
  • [2024/04] Support Llama3 and more VLMs, such as InternVL v1.1, v1.2, MiniGemini, InternLMXComposer2.
  • [2024/04] TurboMind adds online int8/int4 KV cache quantization and inference for all supported devices. Refer here for detailed guide
  • [2024/04] TurboMind latest upgrade boosts GQA, rocketing the internlm2-20b model inference to 16+ RPS, about 1.8x faster than vLLM.
  • [2024/04] Support Qwen1.5-MOE and dbrx.
  • [2024/03] Support DeepSeek-VL offline inference pipeline and serving.
  • [2024/03] Support VLM offline inference pipeline and serving.
  • [2024/02] Support Qwen 1.5, Gemma, Mistral, Mixtral, Deepseek-MOE and so on.
  • [2024/01] OpenAOE seamless integration with LMDeploy Serving Service.
  • [2024/01] Support for multi-model, multi-machine, multi-card inference services. For usage instructions, please refer to here
  • [2024/01] Support PyTorch inference engine, developed entirely in Python, helping to lower the barriers for developers and enable rapid experimentation with new features and technologies.
2023
  • [2023/12] Turbomind supports multimodal input.
  • [2023/11] Turbomind supports loading hf model directly. Click here for details.
  • [2023/11] TurboMind major upgrades, including: Paged Attention, faster attention kernels without sequence length limitation, 2x faster KV8 kernels, Split-K decoding (Flash Decoding), and W4A16 inference for sm_75
  • [2023/09] TurboMind supports Qwen-14B
  • [2023/09] TurboMind supports InternLM-20B
  • [2023/09] TurboMind supports all features of Code Llama: code completion, infilling, chat / instruct, and python specialist. Click here for deployment guide
  • [2023/09] TurboMind supports Baichuan2-7B
  • [2023/08] TurboMind supports flash-attention2.
  • [2023/08] TurboMind supports Qwen-7B, dynamic NTK-RoPE scaling and dynamic logN scaling
  • [2023/08] TurboMind supports Windows (tp=1)
  • [2023/08] TurboMind supports 4-bit inference, 2.4x faster than FP16, the fastest open-source implementation. Check this guide for detailed info
  • [2023/08] LMDeploy has launched on the HuggingFace Hub, providing ready-to-use 4-bit models.
  • [2023/08] LMDeploy supports 4-bit quantization using the AWQ algorithm.
  • [2023/07] TurboMind supports Llama-2 70B with GQA.
  • [2023/07] TurboMind supports Llama-2 7B/13B.
  • [2023/07] TurboMind supports tensor-parallel inference of InternLM.

Introduction

LMDeploy is a toolkit for compressing, deploying, and serving LLM, developed by the MMRazor and MMDeploy teams. It has the following core features:

  • Efficient Inference: LMDeploy delivers up to 1.8x higher request throughput than vLLM, by introducing key features like persistent batch(a.k.a. continuous batching), blocked KV cache, dynamic split&fuse, tensor parallelism, high-performance CUDA kernels and so on.

  • Effective Quantization: LMDeploy supports weight-only and k/v quantization, and the 4-bit inference performance is 2.4x higher than FP16. The quantization quality has been confirmed via OpenCompass evaluation.

  • Effortless Distribution Server: Leveraging the request distribution service, LMDeploy facilitates an easy and efficient deployment of multi-model services across multiple machines and cards.

  • Excellent Compatibility: LMDeploy supports KV Cache Quant, AWQ and Automatic Prefix Caching to be used simultaneously.

Performance

v0 1 0-benchmark

Supported Models

LLMs VLMs
  • Llama (7B - 65B)
  • Llama2 (7B - 70B)
  • Llama3 (8B, 70B)
  • Llama3.1 (8B, 70B)
  • Llama3.2 (1B, 3B)
  • InternLM (7B - 20B)
  • InternLM2 (7B - 20B)
  • InternLM3 (8B)
  • InternLM2.5 (7B)
  • Qwen (1.8B - 72B)
  • Qwen1.5 (0.5B - 110B)
  • Qwen1.5 - MoE (0.5B - 72B)
  • Qwen2 (0.5B - 72B)
  • Qwen2-MoE (57BA14B)
  • Qwen2.5 (0.5B - 32B)
  • Qwen3, Qwen3-MoE
  • Qwen3-Next(80B)
  • Baichuan (7B)
  • Baichuan2 (7B-13B)
  • Code Llama (7B - 34B)
  • ChatGLM2 (6B)
  • GLM-4 (9B)
  • GLM-4-0414 (9B, 32B)
  • CodeGeeX4 (9B)
  • YI (6B-34B)
  • Mistral (7B)
  • DeepSeek-MoE (16B)
  • DeepSeek-V2 (16B, 236B)
  • DeepSeek-V2.5 (236B)
  • DeepSeek-V3 (685B)
  • DeepSeek-V3.2 (685B)
  • Mixtral (8x7B, 8x22B)
  • Gemma (2B - 7B)
  • StarCoder2 (3B - 15B)
  • Phi-3-mini (3.8B)
  • Phi-3.5-mini (3.8B)
  • Phi-3.5-MoE (16x3.8B)
  • Phi-4-mini (3.8B)
  • MiniCPM3 (4B)
  • SDAR (1.7B-30B)
  • gpt-oss (20B, 120B)
  • GLM-4.7-Flash (30B)
  • LLaVA(1.5,1.6) (7B-34B)
  • InternLM-XComposer2 (7B, 4khd-7B)
  • InternLM-XComposer2.5 (7B)
  • Qwen-VL (7B)
  • Qwen2-VL (2B, 7B, 72B)
  • Qwen2.5-VL (3B, 7B, 72B)
  • Qwen3-VL (2B - 235B)
  • DeepSeek-VL (7B)
  • DeepSeek-VL2 (3B, 16B, 27B)
  • InternVL-Chat (v1.1-v1.5)
  • InternVL2 (1B-76B)
  • InternVL2.5(MPO) (1B-78B)
  • InternVL3 (1B-78B)
  • InternVL3.5 (1B-241BA28B)
  • Intern-S1 (241B)
  • Intern-S1-mini (8.3B)
  • Intern-S1-Pro (1TB)
  • Mono-InternVL (2B)
  • ChemVLM (8B-26B)
  • CogVLM-Chat (17B)
  • CogVLM2-Chat (19B)
  • MiniCPM-Llama3-V-2_5
  • MiniCPM-V-2_6
  • Phi-3-vision (4.2B)
  • Phi-3.5-vision (4.2B)
  • GLM-4V (9B)
  • GLM-4.1V-Thinking (9B)
  • Llama3.2-vision (11B, 90B)
  • Molmo (7B-D,72B)
  • Gemma3 (1B - 27B)
  • Llama4 (Scout, Maverick)

LMDeploy has developed two inference engines - TurboMind and PyTorch, each with a different focus. The former strives for ultimate optimization of inference performance, while the latter, developed purely in Python, aims to decrease the barriers for developers.

They differ in the types of supported models and the inference data type. Please refer to this table for each engine's capability and choose the proper one that best fits your actual needs.

Quick Start Open In Colab

Installation

It is recommended installing lmdeploy using pip in a conda environment (python 3.10 - 3.13):

conda create -n lmdeploy python=3.10 -y
conda activate lmdeploy
pip install lmdeploy

The default prebuilt package is compiled on CUDA 12 since v0.3.0.

For the GeForce RTX 50 series, please install the LMDeploy prebuilt package complied with CUDA 12.8

export LMDEPLOY_VERSION=0.12.0
export PYTHON_VERSION=310
pip install https://github.com/InternLM/lmdeploy/releases/download/v${LMDEPLOY_VERSION}/lmdeploy-${LMDEPLOY_VERSION}+cu128-cp${PYTHON_VERSION}-cp${PYTHON_VERSION}-manylinux2014_x86_64.whl --extra-index-url https://download.pytorch.org/whl/cu128

For more information on installing on CUDA 11+ platform, or for instructions on building from source, please refer to the installation guide.

Offline Batch Inference

import lmdeploy
with lmdeploy.pipeline("internlm/internlm3-8b-instruct") as pipe:
    response = pipe(["Hi, pls intro yourself", "Shanghai is"])
    print(response)

[!NOTE] By default, LMDeploy downloads model from HuggingFace. If you would like to use models from ModelScope, please install ModelScope by pip install modelscope and set the environment variable:

export LMDEPLOY_USE_MODELSCOPE=True

If you would like to use models from openMind Hub, please install openMind Hub by pip install openmind_hub and set the environment variable:

export LMDEPLOY_USE_OPENMIND_HUB=True

For more information about inference pipeline, please refer to here.

Tutorials

Please review getting_started section for the basic usage of LMDeploy.

For detailed user guides and advanced guides, please refer to our tutorials:

Third-party projects

  • Deploying LLMs offline on the NVIDIA Jetson platform by LMDeploy: LMDeploy-Jetson

  • Example project for deploying LLMs using LMDeploy and BentoML: BentoLMDeploy

Contributing

We appreciate all contributions to LMDeploy. Please refer to CONTRIBUTING.md for the contributing guideline.

Acknowledgement

Citation

@misc{2023lmdeploy,
    title={LMDeploy: A Toolkit for Compressing, Deploying, and Serving LLM},
    author={LMDeploy Contributors},
    howpublished = {\url{https://github.com/InternLM/lmdeploy}},
    year={2023}
}
@article{zhang2025efficient,
  title={Efficient Mixed-Precision Large Language Model Inference with TurboMind},
  author={Zhang, Li and Jiang, Youhe and He, Guoliang and Chen, Xin and Lv, Han and Yao, Qian and Fu, Fangcheng and Chen, Kai},
  journal={arXiv preprint arXiv:2508.15601},
  year={2025}
}

License

This project is released under the Apache 2.0 license.

Contributors

zh-nj

6 commits

Languages

Python

63.9%

C++

22.6%

Cuda

12.4%