Easy, fast, and cheap LLM serving for everyone
Deprecation Notice:
vllm-cpu-avx512vnniv0.16.0 is the last release of this variant package. Starting with v0.17.0, all CPU ISA variants are unified into a single package with automatic ISA detection at runtime. Migrate to:pip install vllm-cpuSee github.com/MekayelAnik/vllm-cpu for details.
Buy Me a Coffee
Your support encourages me to keep creating/supporting my open-source projects. If you found value in this project, you can buy me a coffee to keep me up all the sleepless nights.
About
vLLM is a fast and easy-to-use library for LLM inference and serving. This PyPl package has VNNI (AVX512+VNNI) inference built in on supported CPUs.
Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has evolved into a community-driven project with contributions from both academia and industry.
vLLM is fast with:
- State-of-the-art serving throughput
- Efficient management of attention key and value memory with PagedAttention
- Continuous batching of incoming requests
- Fast model execution with VNNI on supported CPUs Use this package ONLY IF your CPU have avx512vnni or newer instruction sets
- Quantizations: GPTQ, AWQ, AutoRound, INT4, INT8, and FP8
- Optimized CPU kernels, including integration with FlashAttention and FlashInfer
- Speculative decoding
- Chunked prefill
vLLM is flexible and easy to use with:
- Seamless integration with popular Hugging Face models
- High-throughput serving with various decoding algorithms, including parallel sampling, beam search, and more
- Tensor, pipeline, data and expert parallelism support for distributed inference
- Streaming outputs
- OpenAI-compatible API server
- Support for x86_64, PowerPC CPUs, Arm CPUs and Applie Scilicon (CPU inference). This package does not support any GPU inference. For GPU inference support use the official vLLM PypI
- Prefix caching support
- Multi-LoRA support
vLLM seamlessly supports most popular open-source models on HuggingFace, including:
- Transformer-like LLMs (e.g., Llama)
- Mixture-of-Expert LLMs (e.g., Mixtral, Deepseek-V2 and V3)
- Embedding Models (e.g., E5-Mistral)
- Multi-modal LLMs (e.g., LLaVA)
Find the full list of supported models here.
Important Notes
- Install this package on Linux envirenment only. For Windows you will have to use WSL2 or later
- This package has a Container.io (Docker/Podman etc.) compatible image in Docker Hub
- Apache Licence of main vLLM project
- GPL License of this CPU specific vLLM package
- For versions 0.8.5–0.12.0, use
.post2releases (e.g.,pip install vllm-cpu-avx512vnni==0.12.0.post2) — includes critical CPU platform detection fix
Platform Detection Fix (versions 0.8.5 - 0.12.0)
If you encounter RuntimeError: Failed to infer device type or see UnspecifiedPlatform warnings with versions 0.8.5 to 0.12.0, run this one-time fix after installation:
import os, sys, importlib.metadata as m
v = next((d.metadata['Version'] for d in m.distributions() if d.metadata['Name'].startswith('vllm-cpu')), None)
if v:
p = next((p for p in sys.path if 'site-packages' in p and os.path.isdir(p)), None)
if p:
d = os.path.join(p, 'vllm-0.0.0.dist-info'); os.makedirs(d, exist_ok=True)
open(os.path.join(d, 'METADATA'), 'w').write(f'Metadata-Version: 2.1\nName: vllm\nVersion: {v}+cpu\n')
print(f'Fixed: vllm version set to {v}+cpu')
This creates a package alias so vLLM detects the CPU platform correctly. Only needed once per environment. Versions 0.8.5.post2+ and 0.12.0+ include this fix automatically.
Getting Started
Install vLLM with a single command:
pip install vllm-cpu-avx512vnni --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple
This installs vllm-cpu-avx512vnni with CPU-optimized PyTorch (no CUDA dependencies).
Alternative: Using uv (faster)
uv pip install vllm-cpu-avx512vnni --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple
Install uv on Linux:
curl -LsSf https://astral.sh/uv/install.sh | sh
Docker Images
Pre-built Docker images are available on Docker Hub and GitHub Container Registry.
# Pull from Docker Hub
docker pull mekayelanik/vllm-cpu:avx512vnni-latest
# Or from GitHub Container Registry
docker pull ghcr.io/mekayelanik/vllm-cpu:avx512vnni-latest
# Run OpenAI-compatible API server
docker run -p 8000:8000 \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
mekayelanik/vllm-cpu:avx512vnni-latest \
--model facebook/opt-125m
Available tags: avx512vnni-latest, avx512vnni-<version> (e.g., avx512vnni-0.12.0)
Platforms: linux/amd64
vllm-cpu
This CPU specific vLLM has 5 optimized wheel packages from the upstream vLLM source code:
| Package | Optimizations | Target CPUs |
|---|---|---|
vllm-cpu |
Baseline (no AVX512) | All x86_64 and ARM64 CPUs |
vllm-cpu-avx512 |
AVX512 | Intel Skylake-X and newer |
vllm-cpu-avx512vnni |
AVX512 + VNNI | Intel Cascade Lake and newer |
vllm-cpu-avx512bf16 |
AVX512 + VNNI + BF16 | Intel Cooper Lake and newer |
vllm-cpu-amxbf16 |
AVX512 + VNNI + BF16 + AMX | Intel Sapphire Rapids (4th gen Xeon+) |
Each package is compiled with specific CPU instruction set flags for optimal inference performance.
Check Your CPU & Get Install Command
pkg=vllm-cpu
grep -q avx512f /proc/cpuinfo && pkg=vllm-cpu-avx512
grep -q avx512_vnni /proc/cpuinfo && pkg=vllm-cpu-avx512vnni
grep -q avx512_bf16 /proc/cpuinfo && pkg=vllm-cpu-avx512bf16
grep -q amx_bf16 /proc/cpuinfo && pkg=vllm-cpu-amxbf16
printf "\n\tRUN:\n\t\tuv pip install $pkg\n"
Example list of CPUs with their supported instruction sets
| CPU Architecture (Intel/AMD) | AVX2 | AVX-512 F (Base) | VNNI (INT8) | BF16 (BFloat16) (via AVX-512) | AMX-BF16 (via Tile Unit) |
|---|---|---|---|---|---|
| Intel 4th Gen / AMD Ryzen Zen2 & Newer | Yes | No | No | No | No |
| Intel Skylake-SP / Skylake-X / AMD Zen 4 & Newer | Yes | Yes | No | No | No |
| Intel Cooper Lake (3rd Gen Xeon) / AMD Zen 4 (EPYC) / Ryzen Zen5 & Newer | Yes | Yes | Yes | Yes | No |
| Intel Sapphire Rapids (4th Gen Xeon) & Newer | Yes | Yes | Yes | Yes | Yes |
***Currently no AMD CPU support AMXBF16. AMD expected to include AMXBF16 support from AMD Zen 7 CPUs
Buy Me a Coffee
Your support encourages me to keep creating/supporting my open-source projects. If you found value in this project, you can buy me a coffee to keep me up all the sleepless nights.
Release files for vllm-cpu-avx512vnni 0.16.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vllm_cpu_avx512vnni-0.16.0-cp38-abi3-manylinux_2_28_x86_64.whl | CPython 3.8 | abi3 | Linux glibc 2.28+ x86-64 | Details |
Release files / vllm_cpu_avx512vnni-0.16.0-cp38-abi3-manylinux_2_28_x86_64.whl
| Download URL | vllm_cpu_avx512vnni-0.16.0-cp38-abi3-manylinux_2_28_x86_64.whl |
|---|---|
| Size | 31.6 MB |
| Tags | CPython 3.8 Linux glibc 2.28+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
2ec3e43d89658cc828bcc1e7ce1e1c7a07570456f90fc15e3e372b0051b0b14d
|
|
BLAKE2b-256 checksum How to use checksums |
370495f6c1ce65fe978078514ae39985b7806ccd96ff3df3717d28d67ac50cba
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|