lmcache-cli

A LLM serving engine extension to reduce TTFT and increase throughput, especially under long-context scenarios.

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

deng451e

These details have not been verified by PyPI

Project links

homepage

Project description

lmcache logo

A KV Cache Management Layer for Scalable LLM Inference

Blog | Documentation | Join Slack | Community Meeting | Roadmap

Updates

[2026/05] 🔥 Agentic workload benchmark on AMD MI300X (blog).
[2026/04] 🔥 LMCache's new multiprocess(MP) architecture release (blog).
[2026/03] LMCache at GTC 2026 (post).
[2026/01] LMCache multi-node P2P CPU memory sharing, from experimental feature to production (blog).

[2025/11] LMCache x CoreWeave accelerate efficient LLM inference for Cohere (blog).
[2025/10] LMCache joins the PyTorch Foundation and Tensormesh unveiled (blog, PyTorch).
[2025/09] NVIDIA Dynamo integrates LMCache, accelerating LLM inference (blog).
[2025/08] 🎉 LMCache hits 5,000+ GitHub stars (blog).
[2025/08] LMCache supports gpt-oss (20B/120B) on day 1 (blog).
[2025/07] Get faster LLM inference and cheaper responses with LMCache and Redis (Redis blog).
[2025/07] LMCache extends its turbo-boost to multimodal models in vLLM V1 (blog).
[2025/06] LLM Production Stack goes cross-hardware: AMD, Arm and Ascend (blog).

About

LMCache is a KV cache management layer for LLM inference. It turns KV cache from a temporary state into reusable AI-native knowledge that can be stored persistently, reused across multiple serving engines, monitored with an observability stack, and transformed for better generation quality. As a result, LMCache reduces TTFT (time-to-first-token) and improves throughput, especially for long-context agentic, multi-turn conversation, and knowledge-augmented workloads (e.g., RAG).

LMCache is vendor-neutral. It can be used as a KV cache layer for a range of mainstream open-source serving engines, inference frameworks, hardware vendors, storage systems, and infrastructure providers. The vendor neutrality allows users to freely switch between serving engines and storage vendors, while reusing the stored KV caches.

LMCache Deployment Modes

Key features

Engine-independent deployment: LMCache, as a standalone daemon process, manages KV cache independently from the inference engine process, so that KV cache will not be lost even if the inference engine crashes (i.e., no fate-sharing with engines).
Persistent, tiered KV cache offloading and reuse: Move KV caches out of GPU memory into a tiered storage hierarchy spanning CPU memory, local storage, and remote backends, enabling reuse across requests, sessions, and engine instances to reduce repeated prefill computation and improve TTFT.
Production-level KV cache observability: LMCache provides a rich set of KV cache observability metrics, including typical Kubernetes metrics (health monitoring, performance diagnostics), KV-cache-specific metrics (request-level and token-level prefix cache hits, lifecycle, request-level KV cache performance), management metrics (user-specific usage), and more.
Pluggable storage and transport backends: Easily integrate remote storage and KV transfer backends through a unified interface, enabling KV cache offloading and sharing across storage providers. Through this interface, LMCache supports storage backends including CPU RAM, local disk (SSD), Redis/Valkey, Mooncake, InfiniStore, S3-compatible object storage, NIXL, and GDS.
Non-prefix KV reuse: Extend KV reuse beyond prefix caching by reusing cached KV blocks at any position in the prompt. This leverages CacheBlend to selectively recompute tokens for quality recovery.
PD disaggregation and KV transfer: Support KV cache transfer from prefill workers to decode workers over NVLink, RDMA, or TCP through transport layers such as NIXL.
Pluggable KV transformation: A simple interface for researchers to write compression, token dropping, and custom serialization through a flexible SERDE interface.

LMCache is becoming an integral layer in the LLM inference ecosystem, with community-driven integration with serving engines, inference frameworks, hardware vendors, storage systems, and infrastructure providers:

LMCache ecosystem

Getting Started

To use LMCache, simply install lmcache from your package manager, e.g. pip:

pip install lmcache

For more setup options and examples, see:

Contributing

We welcome and value contributions and collaborations. Join us in improving LMCache. Check out the Contributing Guide or join our Slack community to get started.

Adoption and Partnerships

LMCache has a growing community of developers, researchers, industry adopters, and partners building the next generation of efficient LLM inference systems.

LMCache Adoption and Partnerships

As an independent open-source project, LMCache is becoming the de-facto standard for KV Cache management in LLM inference. Its continued development and community work are supported in part by Tensormesh.

Citation

LMCache builds on research in KV cache management, including cache reuse, offloading, compression, and serving optimization. If you use LMCache in your research, please cite the LMCache paper and related work.

@article{cheng2025lmcache,
  title={LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference},
  author={Cheng, Yihua and Liu, Yuhan and Yao, Jiayi and An, Yuwei and Chen, Xiaokun and Feng, Shaoting and Huang, Yuyang and Shen, Samuel and Du, Kuntai and Jiang, Junchen},
  journal={arXiv preprint arXiv:2510.09665},
  year={2025}
}

License

The LMCache codebase is licensed under Apache License 2.0. See the LICENSE file for details.

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

deng451e

These details have not been verified by PyPI

Project links

homepage

Release history Release notifications | RSS feed

This version

0.5.1.dev0 pre-release

Jun 23, 2026

0.5.0rc2.dev0 pre-release

Jun 23, 2026

0.4.8.dev42 pre-release

Jun 18, 2026

0.4.8.dev38 pre-release

Jun 18, 2026

0.4.8.dev31 pre-release

Jun 17, 2026

0.4.8.dev24 pre-release

Jun 16, 2026

0.4.8.dev21 pre-release

Jun 16, 2026

0.4.6.dev0 pre-release

May 14, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

lmcache_cli-0.5.1.dev0-py3-none-any.whl (2.0 MB view details)

Uploaded Jun 23, 2026 Python 3

File details

Details for the file lmcache_cli-0.5.1.dev0-py3-none-any.whl.

File metadata

Download URL: lmcache_cli-0.5.1.dev0-py3-none-any.whl
Upload date: Jun 23, 2026
Size: 2.0 MB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for lmcache_cli-0.5.1.dev0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`4904f6fe7ef5724d670d69705d149df6b97caa2decdf19d87051502aee32a96f`
MD5	`d0026df0b2377910df64a5b5d84b9c99`
BLAKE2b-256	`ab302be451393f34b0eb2350641f07f73a67ad6cc0ca62d3e1d33aef6a21ad79`

See more details on using hashes here.

Provenance

The following attestation bundles were made for lmcache_cli-0.5.1.dev0-py3-none-any.whl:

Publisher: publish.yml on LMCache/LMCache

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: lmcache_cli-0.5.1.dev0-py3-none-any.whl
- Subject digest: 4904f6fe7ef5724d670d69705d149df6b97caa2decdf19d87051502aee32a96f
- Sigstore transparency entry: 1930950949
- Sigstore integration time: Jun 23, 2026
Source repository:
- Permalink: LMCache/LMCache@389b9cfcb5cec502dbba6b5d13725bf2e024610d
- Branch / Tag: refs/tags/v0.5.0
- Owner: https://github.com/LMCache
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@389b9cfcb5cec502dbba6b5d13725bf2e024610d
- Trigger Event: release

lmcache-cli 0.5.1.dev0

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

A KV Cache Management Layer for Scalable LLM Inference

Blog | Documentation | Join Slack | Community Meeting | Roadmap

Updates

About

Key features

Getting Started

Contributing

Adoption and Partnerships

Citation

License

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distributions

Built Distribution

File details

File metadata

File hashes

Provenance