Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

UCM

| Documentation | Website | RoadMap | 中文 |

DeepWiki


Overview

The core principle of Unified Cache Manager (UCM) is to persist the LLM KVCache and replace redundant computations through multiple retrieval mechanisms. UCM not only supports prefix caching but also offers a variety of training-free sparse attention retrieval methods, delivering higher performance when handling extremely long sequence inference tasks. Additionally, UCM provides a PD disaggregation solution based on a storage-compute separation architecture, which enables more straightforward and flexible management of heterogeneous computing resources. When integrated with vLLM, UCM achieves a 3-10x reduction in inference latency across various scenarios, including multi-turn dialogue and long-context reasoning tasks.

Motivation

With the increase of model size, the KV cache became larger and sparser, especially for long sequence requests. To reduce the GPU memory used, offload full KV to external storage and only keep partial or compressed KV in GPU memory became the popular direction. This can also reduce the GPU calculation, increase the sequence length and batch size of decoding.

Sparse KV cache have many different choices. Recently paper point out that there is no common way can fit all scenarios and all models. So better to build a common framework then different sparse algorithms can be plugin to it like KV connector for PC.

architecture.png

All gray boxes in the diagram represent existing classes in vLLM version 0.9.2, while the green boxes indicate newly added components by UCM. The light green boxes demonstrate potential future subclass extensions based on this framework.

UcmSparseBase is the base class of different sparse algorithms. Just like KV connector design, it will hook few places of scheduler and layer.py to do additional load, dump and calculate sparse KV blocks.

SparseKVManager allows users to define custom KV block allocations for different algorithms. To keep all implementations unified under the SparseKVBase framework, the system calls the SparseKVBase base class, while the actual implementation occurs in subclasses of sparse algorithms.

KVStoreBase helps decouple sparse algorithms from external storage. It defines methods for communicating with external storage, enabling any sparse algorithm to work seamlessly with any external storage system. The core concept here involves identifying blocks through IDs and offsets. This approach is not only suitable for sparse scenarios but also naturally accommodates prefix caching. The KVStoreConnector links it with the current KVConnectorBase_V1 to provide PC (Prefix Caching) functionality. For example, NFSStore serves as a reference implementation that provides the capability to store KVCache in either a local filesystem for single-machine scenarios or through NFS mount points in multi-server environments.


Support Features

  • Prefix Cache
  • Cache Blend
  • Model Window Extrapolation
  • Prefill Offload
  • Sparse Attention
  • Sparse Attention Offload
  • Heterogeneous PD Disaggregation

Quick Start

please refer to Quick Start for vLLM and Quick Start for vLLM-Ascend.


Branch

Branch Status vLLM version vLLM-Ascend version
main Maintained v0.27.1 nightly-0.26.0
develop Maintained v0.27.1 nightly-0.26.0

Contact Us

  1. For technical questions and feature requests, please use GitHub Issues.
  2. WeChat technical discussion group: Scan the QR code below.
wechat-gh

License

UCM is licensed under the MIT with additional conditions. Please read the LICENSE file for details.

Metadata

Release files for uc-manager-cann910-a5 0.9.0rc1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for uc-manager-cann910-a5 0.9.0rc1
File Interpreter ABI Platform
uc_manager_cann910_a5-0.9.0rc1-cp312-cp312-manylinux_2_34_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.34+ x86-64 Details
uc_manager_cann910_a5-0.9.0rc1-cp312-cp312-manylinux_2_34_aarch64.whl CPython 3.12 CPython 3.12 Linux glibc 2.34+ ARM64 Details

Total release size: 10.5 MB

Release files / uc_manager_cann910_a5-0.9.0rc1-cp312-cp312-manylinux_2_34_x86_64.whl

Download URL uc_manager_cann910_a5-0.9.0rc1-cp312-cp312-manylinux_2_34_x86_64.whl
Size 5.4 MB
Tags CPython 3.12 Linux glibc 2.34+ x86-64
SHA-256 checksum
How to use checksums
b49e3cb9d8d1a8ab62b1895c690ab00b276db92579e2f5d47354127664949d80
BLAKE2b-256 checksum
How to use checksums
89d7052d50264f564501a73be411fbe86e3025bdaa9ac73aa89e5b5e6107a97d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.14

Release files / uc_manager_cann910_a5-0.9.0rc1-cp312-cp312-manylinux_2_34_aarch64.whl

Download URL uc_manager_cann910_a5-0.9.0rc1-cp312-cp312-manylinux_2_34_aarch64.whl
Size 5.1 MB
Tags CPython 3.12 Linux glibc 2.34+ ARM64
SHA-256 checksum
How to use checksums
9fccbab4ccaf3cc685b796317338bcbd81b27b434748b8dbb351430aa5c929d9
BLAKE2b-256 checksum
How to use checksums
d006a963cb44ed387adc8db70176b9b20bda221555c1f3aff502e0e8e806f6df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.14

Release history Release notifications | RSS feed

This release

0.9.0rc1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page