Skip to main content

Peer-to-peer RDMA zero-copy L3 KV-cache backend for SGLang HiCache

Project description

PeerCache

CI Docs License

A lightweight, peer-to-peer L3 storage backend for SGLang HiCache, built for PD-disaggregated (prefill/decode) inference: prefill workers publish KV pages, decode workers read them back over RDMA with zero CPU copies.

Docs: https://flymysql.github.io/PeerCache/

PeerCache gives you Mooncake-style RDMA zero-copy KV-cache sharing across nodes, but without the centralized master + metadata services. Instead it uses:

  • Embedded service discovery — no separate meta process. One node (chosen by discovery_addr) auto-hosts the discovery service in-process; nodes register their endpoint, heartbeat, and pull the live membership list.
  • A consistent-hash distributed directory (DHT) — the mapping key -> {data_node, remote_addr, rkey, length} is sharded across all nodes by hashing the key. There is no central metadata store.
  • Data stays local on writeset() copies the page into a node-local published pool (a host memcpy, no network, no master) and pushes only a tiny location record to the directory.
  • One-sided RDMA READ on readget() looks up the directory, then issues a zero-copy IBV_WR_RDMA_READ straight into SGLang's registered host buffer.
  • Disk persistence tier (L4) — pages evicted from memory spill to disk (default /data/peercache/, 100GB) and are promoted back into the pool on a later read (locally or by a remote reader).
  • Built-in monitoring — Prometheus /metrics + an embedded HTML dashboard (default port 31997): hit rate, throughput, latency p50/p99, mem/disk usage.
write:  set() ── local memcpy ──> published pool MR
                └── PUT key->{node,addr,rkey,len} ──> directory shard (hash(key))
read:   get() ── GET key ──> directory shard ──> {node,addr,rkey,len}
                └── one-sided RDMA READ ──> local host buffer (zero copy)

Why simpler than Mooncake?

Mooncake PeerCache
metadata central master + metadata service sharded directory (consistent hash)
data placement dedicated managed pool stays on producing node
coordination master allocates / tracks objects only service discovery on meta node
transfer RDMA zero-copy RDMA zero-copy (one-sided READ)

Architecture

  • C++ data plane (cpp/): raw libibverbs + librdmacm. RC QPs, one-sided READ/WRITE, CQ polling, lazy per-peer connection pooling. Exposed to Python via pybind11 as the _peercache module.
  • Python control plane (python/peercache/): TCP RPC, service discovery, consistent-hash ring, distributed directory, and the published-pool with LRU.
  • TCP fallback transport: a pure-Python transport that mirrors the RDMA API so the design can be validated end-to-end on machines without RDMA hardware.

Two-MR model (correctness)

SGLang's host KV buffer is the L2 tier and is evicted/overwritten by HiCache, so we cannot register its address into the directory directly (dangling reference). Each node therefore registers two memory regions:

  1. Receive MR = mem_pool_host.kv_buffer — destination of one-sided READ on get.
  2. Published pool MR = a backend-owned host pool with LRU — source of READ on remote nodes. set memcpys the page into this pool (node-local, no network) and publishes its addr+rkey+len to the directory. Eviction from the pool deletes the corresponding directory entry, so a published address stays valid until evicted.

Install

# Linux with RDMA (Mellanox OFED / rdma-core dev headers installed)
pip install .

# Without RDMA (control-plane + TCP fallback only, e.g. for tests on a laptop)
pip install -e . --config-settings=cmake.define.PEERCACHE_NO_RDMA=ON

Run with SGLang

The meta service is embedded — there is no separate meta process. Point discovery_addr at one node's IP on every node; the node whose IP matches auto-starts the discovery service in-process.

# On every SGLang node, set discovery_addr to the SAME node's IP (say node-0).
# node-0 detects the IP is itself and hosts the embedded meta automatically.
python -m sglang.launch_server --enable-hierarchical-cache \
  --hicache-storage-backend dynamic \
  --hicache-storage-backend-extra-config \
  '{"backend_name":"peercache","module_path":"peercache.store","class_name":"PeerCacheStore","discovery_addr":"NODE0_IP:9100","protocol":"rdma","device_name":"mlx5_0","global_segment_size":"4gb"}'

(Optionally, you can still run a standalone meta with peercache-meta --bind 0.0.0.0:9100 if you prefer a dedicated discovery host.)

See examples/sglang_launch.md for details.

Test

pip install pytest
PYTHONPATH=python pytest tests/ -v

Maintainer setup (one-time)

  • GitHub Pages: Settings → Pages → Build and deployment → Source = GitHub Actions. The Docs workflow then publishes to https://flymysql.github.io/PeerCache/ on every push to main.
  • PyPI Trusted Publishing: on the PyPI peercache project, add a GitHub publisher (owner flymysql, repo PeerCache, workflow release.yml, environment pypi). Tagging vX.Y.Z then builds the sdist, attaches it to a GitHub Release, and publishes to PyPI. Until configured, the PyPI step is non-blocking and the GitHub Release still ships the package.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

peercache-0.2.0.tar.gz (73.8 kB view details)

Uploaded Source

File details

Details for the file peercache-0.2.0.tar.gz.

File metadata

  • Download URL: peercache-0.2.0.tar.gz
  • Upload date:
  • Size: 73.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for peercache-0.2.0.tar.gz
Algorithm Hash digest
SHA256 fae502e9e1bfd18b915ae801b97db894e3c9530a91bca0111fd6395f0a6df820
MD5 57b99aac562c5b11ab1202065b105997
BLAKE2b-256 11cf1e1422f3b32ae8a5d7c6db2c5a63767fe6e9c3048ecad2ce5b10ce20038c

See more details on using hashes here.

Provenance

The following attestation bundles were made for peercache-0.2.0.tar.gz:

Publisher: release.yml on flymysql/PeerCache

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page