memkv-sglang
sglang HiCacheStorage backend that persists prefix KV pages in a
remote MemKV cluster. Loaded as a vendor plugin via sglang's built-in
dynamic storage backend dispatch — no patches to sglang's tree.
Build
cd sglang-plugin
pip install maturin
maturin develop --release # local dev install
# or
maturin build --release # wheel under target/wheels/
pip install target/wheels/memkv_sglang-*.whl
The wheel bundles a native PyO3 extension built from the same
memkv-client crate the NIXL plugin uses, so RDMA/TCP transport
selection works the same way.
Configure the MemKV connection
The plugin reads the standard MemKV config chain — MEMKV_CONFIG
yaml first, then MEMKV_* env vars:
export MEMKV_SERVERS="10.0.0.10:9900,10.0.0.11:9900"
export MEMKV_RDMA_DEVICES="mlx5_0,mlx5_1"
export MEMKV_AUTH_KEY="<64-hex>"
# optional:
# export MEMKV_TRANSPORT=auto # rdma | tcp | auto (default)
# export MEMKV_CONFIG=/etc/memkv.yaml
Launch sglang against MemKV
Use sglang's dynamic storage backend, pointing it at the class in
this package:
python -m sglang.launch_server \
--model-path meta-llama/Llama-3-8B \
--enable-hierarchical-cache \
--hicache-storage-backend dynamic \
--hicache-storage-backend-extra-config '{
"backend_name": "memkv",
"module_path": "memkv_sglang.backend",
"class_name": "MemKVHiCacheStorage"
}'
sglang's StorageBackendFactory._create_dynamic_backend imports the
class and constructs it as MemKVHiCacheStorage(storage_config, kwargs).
What's implemented
| Method | Status |
|---|---|
get / batch_get |
yes; RDMA zero-copy direct into the target tensor when eligible (Linux + CPU + contiguous), bytes path otherwise |
set / batch_set |
yes |
exists / batch_exists |
yes (router-aware, one batched RPC per server) |
batch_exists_v2 |
yes (per-pool hit policies) |
batch_get_v2 |
yes; RDMA zero-copy into the dummy flat page when eligible |
batch_set_v2 |
yes |
clear |
no-op (server manages retention) |
Layout
sglang-plugin/
├── Cargo.toml # cdylib + pyo3 + memkv-client
├── pyproject.toml # maturin
├── src/lib.rs # PyO3 wrapper around memkv-client::Engine
└── python/memkv_sglang/
├── __init__.py # re-exports Client
└── backend.py # MemKVHiCacheStorage(HiCacheStorage)
Metadata
Release files for memkv-sglang 1.0.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| memkv_sglang-1.0.6-cp38-abi3-manylinux_2_28_x86_64.whl | CPython 3.8 | abi3 | Linux glibc 2.28+ x86-64 | Details |
| memkv_sglang-1.0.6-cp38-abi3-manylinux_2_28_aarch64.whl | CPython 3.8 | abi3 | Linux glibc 2.28+ ARM64 | Details |
Total release size: 2.9 MB
Release files / memkv_sglang-1.0.6-cp38-abi3-manylinux_2_28_x86_64.whl
| Download URL | memkv_sglang-1.0.6-cp38-abi3-manylinux_2_28_x86_64.whl |
|---|---|
| Size | 1.5 MB |
| Tags | CPython 3.8 Linux glibc 2.28+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
57f2b8297bdada3212fbf6774a2ec975ccf7637c4eb88d6ff486b91772d36135
|
|
BLAKE2b-256 checksum How to use checksums |
edb2496e5af4c21dcac4c636bd6cfd11df9a0720ba56a68d7c28befd8e5779d9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.20
|
Release files / memkv_sglang-1.0.6-cp38-abi3-manylinux_2_28_aarch64.whl
| Download URL | memkv_sglang-1.0.6-cp38-abi3-manylinux_2_28_aarch64.whl |
|---|---|
| Size | 1.4 MB |
| Tags | CPython 3.8 Linux glibc 2.28+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
9de4f12475fcf5ed2632bb5a7a5e4dfd85f7c75a3b22d19e6e59da5e33b4c1e8
|
|
BLAKE2b-256 checksum How to use checksums |
760a2b23dd4ff32fa654ff554fa601ecea8cd6d91d922519fe501645b211631d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.20
|