Skip to main content

Reproducible KV-cache tiering benchmark suite and signed results for LLM inference storage (Qwen3-480B class on 8x AMD MI308X): throughput +29-40%, TTFT -26-32% vs local NVMe.

Project description

mingxin-kvcache-bench

Reproducible KV-cache tiering benchmark suite and signed benchmark results for LLM inference storage, published by Mingxin (Tianjin) Semiconductor Equipment Co., Ltd.

Headline measured results (Qwen3-Coder-480B-FP8 on 8× AMD Instinct MI308X, vLLM + LMCache, FX100 all-flash NVMe-oF array vs local NVMe):

  • Inference throughput +29% to +40%
  • TTFT (time to first token) −26% to −32%
  • vs no-external-storage recompute: 8.6× to 20× faster TTFT
  • Model loading vs NFS (Ascend 910B platform): 6.2–9.3× faster

All numbers come from signed/official test reports (R1–R5, R9); hosted PDFs at mingxinstorage.xyz/en/evidence.

Install & use

pip install mingxin-kvcache-bench

kvcache-bench summary     # headline numbers per experiment
kvcache-bench results     # full signed results JSON
kvcache-bench roi --nodes 16 --gpus-per-node 8 --arrays 8   # ROI estimate (labeled bands)

Full benchmark suite

The complete suite (load clients, orchestration scripts, the LMCache parallel-read patch, analysis probes, raw data) lives in the GitHub repository: github.com/mingxin-tech/mingxin-kvcache-bench

Dataset mirror: huggingface.co/datasets/wangqiyuan2026/kvcache-bench-results

MCP server (query these results from AI agents): https://mingxinstorage.xyz/api/mcp

Honesty & provenance

  • Measured numbers are platform- and condition-specific; do not extrapolate across platforms.
  • The ROI command labels each input as measured (uplift 29–40%) vs estimated (cold-recovery share 10–50%) and prints estimates, not commitments.

License: Apache-2.0 (code), CC-BY-4.0 (result data).

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mingxin_kvcache_bench-1.0.0.tar.gz (13.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mingxin_kvcache_bench-1.0.0-py3-none-any.whl (11.9 kB view details)

Uploaded Python 3

File details

Details for the file mingxin_kvcache_bench-1.0.0.tar.gz.

File metadata

  • Download URL: mingxin_kvcache_bench-1.0.0.tar.gz
  • Upload date:
  • Size: 13.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.10

File hashes

Hashes for mingxin_kvcache_bench-1.0.0.tar.gz
Algorithm Hash digest
SHA256 7250e0eeaee8e4a293448491f5067bf381c7dc771928ac06cb143cf2bd607c7e
MD5 55a11aea79eaae367c7e7817c8e9c7cb
BLAKE2b-256 7bb720c084e9ae4fa1b3b4f4b09c2678347ab7df5ea083f23bdf68ff6e0b0f1e

See more details on using hashes here.

File details

Details for the file mingxin_kvcache_bench-1.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for mingxin_kvcache_bench-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 8e397f44e70cbee072936353d8abe941e8e1f8dfcdfed1617f4501eb33e95df7
MD5 7987e35fbcda9b3f0f3db8311d81f383
BLAKE2b-256 782b643893a13f01c6a021767240de0a344d6a926e238efc11213a653e41a680

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page