LLM gateway traffic profiler: prefix-reuse analysis, cache-hit-rate upper bounds, and anonymized replayable trace export
Project description
llmtrafficlens
LLM gateway traffic profiler — analyze prefix-reuse structure in your gateway logs, estimate the prefix-cache hit-rate ceiling before you build KV-cache-aware routing, and export anonymized replayable traces in community-standard formats.
Every KV-cache product tells you how much you saved after deployment.
llmtrafficlens answers the prior question: is your traffic worth it, and
what is the upper bound?
What it does
- Traffic profile — input/output token distributions, streaming ratio, model mix, error rates.
- Prefix structure — distinct prefixes, shared-prefix groups, and a dual-metric leaderboard: top prefixes by reuse count vs by reusable token volume. The two heads routinely differ: a 15-token prefix reused 870× is worthless to a cache, a 19K-token prefix reused 95× dominates it.
- Session-shape detection — presence rate of session identifiers
(
prompt_cache_key,metadata.*): tells you whether reuse is cross-request shared-prefix shaped (needs prefix-aware routing) or intra-session shaped (session affinity suffices). - Reuse potential — ECHR (engine cache hit rate) upper bound under infinite cache + perfect routing, plus an LRU capacity/hit-rate curve.
- Anonymized export — qwen-bailian-superset or Mooncake JSONL (salted chained block hashes; no text leaves your machine), directly replayable by NVIDIA AIPerf, SGLang bench_serving, and AIBrix.
- Generator fitting — emit SGLang
generated-shared-prefixCLI args (zipf popularity fitted to your histogram) with a fidelity report.
Install
pip install llmtrafficlens
Zero runtime dependencies. Python ≥ 3.10.
Quickstart on a public trace
Validate the analyzer against the public Mooncake trace (FAST'25) — the result should match the official KV Cache Hit Rate Simulator's infinite-capacity ceiling:
curl -LO https://raw.githubusercontent.com/kvcache-ai/Mooncake/main/FAST25-release/traces/conversation_trace.jsonl
llmtrafficlens validate conversation_trace.jsonl --format mooncake
llmtrafficlens profile conversation_trace.jsonl --format mooncake -o mooncake-report
Profile your own gateway logs
Input: a CSV with a request_json column holding the OpenAI-format request
body (optional response_json, model_name, status_code columns are
picked up when present).
llmtrafficlens profile gateway.csv -o report
# -> report.json + report.html
All processing is local. The report contains aggregates only — raw prompts never appear in any output.
Token counts use a chars/4 approximation unless the log carries usage
(then exact counts are used and marked in the report). A keyed hash salt is
generated at .ltl-salt on first run; keep it stable to make runs and
exports comparable, and treat it as a secret.
Export a replayable anonymized trace
# primary: bailian superset (16-token blocks, session structure, stream/status)
llmtrafficlens export gateway.csv --to bailian -o trace.jsonl
# downgrade: Mooncake four-field format
llmtrafficlens export gateway.csv --to mooncake -o mooncake.jsonl
Replay examples:
# NVIDIA AIPerf (timestamped replay)
aiperf profile --custom-dataset-type mooncake_trace --input-file mooncake.jsonl --fixed-schedule ...
# SGLang bench_serving
python -m sglang.bench_serving --dataset-name mooncake --dataset-path mooncake.jsonl ...
Fit a load generator to your traffic
llmtrafficlens fit gateway.csv --target sglang-gsp
Emits --gsp-* arguments (group count, zipf alpha, prompt lengths) plus a
fidelity report: the theoretical ECHR the synthetic config would
produce vs your measured upper bound, so the flattening loss is explicit.
Reading the numbers
- ECHR upper bound counts cross-request shared prefixes only, under infinite cache / perfect routing / no eviction. Real deployments land below it; intra-session reuse is invisible unless your sample carries session identifiers.
- Session identifier presence 0% means session-affinity routing cannot capture any of the observed reuse — a prefix-aware policy is required.
- Block size defaults to 16 tokens (qwen-bailian convention; finer and engine-page aligned). Mooncake input implies 512.
Positioning
Observability platforms (Langfuse, Helicone, Datadog LLM) stop at
token/cost counting. Trace analyzers (Mooncake simulator, AIPerf
analyze-trace) require pre-hashed traces. Simulators (Vidur,
SplitwiseSim) consume traces nobody produces from production logs.
llmtrafficlens covers the missing pipeline: raw gateway logs →
structural analysis → decision numbers → standard replayable trace.
License
Apache-2.0
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llmtrafficlens-0.1.0.tar.gz.
File metadata
- Download URL: llmtrafficlens-0.1.0.tar.gz
- Upload date:
- Size: 17.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
38726df0117ea7ff9cf145ef8c6863de156d220f8479fa02f2ea4722825c88de
|
|
| MD5 |
df595d5b40e37ea951461e541c090d5b
|
|
| BLAKE2b-256 |
0a8af91efa9dc9e54166f0be8fe2907cf1e843778023a77ea0d6e87d7334d7e8
|
File details
Details for the file llmtrafficlens-0.1.0-py3-none-any.whl.
File metadata
- Download URL: llmtrafficlens-0.1.0-py3-none-any.whl
- Upload date:
- Size: 21.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
04dc49bed2b0483388d6f5a5d3a273d428daa6c80e801e24a43d140b75d15cc2
|
|
| MD5 |
0b27bcf68c3b13b796ebee5470463fd3
|
|
| BLAKE2b-256 |
87b87bbaa2ce113681e550a5e42a6e766168921a26a003f0a63be3c66faac3bb
|