KVCalib
KVCalib turns "how aggressively can I evict my KV cache" into a decision backed by a distribution-free statistical guarantee, calibrated on your own workload, instead of a benchmark number computed on someone else's traffic.
It is a thin calibration/monitoring layer on top of kvpress
(NVIDIA's KV-cache eviction backends and benchmark harness) -- not a new eviction
heuristic, not a fork of kvpress, and not a competitor to it. See
docs/architecture.md for the full design rationale and
docs/understanding_the_guarantee.md for what
the guarantee does and doesn't promise, including how it differs from H2O's own
approximation-ratio guarantee.
Status: Phase 2 complete. Multi-backend: H2O, SnapKV, and StreamingLLM are all
validated (monotonicity + repeat-split coverage) on exact-match loss; PyramidKV is
implemented but documented as excluded on non-CUDA hardware (see
docs/monotonicity_reports/pyramidkv_exact_match.md).
Cross-backend comparison: SnapKV achieves the smallest certified budget at alpha=0.80
(see Cross-backend comparison). Metric-degradation loss
is implemented but parked, not validated -- its own monotonicity check failed on
real data (see
docs/monotonicity_reports/h2o_metric_degradation.md).
An ACI-based drift monitor (kvcalib.monitor.aci) is implemented and validated
against a real distribution shift (see Negative controls).
Phase 3 complete: MondrianRiskControl adds group-conditional (Mondrian-style)
calibration, analysis/reporting only for now, validated on real data at both a
genuine budget tightening and a genuine calibration-data cost (see
Group-conditional coverage). Phase 4 complete:
kvcalib.telemetry.otel_export adds opt-in OpenTelemetry span/event export (see
Install); kvcalib.monitor.streaming.StreamingCalibrator adds
windowed recalibration triggered by the existing drift monitor, not a new online
calibration algorithm; score-threshold eviction knobs were considered and scoped
out (all four backends and kvpress's shared scoring base are rank-based, with
no score-cutoff quantity to expose); packaging polish and a PyPI Trusted
Publishing release workflow round out the phase. See
KVCALIB_PROJECT_SPEC.md for full phase-by-phase
detail and CHANGELOG.md for the versioned history.
Install
pip install "kvcalib[kvpress]"
Not yet published to PyPI -- until the first release goes out, install from a checkout instead:
pip install -e ".[kvpress]" # from a checkout, until the PyPI release is published
Known upstream issue on Python 3.13: kvpress 0.5.4 (latest as of this writing)
transitively depends on google-fire, which still imports the pipes stdlib module
removed in Python 3.13 -- import kvpress fails with ModuleNotFoundError: No module named 'pipes' on 3.13, regardless of anything in this project. This is a kvpress/
fire compatibility gap, not a KVCalib issue -- kvcalib's own package imports lazily
and does not hit it. Until it's fixed upstream, either use Python 3.10-3.12 for the
[kvpress] extra, or work around it locally with a one-line shim:
echo 'from shlex import quote' > "$(python3 -c "import sysconfig; print(sysconfig.get_paths()['purelib'])")/pipes.py".
Known upstream limitation: PyramidKVBackend cannot run generation on non-CUDA
hardware. PyramidKV gives each transformer layer a different KV-cache length by
design; transformers' eager/sdpa/flex_attention paths build one causal mask
(sized for one assumed cache length) and reuse it for every layer, which crashes
whenever a prompt/budget pair produces non-uniform per-layer lengths (common, not an
edge case). Only flash_attention_2/_3 are exempt, and both are hard-gated in
transformers to CUDA/HIP/Cambricon-MLU hardware -- unavailable on Apple Silicon
regardless of attention-kernel package. See
docs/monotonicity_reports/pyramidkv_exact_match.md
for the full investigation, including two remediation paths (PyTorch 2.13's MPS
FlexAttention, the mps-flash-attn package) that were concretely tried and ruled out
rather than assumed incompatible. PyramidKVBackend is implemented and its
budget-to-ratio translation is correct -- this is a kvpress/transformers
limitation on this hardware, not a bug in this project -- but it is excluded from
Phase 2's coverage validation and cross-backend comparison as a result.
Calibration and deployment require kvpress + transformers + torch and an actual
model -- there is no synthetic substitute for "run the model," since the whole point
is measuring eviction's real effect. The core calibration math
(kvcalib.core) has no such dependency and can be tested and used standalone (e.g.
against precomputed loss curves).
Optional: pip install "kvcalib[otel]" (or .[otel] from a checkout) adds
opentelemetry-api for kvcalib.telemetry.otel_export's span/event export
(export_calibration_span, export_generation_event) -- this is opt-in and
caller-driven, not required for calibration or deployment.
Quickstart
from transformers import AutoModelForCausalLM, AutoTokenizer
from kvcalib import CalibrationPrompt, Calibrator
from kvcalib.backends import H2OBackend
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct", attn_implementation="eager")
calibration_set = [
CalibrationPrompt(prompt="...", reference_answer="..."),
# ...prompts representative of your own workload...
]
calibrator = Calibrator(alpha=0.1, backend=H2OBackend())
calibrator.fit(calibration_set, model, budget_grid=[8, 16, 32, 64, 512], tokenizer=tokenizer)
calibrator.save("calibration.json")
print(calibrator.guarantee)
# --- at deploy time ---
calibrator = Calibrator.load("calibration.json")
cache = calibrator.risk_controlled_cache() # drop-in kvpress-compatible press
from kvpress import KVPressTextGenerationPipeline
pipe = KVPressTextGenerationPipeline(model=model, tokenizer=tokenizer)
print(pipe("Context: ...", question="...", press=cache, max_new_tokens=10))
Running examples/single_backend_quickstart.py end to end (a tiny, 14-prompt
calibration set of short factual questions) prints:
Expected exact_match accuracy loss from evicting to budget=8 tokens is <= 0.1,
assuming calibration and live traffic are exchangeable. Calibrated on n=14 prompts
on 2026-07-22 using the h2o backend. This guarantee holds only if calibration and
live traffic are exchangeable [...]
Generated: Madrid
See examples/ for this quickstart, a cross-backend comparison, and a fast
(no-GPU) walkthrough of the negative-control refusal path.
Coverage validation
This is the project's actual evidence, not a benchmark score borrowed from someone
else's traffic: real model + real kvpress H2O eviction, run on LongBench
multifieldqa_en, then repeat-split many times from that one real, measured loss
matrix (see tests/coverage_validation/, pytest tests/coverage_validation -m gpu).
Setup: Qwen2.5-0.5B-Instruct, H2O backend, exact-match-with-containment loss,
n=39 calibration prompts (all multifieldqa_en records whose full document fits
under the 4000-token budget grid ceiling -- see
docs/monotonicity_reports/h2o_exact_match.md
for why), 70/30 calibration/test split, 200 repeat-splits per alpha.
| alpha target | result | mean realized loss (held-out) | n_calib | n_test |
|---|---|---|---|---|
| 0.01 | refused -- InsufficientCalibrationDataError (n_calib=27 < required 100) |
-- | 27 | 12 |
| 0.05 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.10 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.20 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.60 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.80 | genuine, non-trivial budget selection -- 512, 1024, or 4096 depending on the split | 0.722 | 27 | 12 |
These are the real, reported numbers -- not illustrative ones. Two things worth being upfront about, since both are legitimate outcomes this tool is specifically designed to surface rather than hide:
- alpha=0.01 is refused, not silently miscalibrated. With only 27 calibration
prompts in a 70/30 split of this 39-prompt sample, the derived minimum-n floor for
alpha=0.01 is 100 (
kvcalib.core.validation.minimum_calibration_size); the guard fires and the tool says so, rather than returning a technically-computed but meaningless budget. - alpha=0.05/0.10/0.20/0.60 all fall back to the largest budget in the grid, with an identical realized loss, in every one of the 200 repeat-splits. Qwen2.5-0.5B- Instruct's own baseline exact-match accuracy on this task -- even with no eviction at all -- has a loss around 0.56 (see the monotonicity report). None of these four targets is achievable by any budget for this model on this task, and KVCalib correctly refuses to pretend otherwise: it returns the least-aggressive budget in the grid and says so in the guarantee text, rather than quietly certifying a number the model can't back up. (An earlier draft of this table described alpha=0.60 as "calibrating a non-trivial budget" -- checking directly which budget got selected across all 200 splits showed that was wrong: it selects 4096 every time, identically to 0.05/0.10/0.20. Left here as a correction rather than silently fixed, since it's exactly the kind of claim this table exists to get right.) alpha=0.80 is included specifically to also show what happens once the target is above that floor: real, non-trivial budget selection (512, 1024, or 4096 depending on the split), with realized loss tracking the target as it relaxes.
A more capable pilot model would very likely clear the smaller alpha targets too; this table reports what a real 0.5B model on real hardware in this session actually does, which is the entire point of calibrating on your own workload rather than trusting someone else's benchmark number.
Cross-backend comparison
The same repeat-split protocol (real generation, real eviction, 200 repeat-splits, 70/30 calibration/test) extended to SnapKV and StreamingLLM: at alpha in {0.05, 0.10, 0.20}, all three backends fall back to the largest grid budget in every split, identically -- this model's own baseline exact-match floor on this task, not a backend difference. At alpha=0.80, the one target already known to clear that floor, the backends separate:
| rank | backend | median selected budget | mean selected budget | fraction at max budget | mean realized loss |
|---|---|---|---|---|---|
| 1 | SnapKV | 256 | 349 | 1% | 0.743 |
| 2 | H2O | 1024 | 1782 | 28% | 0.722 |
| 3 | StreamingLLM | 4096 | 2606 | 51.5% | 0.684 |
SnapKV achieves the smallest certified budget at alpha=0.80 by a wide margin --
consistent with SnapKV's content-scored eviction outperforming both H2O's cumulative-
attention scoring and StreamingLLM's purely positional sink+recency scheme on this
single-document QA task (see the per-backend monotonicity reports for the underlying
per-budget curves). See
docs/coverage_validation/cross_backend_comparison.md
for the full per-backend repeat-split tables and discussion.
Group-conditional coverage
MondrianRiskControl (Phase 3) calibrates a separate budget per group instead of
one pooled budget -- see
docs/group_conditional_risk_control.md
for what this does and doesn't change (analysis/reporting only: it does not change
what risk_controlled_cache() serves). Validated on the same real H2O +
multifieldqa_en matrix, grouped by context-length bucket ("short"/"long", split
at the median): at alpha=0.80, the pooled budget (1782 tokens) is the wrong number
for either group -- short-context prompts only need 1508 (15% less), long-context
prompts need 2171 (22% more). At alpha=0.05, this dataset shows the real cost:
n_min=20 exceeds the entire 17-member "long" group before any split is even
drawn, a structural floor, not unlucky sampling, so that group correctly falls back
to the pooled guarantee instead of fabricating one. See
docs/coverage_validation/group_conditional_coverage.md
for the full repeat-split table and discussion.
Negative controls
All three are run on real model output and detailed in
docs/negative_controls.md; the first two are mandatory
per Sec. 7.2 of the project brief, the third is Phase 2's drift monitor validated
against the same real distribution shift as the second:
| control | result |
|---|---|
| Broken monotonicity (real per-prompt losses, budget assignment shuffled per prompt) | MonotonicityViolationError raised; calibration refused |
Broken exchangeability (calibrated on multifieldqa_en at alpha=0.80, tested on hotpotqa) |
Calibrated budget 1024 (a genuine ~75% eviction, not a fallback); realized loss 0.867 on hotpotqa vs. target alpha=0.80 -- guarantee measurably broken, as expected |
Drift monitor (monitor.aci.ACIDriftMonitor, same shift replayed live-traffic-style) |
Silent across all 39 in-control rounds; raised GuaranteeInvalidatedError at round 28 of 60 on the shifted replay |
The exchangeability control specifically calibrates at an alpha that selects a
genuinely compressed budget (see docs/negative_controls.md for why alpha=0.10,
used in an earlier draft, was a weaker version of this control: it fell back to the
largest budget in the grid, so there was no real eviction decision in effect for the
distribution shift to actually invalidate).
License
Apache 2.0.
Metadata
Release files for kvcalib 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| kvcalib-0.4.0.tar.gz | 130.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| kvcalib-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 180.3 kB
Release files / kvcalib-0.4.0.tar.gz
| Download URL | kvcalib-0.4.0.tar.gz |
|---|---|
| Size | 130.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
835e45edbca63fc3902c48fba6c906bc222437040707973e105541c084647535
|
|
BLAKE2b-256 checksum How to use checksums |
537da8ea04bdd1478eddf3f9a5434399ceb6e8ef9238ad65c88d22481ef86e25
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.
Transparency logRelease files / kvcalib-0.4.0-py3-none-any.whl
| Download URL | kvcalib-0.4.0-py3-none-any.whl |
|---|---|
| Size | 49.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
449fc9d3ba56ed09ed6d58f6544f7e44a10e19b6f3ca044ff61e076057f4e37a
|
|
BLAKE2b-256 checksum How to use checksums |
fec02833a9df3c17086f7d86228b703c6c40261567d28f840407c982c811eeb5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.
Transparency log