Skip to main content

KVCalib

KVCalib turns "how aggressively can I evict my KV cache" into a decision backed by a distribution-free statistical guarantee, calibrated on your own workload, instead of a benchmark number computed on someone else's traffic.

It is a thin calibration/monitoring layer on top of kvpress (NVIDIA's KV-cache eviction backends and benchmark harness) -- not a new eviction heuristic, not a fork of kvpress, and not a competitor to it. See docs/architecture.md for the full design rationale and docs/understanding_the_guarantee.md for what the guarantee does and doesn't promise, including how it differs from H2O's own approximation-ratio guarantee.

Status: Phase 2 complete. Multi-backend: H2O, SnapKV, and StreamingLLM are all validated (monotonicity + repeat-split coverage) on exact-match loss; PyramidKV is implemented but documented as excluded on non-CUDA hardware (see docs/monotonicity_reports/pyramidkv_exact_match.md). Cross-backend comparison: SnapKV achieves the smallest certified budget at alpha=0.80 (see Cross-backend comparison). Metric-degradation loss is implemented but parked, not validated -- its own monotonicity check failed on real data (see docs/monotonicity_reports/h2o_metric_degradation.md). An ACI-based drift monitor (kvcalib.monitor.aci) is implemented and validated against a real distribution shift (see Negative controls). Phase 3 complete: MondrianRiskControl adds group-conditional (Mondrian-style) calibration, analysis/reporting only for now, validated on real data at both a genuine budget tightening and a genuine calibration-data cost (see Group-conditional coverage). Phase 4 complete: kvcalib.telemetry.otel_export adds opt-in OpenTelemetry span/event export (see Install); kvcalib.monitor.streaming.StreamingCalibrator adds windowed recalibration triggered by the existing drift monitor, not a new online calibration algorithm; score-threshold eviction knobs were considered and scoped out (all four backends and kvpress's shared scoring base are rank-based, with no score-cutoff quantity to expose); packaging polish and a PyPI Trusted Publishing release workflow round out the phase. See KVCALIB_PROJECT_SPEC.md for full phase-by-phase detail and CHANGELOG.md for the versioned history.

Install

pip install "kvcalib[kvpress]"

Not yet published to PyPI -- until the first release goes out, install from a checkout instead:

pip install -e ".[kvpress]"   # from a checkout, until the PyPI release is published

Known upstream issue on Python 3.13: kvpress 0.5.4 (latest as of this writing) transitively depends on google-fire, which still imports the pipes stdlib module removed in Python 3.13 -- import kvpress fails with ModuleNotFoundError: No module named 'pipes' on 3.13, regardless of anything in this project. This is a kvpress/ fire compatibility gap, not a KVCalib issue -- kvcalib's own package imports lazily and does not hit it. Until it's fixed upstream, either use Python 3.10-3.12 for the [kvpress] extra, or work around it locally with a one-line shim: echo 'from shlex import quote' > "$(python3 -c "import sysconfig; print(sysconfig.get_paths()['purelib'])")/pipes.py".

Known upstream limitation: PyramidKVBackend cannot run generation on non-CUDA hardware. PyramidKV gives each transformer layer a different KV-cache length by design; transformers' eager/sdpa/flex_attention paths build one causal mask (sized for one assumed cache length) and reuse it for every layer, which crashes whenever a prompt/budget pair produces non-uniform per-layer lengths (common, not an edge case). Only flash_attention_2/_3 are exempt, and both are hard-gated in transformers to CUDA/HIP/Cambricon-MLU hardware -- unavailable on Apple Silicon regardless of attention-kernel package. See docs/monotonicity_reports/pyramidkv_exact_match.md for the full investigation, including two remediation paths (PyTorch 2.13's MPS FlexAttention, the mps-flash-attn package) that were concretely tried and ruled out rather than assumed incompatible. PyramidKVBackend is implemented and its budget-to-ratio translation is correct -- this is a kvpress/transformers limitation on this hardware, not a bug in this project -- but it is excluded from Phase 2's coverage validation and cross-backend comparison as a result.

Calibration and deployment require kvpress + transformers + torch and an actual model -- there is no synthetic substitute for "run the model," since the whole point is measuring eviction's real effect. The core calibration math (kvcalib.core) has no such dependency and can be tested and used standalone (e.g. against precomputed loss curves).

Optional: pip install "kvcalib[otel]" (or .[otel] from a checkout) adds opentelemetry-api for kvcalib.telemetry.otel_export's span/event export (export_calibration_span, export_generation_event) -- this is opt-in and caller-driven, not required for calibration or deployment.

Quickstart

from transformers import AutoModelForCausalLM, AutoTokenizer
from kvcalib import CalibrationPrompt, Calibrator
from kvcalib.backends import H2OBackend

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct", attn_implementation="eager")

calibration_set = [
    CalibrationPrompt(prompt="...", reference_answer="..."),
    # ...prompts representative of your own workload...
]

calibrator = Calibrator(alpha=0.1, backend=H2OBackend())
calibrator.fit(calibration_set, model, budget_grid=[8, 16, 32, 64, 512], tokenizer=tokenizer)
calibrator.save("calibration.json")
print(calibrator.guarantee)

# --- at deploy time ---
calibrator = Calibrator.load("calibration.json")
cache = calibrator.risk_controlled_cache()  # drop-in kvpress-compatible press

from kvpress import KVPressTextGenerationPipeline
pipe = KVPressTextGenerationPipeline(model=model, tokenizer=tokenizer)
print(pipe("Context: ...", question="...", press=cache, max_new_tokens=10))

Running examples/single_backend_quickstart.py end to end (a tiny, 14-prompt calibration set of short factual questions) prints:

Expected exact_match accuracy loss from evicting to budget=8 tokens is <= 0.1,
assuming calibration and live traffic are exchangeable. Calibrated on n=14 prompts
on 2026-07-22 using the h2o backend. This guarantee holds only if calibration and
live traffic are exchangeable [...]
Generated: Madrid

See examples/ for this quickstart, a cross-backend comparison, and a fast (no-GPU) walkthrough of the negative-control refusal path.

Coverage validation

This is the project's actual evidence, not a benchmark score borrowed from someone else's traffic: real model + real kvpress H2O eviction, run on LongBench multifieldqa_en, then repeat-split many times from that one real, measured loss matrix (see tests/coverage_validation/, pytest tests/coverage_validation -m gpu).

Setup: Qwen2.5-0.5B-Instruct, H2O backend, exact-match-with-containment loss, n=39 calibration prompts (all multifieldqa_en records whose full document fits under the 4000-token budget grid ceiling -- see docs/monotonicity_reports/h2o_exact_match.md for why), 70/30 calibration/test split, 200 repeat-splits per alpha.

alpha target result mean realized loss (held-out) n_calib n_test
0.01 refused -- InsufficientCalibrationDataError (n_calib=27 < required 100) -- 27 12
0.05 falls back to largest budget (4096) in all 200 splits 0.556 27 12
0.10 falls back to largest budget (4096) in all 200 splits 0.556 27 12
0.20 falls back to largest budget (4096) in all 200 splits 0.556 27 12
0.60 falls back to largest budget (4096) in all 200 splits 0.556 27 12
0.80 genuine, non-trivial budget selection -- 512, 1024, or 4096 depending on the split 0.722 27 12

These are the real, reported numbers -- not illustrative ones. Two things worth being upfront about, since both are legitimate outcomes this tool is specifically designed to surface rather than hide:

  • alpha=0.01 is refused, not silently miscalibrated. With only 27 calibration prompts in a 70/30 split of this 39-prompt sample, the derived minimum-n floor for alpha=0.01 is 100 (kvcalib.core.validation.minimum_calibration_size); the guard fires and the tool says so, rather than returning a technically-computed but meaningless budget.
  • alpha=0.05/0.10/0.20/0.60 all fall back to the largest budget in the grid, with an identical realized loss, in every one of the 200 repeat-splits. Qwen2.5-0.5B- Instruct's own baseline exact-match accuracy on this task -- even with no eviction at all -- has a loss around 0.56 (see the monotonicity report). None of these four targets is achievable by any budget for this model on this task, and KVCalib correctly refuses to pretend otherwise: it returns the least-aggressive budget in the grid and says so in the guarantee text, rather than quietly certifying a number the model can't back up. (An earlier draft of this table described alpha=0.60 as "calibrating a non-trivial budget" -- checking directly which budget got selected across all 200 splits showed that was wrong: it selects 4096 every time, identically to 0.05/0.10/0.20. Left here as a correction rather than silently fixed, since it's exactly the kind of claim this table exists to get right.) alpha=0.80 is included specifically to also show what happens once the target is above that floor: real, non-trivial budget selection (512, 1024, or 4096 depending on the split), with realized loss tracking the target as it relaxes.

A more capable pilot model would very likely clear the smaller alpha targets too; this table reports what a real 0.5B model on real hardware in this session actually does, which is the entire point of calibrating on your own workload rather than trusting someone else's benchmark number.

Cross-backend comparison

The same repeat-split protocol (real generation, real eviction, 200 repeat-splits, 70/30 calibration/test) extended to SnapKV and StreamingLLM: at alpha in {0.05, 0.10, 0.20}, all three backends fall back to the largest grid budget in every split, identically -- this model's own baseline exact-match floor on this task, not a backend difference. At alpha=0.80, the one target already known to clear that floor, the backends separate:

rank backend median selected budget mean selected budget fraction at max budget mean realized loss
1 SnapKV 256 349 1% 0.743
2 H2O 1024 1782 28% 0.722
3 StreamingLLM 4096 2606 51.5% 0.684

SnapKV achieves the smallest certified budget at alpha=0.80 by a wide margin -- consistent with SnapKV's content-scored eviction outperforming both H2O's cumulative- attention scoring and StreamingLLM's purely positional sink+recency scheme on this single-document QA task (see the per-backend monotonicity reports for the underlying per-budget curves). See docs/coverage_validation/cross_backend_comparison.md for the full per-backend repeat-split tables and discussion.

Group-conditional coverage

MondrianRiskControl (Phase 3) calibrates a separate budget per group instead of one pooled budget -- see docs/group_conditional_risk_control.md for what this does and doesn't change (analysis/reporting only: it does not change what risk_controlled_cache() serves). Validated on the same real H2O + multifieldqa_en matrix, grouped by context-length bucket ("short"/"long", split at the median): at alpha=0.80, the pooled budget (1782 tokens) is the wrong number for either group -- short-context prompts only need 1508 (15% less), long-context prompts need 2171 (22% more). At alpha=0.05, this dataset shows the real cost: n_min=20 exceeds the entire 17-member "long" group before any split is even drawn, a structural floor, not unlucky sampling, so that group correctly falls back to the pooled guarantee instead of fabricating one. See docs/coverage_validation/group_conditional_coverage.md for the full repeat-split table and discussion.

Negative controls

All three are run on real model output and detailed in docs/negative_controls.md; the first two are mandatory per Sec. 7.2 of the project brief, the third is Phase 2's drift monitor validated against the same real distribution shift as the second:

control result
Broken monotonicity (real per-prompt losses, budget assignment shuffled per prompt) MonotonicityViolationError raised; calibration refused
Broken exchangeability (calibrated on multifieldqa_en at alpha=0.80, tested on hotpotqa) Calibrated budget 1024 (a genuine ~75% eviction, not a fallback); realized loss 0.867 on hotpotqa vs. target alpha=0.80 -- guarantee measurably broken, as expected
Drift monitor (monitor.aci.ACIDriftMonitor, same shift replayed live-traffic-style) Silent across all 39 in-control rounds; raised GuaranteeInvalidatedError at round 28 of 60 on the shifted replay

The exchangeability control specifically calibrates at an alpha that selects a genuinely compressed budget (see docs/negative_controls.md for why alpha=0.10, used in an earlier draft, was a weaker version of this control: it fell back to the largest budget in the grid, so there was no real eviction decision in effect for the distribution shift to actually invalidate).

License

Apache 2.0.

Metadata

Release files for kvcalib 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for kvcalib 0.4.0
File Size Uploaded
kvcalib-0.4.0.tar.gz 130.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for kvcalib 0.4.0
File Interpreter ABI Platform
kvcalib-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 180.3 kB

Release files / kvcalib-0.4.0.tar.gz

Download URL kvcalib-0.4.0.tar.gz
Size 130.8 kB
Tags Source
SHA-256 checksum
How to use checksums
835e45edbca63fc3902c48fba6c906bc222437040707973e105541c084647535
BLAKE2b-256 checksum
How to use checksums
537da8ea04bdd1478eddf3f9a5434399ceb6e8ef9238ad65c88d22481ef86e25
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release files / kvcalib-0.4.0-py3-none-any.whl

Download URL kvcalib-0.4.0-py3-none-any.whl
Size 49.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
449fc9d3ba56ed09ed6d58f6544f7e44a10e19b6f3ca044ff61e076057f4e37a
BLAKE2b-256 checksum
How to use checksums
fec02833a9df3c17086f7d86228b703c6c40261567d28f840407c982c811eeb5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 23, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page