DeltaCert
Certify any change to your LLM serving stack — quantization, engine upgrades, batch size, model updates — with a mathematical bound, before you deploy. Unlock the cost savings (2x KV-cache batch capacity, 75% VRAM reduction) that fear of silent regressions currently blocks, and catch the two failure classes standard benchmarks and single-pass checks both miss: benchmark-blind long-generation forking, and feedback-driven collapse that passes every logit-level check and only shows up in the model's own output.
Built by Threvo Labs.
pip install deltacert
The catch — two ways a change can lie to you, both caught
Way 1: a benchmark says "fine" but generations fork anyway.
| Config | Short eval said | d_COMM | safe_until_token | failure_after_token | Verdict |
|---|---|---|---|---|---|
| Llama-3.1-8B fp16 → nf4 | GSM8K 5-shot exact-match +1.0% (looks safe) | 0.00 | 31 | 14 | unsafe |
| Qwen2.5-7B fp16 → nf4 | GSM8K exact-match +0.0% (looks safe) | 0.00 | 17 | 18 | unsafe |
Same pattern, two model families, two labs, two tokenizers. GSM8K missed both. DeltaCert's trajectory certification caught both — generations fork from the fp16 reference within ~15-18 tokens on long-form coding tasks, well before a short-form benchmark would ever see it.
Way 2 — the one that made us rewrite our own default mode: a config passes every teacher-forced check and is still destroyed.
| Config | Single-position | Trajectory (30,643 positions) | Downstream reality | Verdict |
|---|---|---|---|---|
| Qwen2.5-7B fp8 KV-cache | safe (d=3.72, cosines ≥0.9998) | safe (d≥3.52 at every position) | GSM8K 0.88 → 0.00, 0/100 correct | unsafe |
| Llama-3.1-8B fp8 KV-cache, identical flag | safe | safe | GSM8K within noise | genuinely safe |
The identical engine flag is benign on Llama and catastrophic on Qwen. Both teacher-forced modes — single-position and full trajectory — certify the Qwen collapse safe, because the damage doesn't live in the logits; it lives in the autoregressive feedback loop, which teacher forcing structurally can't see. Catching this needed a third instrument: a free-running collector that runs the deployed decode policy on both engines and measures the actual output process, with a McNemar-exact-test guard so a benign fork's ordinary spontaneous-repetition rate can't be mistaken for caused collapse. It fires decisively on Qwen (79% excess degeneration, p≈10⁻¹⁰) and stays quiet on Llama (2.3%, p=0.63) — sensitivity on the real failure, specificity on the real clean case, same instrument, same thresholds.
Robustness-checked: excluding all 7 references with degenerate repetition, every trajectory statistic is identical (cert_trajectory_clean7.json); the fp8-KV single-position measurement was independently reproduced on a different host/stack to three decimal places.
Reproduce it yourself:
deltacert generate-cases --model meta-llama/Llama-3.1-8B-Instruct --output cases.jsonl
deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int4 \
--checks trajectory --trajectory-cases cases.jsonl
Every number above traces to a real certificate in validation_results/; every row in the table below reproduces with one script.
The proof
Seven real flagship tests, each a full before/after comparison on a real model, real GPU, real downstream benchmark:
| Change | Business gain | d_COMM | Downstream effect | Verdict |
|---|---|---|---|---|
| Llama-3.1-8B batch=1 → batch=64 | 64 concurrent requests, same GPU | 6.21 | GSM8K -1.0 pt | ✅ Safe |
| vLLM 0.8.5 → vLLM 0.9.0 | take the upgrade same-week, not months later | 16.09 | GSM8K -1.0 pt | ✅ Safe |
| KV cache default → fp8 (vLLM native) | 2x concurrent capacity | 4.83 | GSM8K -1.0 pt | ✅ Safe |
| gpt-4o-mini pinned snapshot → current alias | same-day provider-drift check | 6.65 | canary acc +0.0 pt | ✅ Safe |
| Standard decode → speculative decode (k=5) | claimed ~2x throughput | 15.38 | GSM8K +0.0 pt, measured 0.28x (slower) | ✅ Safe on quality, not on speed |
| Llama-3.1-8B fp16 → nf4 (W4) | +60% VRAM reduction | 0.00 | forks at token 14 on long generations | ❌ Unsafe |
| Llama-3.1-8B fp16 → GPTQ int4 | +75% VRAM reduction | 1.16 | GSM8K -8.0 pts | ❌ Unsafe |
Five safe, two unsafe. A tool that only ever says "safe" isn't measuring anything — the two unsafe rows above are DeltaCert doing its job.
How it works
DeltaCert compares output distributions before and after a change. d_COMM ("commutator distance," from the operator-algebraic commutator bound it's derived from) is computed from the cosine similarity c between two runs:
Δ = 4c√(1-c²) (commutator magnitude)
d = -log(Δ/2) (algebraic distance)
divergence_bound = 2·exp(-d)
d is an algebraic distance; 2e⁻ᵈ is the certified bound on output divergence — deterministic, minutes to compute, checkable at every token position, no eval harness or labeled data required.
"Certified" throughout this document means: measured against a calibrated threshold with a stated bound — not a guarantee of downstream quality.
- Full derivation, clamp behavior, per-method calibration (with sample sizes disclosed), and the top-k logprobs caveat: see
SPEC.md. - This README asserts. The spec defends. Nothing here is a proof.
Getting started
pip install deltacert
deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int8
That uses DeltaCert's shipped reference calibration (from the 7-test suite above). For a threshold tuned to your own model and workload, run the sweep yourself:
deltacert capture --model your-model --output baseline.npz
deltacert capture --model your-model --quantization int8 --output candidate.npz
deltacert calibrate --baseline baseline.npz --candidates candidate.npz \
--names int8 --downstream-file your_evals.json
deltacert certify always tells you when it's using the shipped calibration instead of your own.
For KV-cache and any change whose damage can be feedback-driven, add a free-running check (separate subcommand — its certificate is McNemar/degeneration-based, not d_COMM-based):
deltacert free-running --model your-model --kv-cache-dtype fp8 --output cert_free_running.json
Integrations
- vLLM plugin — official
vllm.general_pluginsentry point, already wired in this package. Opt-in only: complete no-op unlessDELTACERT_ENFORCE=1is set, sopip install deltacertis safe in a shared image; when enabled, a serving engine refuses to start on an uncertified change. - CI/CD gate —
python -m deltacert.integrations.cicd_hook --cert ./cert.jsonexits 1 (blocks the pipeline) if not certified, 0 if certified. Works with GitHub Actions, GitLab CI, Jenkins, or any CI that checks exit codes. - HuggingFace auto-wiring —
from deltacert.integrations.hf_integration import auto_certifypicks the right collectors for you from what's active in your config (quantization, LoRA, prefix cache) instead of callingcertify_system()with raw parameters yourself.
Verifying a certificate
Optional — certificates work unsigned; signing adds tamper-evidence for sharing certs across teams or with auditors.
deltacert keygen --private-key mykey.pem --public-key mykey.pub
deltacert sign --cert cert.json --key-file mykey.pem
deltacert verify --cert cert.json --key-file mykey.pub
verify exits 0 if the signature is valid, 1 if the certificate was modified after signing or signed by a different key. Every certificate also carries a validation_status field (flagship_validated vs implemented_pending_validation) — signed as part of the payload, so a signature can never make an unvalidated collector's result look more trustworthy than it is.
Keep your private key secret — never commit it, never share it. Only the public key is meant to be distributed.
All 13 reference certificates in validation_results/ are signed with Threvo's key (deltacert-public.pem, committed in this repo). Don't take our numbers on faith — check them yourself:
deltacert verify --cert validation_results/weight_quant/cert_nf4.json --key-file deltacert-public.pem
What it certifies
DeltaCert ships all 21 collectors described in the design — the code is real and implemented, and doesn't get deleted just because a given check hasn't been run in a full end-to-end validation yet. 8 have real flagship validation results behind them so far (the tables above, including the free-running collector that catches feedback-driven failures teacher-forced checks miss); the rest are working code with the same math, not yet run through that process.
| # | Check | CLI-drivable | Status |
|---|---|---|---|
| 1 | weight_quant |
✅ | ✅ validated |
| 2 | kv_cache_quant |
✅ | ✅ validated |
| 3 | batch_divergence |
✅ (needs vLLM) | ✅ validated |
| 4 | spec_decoding |
✅ (needs vLLM) | ✅ validated |
| 5 | engine_swap |
✅ | ✅ validated |
| 6 | provider_drift |
✅ | ✅ validated |
| 7 | trajectory |
✅ | ✅ validated |
| 8 | free_running |
✅ (needs vLLM) | ✅ validated |
| 9 | activation_quant |
✅ | 🔬 implemented — validation run pending |
| 10 | prefix_cache |
✅ | 🔬 implemented — validation run pending |
| 11 | lora |
✅ | 🔬 implemented — validation run pending |
| 12 | model_swap |
✅ | 🔬 implemented — validation run pending |
| 13 | prompt_swap |
✅ | 🔬 implemented — validation run pending |
| 14 | sparse_attention |
Python API only | 🔬 implemented — validation run pending |
| 15 | moe_token_dropping |
Python API only | 🔬 implemented — validation run pending |
| 16 | neuron_skipping |
Python API only | 🔬 implemented — validation run pending |
| 17 | allreduce_tp |
Python API only | 🔬 implemented — validation run pending |
| 18 | alltoall_ep |
Python API only | 🔬 implemented — validation run pending |
| 19 | pipeline_parallel |
Python API only | 🔬 implemented — validation run pending |
| 20 | kv_transfer |
Python API only | 🔬 implemented — validation run pending |
| 21 | gradient_compress |
Python API only | 🔬 implemented — validation run pending |
"Python API only" means the check needs code you supply (a custom compress_fn, attention mask, etc.) — the CLI can't conjure that for you; see import deltacert as dc; dc.certify_system(...).
Limitations (stated up front, not discovered by you later)
- Validated on two open model families (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct) plus one hosted API (gpt-4o-mini); serving-time flagships beyond weight/KV-cache quantization are validated on Llama only.
- Per-method calibration (e.g. the bnb/GPTQ thresholds) derives from n=6 configs per model — not a settled constant. Run
deltacert calibrateon your own model/workload rather than trusting the shipped default for anything production-critical. - The provider_drift result above is a same-day proxy (pinned snapshot vs. current alias), not the real weekly-cadence drift measurement, which needs two runs across real time.
d_commis a reliable within-method damage indicator but is not directly comparable across different compression methods — seeSPEC.mdfor the bnb-vs-GPTQ false-negative this caused and how it's handled.- The feedback-driven failure class (fp8 KV-cache on Qwen) currently has one confirmed member after a pre-registered four-candidate hunt; whether other configurations populate it is still open.
Roadmap
- More models, more downstream tasks — firming up per-method calibration beyond n=5
- Full validation pass on the remaining 13 collectors
- Real weekly-cadence provider_drift run (beyond the same-day proxy)
- Cross-backend certification: extend
captureto TensorRT-LLM / SGLang (comparison logic is already backend-agnostic)
Citation
If you use DeltaCert, cite it as:
@software{deltacert2026,
title = {DeltaCert: Calibrated Divergence Certification for LLM Serving Systems},
author = {Shorya},
year = {2026},
url = {https://pypi.org/project/deltacert/}
}
License
Apache-2.0. See LICENSE.
Contact
- Issues and questions: open a GitHub issue
- Full validation data:
validation_results/in this repo
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file deltacert-1.2.0.tar.gz.
File metadata
- Download URL: deltacert-1.2.0.tar.gz
- Upload date:
- Size: 93.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7ed7650312ad00339dfe78395aeaca7be4715e2e8c0ccbff98a5e4f0cfa31023
|
|
| MD5 |
d3054f7b2db72ebb24aca9ea5b562014
|
|
| BLAKE2b-256 |
6f7c1ad04dc4b2ec669870e6501fe046d3bb9a01fcd2941b1d9653ca8177cfcb
|
File details
Details for the file deltacert-1.2.0-py3-none-any.whl.
File metadata
- Download URL: deltacert-1.2.0-py3-none-any.whl
- Upload date:
- Size: 79.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9e5401c22a2833520d800fafad32da8a634bb870d618e364dc5641f9098986c6
|
|
| MD5 |
7d11df72dbcb9c79be4a4bccfddd9cb5
|
|
| BLAKE2b-256 |
5fc902ecf7134196064771e088139f236c50f0e981622983e65aff50e92f0b72
|