DeltaCert
Certify any change to your LLM serving stack — quantization, engine upgrades, batch size, model updates — with a mathematical bound, before you deploy.
Built by Threvo Labs.
pip install deltacert
The catch
Standard benchmarks said this 4-bit quantized model was fine:
| Config | Short eval said | d_COMM | safe_until_token | failure_after_token | Verdict |
|---|---|---|---|---|---|
| Llama-3.1-8B fp16 → nf4 | GSM8K 5-shot exact-match +1.0% (looks safe) | 0.00 | 31 | 14 | unsafe |
GSM8K said +1.0% — looks safe. DeltaCert flagged it unsafe. The generated text actually forks from the fp16 model at token 14, on long-form coding generations, well before the math bound (token 31) would even flag it. Same nf4 config, three independent signals, one real finding: a standard short-form benchmark missed a regression that shows up on longer generations.
Robustness-checked: excluding all 7 references with degenerate repetition, every statistic is identical (cert_trajectory_clean7.json).
Reproduce it yourself:
deltacert generate-cases --model meta-llama/Llama-3.1-8B-Instruct --output cases.jsonl
deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int4 \
--checks trajectory --trajectory-cases cases.jsonl
Every number above traces to a real certificate in validation_results/; every row in the table below reproduces with one script.
The proof
Seven real flagship tests, each a full before/after comparison on a real model, real GPU, real downstream benchmark:
| Change | Business gain | d_COMM | Downstream effect | Verdict |
|---|---|---|---|---|
| Llama-3.1-8B batch=1 → batch=64 | 64 concurrent requests, same GPU | 6.21 | GSM8K -1.0 pt | ✅ Safe |
| vLLM 0.8.5 → vLLM 0.9.0 | take the upgrade same-week, not months later | 16.09 | GSM8K -1.0 pt | ✅ Safe |
| KV cache default → fp8 (vLLM native) | 2x concurrent capacity | 4.83 | GSM8K -1.0 pt | ✅ Safe |
| gpt-4o-mini pinned snapshot → current alias | same-day provider-drift check | 6.65 | canary acc +0.0 pt | ✅ Safe |
| Standard decode → speculative decode (k=5) | claimed ~2x throughput | 15.38 | GSM8K +0.0 pt, measured 0.28x (slower) | ✅ Safe on quality, not on speed |
| Llama-3.1-8B fp16 → nf4 (W4) | +60% VRAM reduction | 0.00 | forks at token 14 on long generations | ❌ Unsafe |
| Llama-3.1-8B fp16 → GPTQ int4 | +75% VRAM reduction | 1.16 | GSM8K -8.0 pts | ❌ Unsafe |
Five safe, two unsafe. A tool that only ever says "safe" isn't measuring anything — the two unsafe rows above are DeltaCert doing its job.
How it works
DeltaCert compares output distributions before and after a change. From the cosine similarity c between two runs, it computes:
Δ = 4c√(1-c²) (commutator magnitude)
d = -log(Δ/2) (algebraic distance)
divergence_bound = 2·exp(-d)
d is an algebraic distance; 2e⁻ᵈ is the certified bound on output divergence — deterministic, minutes to compute, checkable at every token position, no eval harness or labeled data required.
"Certified" throughout this document means: measured against a calibrated threshold with a stated bound — not a guarantee of downstream quality.
- Full derivation, clamp behavior, per-method calibration (with sample sizes disclosed), and the top-k logprobs caveat: see
SPEC.md. - This README asserts. The spec defends. Nothing here is a proof.
Getting started
pip install deltacert
deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int8
That uses DeltaCert's shipped reference calibration (from the 7-test suite above). For a threshold tuned to your own model and workload, run the sweep yourself:
deltacert capture --model your-model --output baseline.npz
deltacert capture --model your-model --quantization int8 --output candidate.npz
deltacert calibrate --baseline baseline.npz --candidates candidate.npz \
--names int8 --downstream-file your_evals.json
deltacert certify always tells you when it's using the shipped calibration instead of your own.
Integrations
- vLLM plugin — official
vllm.general_pluginsentry point, already wired in this package. Opt-in only: complete no-op unlessDELTACERT_ENFORCE=1is set, sopip install deltacertis safe in a shared image; when enabled, a serving engine refuses to start on an uncertified change. - CI/CD gate —
python -m deltacert.integrations.cicd_hook --cert ./cert.jsonexits 1 (blocks the pipeline) if not certified, 0 if certified. Works with GitHub Actions, GitLab CI, Jenkins, or any CI that checks exit codes. - HuggingFace auto-wiring —
from deltacert.integrations.hf_integration import auto_certifypicks the right collectors for you from what's active in your config (quantization, LoRA, prefix cache) instead of callingcertify_system()with raw parameters yourself.
Verifying a certificate
Optional — certificates work unsigned; signing adds tamper-evidence for sharing certs across teams or with auditors.
deltacert keygen --private-key mykey.pem --public-key mykey.pub
deltacert sign --cert cert.json --key-file mykey.pem
deltacert verify --cert cert.json --key-file mykey.pub
verify exits 0 if the signature is valid, 1 if the certificate was modified after signing or signed by a different key. Every certificate also carries a validation_status field (flagship_validated vs implemented_pending_validation) — signed as part of the payload, so a signature can never make an unvalidated collector's result look more trustworthy than it is.
Keep your private key secret — never commit it, never share it. Only the public key is meant to be distributed.
All 13 reference certificates in validation_results/ are signed with Threvo's key (deltacert-public.pem, committed in this repo). Don't take our numbers on faith — check them yourself:
deltacert verify --cert validation_results/weight_quant/cert_nf4.json --key-file deltacert-public.pem
What it certifies
DeltaCert ships all 20 collectors described in the design — the code is real and implemented, and doesn't get deleted just because a given check hasn't been run in a full end-to-end validation yet. Only 7 have real flagship validation results behind them so far (the table above); the rest are working code with the same math, not yet run through that process.
| # | Check | CLI-drivable | Status |
|---|---|---|---|
| 1 | weight_quant |
✅ | ✅ validated |
| 2 | kv_cache_quant |
✅ | ✅ validated |
| 3 | batch_divergence |
✅ (needs vLLM) | ✅ validated |
| 4 | spec_decoding |
✅ (needs vLLM) | ✅ validated |
| 5 | engine_swap |
✅ | ✅ validated |
| 6 | provider_drift |
✅ | ✅ validated |
| 7 | trajectory |
✅ | ✅ validated |
| 8 | activation_quant |
✅ | 🔬 implemented — validation run pending |
| 9 | prefix_cache |
✅ | 🔬 implemented — validation run pending |
| 10 | lora |
✅ | 🔬 implemented — validation run pending |
| 11 | model_swap |
✅ | 🔬 implemented — validation run pending |
| 12 | prompt_swap |
✅ | 🔬 implemented — validation run pending |
| 13 | sparse_attention |
Python API only | 🔬 implemented — validation run pending |
| 14 | moe_token_dropping |
Python API only | 🔬 implemented — validation run pending |
| 15 | neuron_skipping |
Python API only | 🔬 implemented — validation run pending |
| 16 | allreduce_tp |
Python API only | 🔬 implemented — validation run pending |
| 17 | alltoall_ep |
Python API only | 🔬 implemented — validation run pending |
| 18 | pipeline_parallel |
Python API only | 🔬 implemented — validation run pending |
| 19 | kv_transfer |
Python API only | 🔬 implemented — validation run pending |
| 20 | gradient_compress |
Python API only | 🔬 implemented — validation run pending |
"Python API only" means the check needs code you supply (a custom compress_fn, attention mask, etc.) — the CLI can't conjure that for you; see import deltacert as dc; dc.certify_system(...).
Limitations (stated up front, not discovered by you later)
- The 7 validated results above all use one model family (Llama-3.1-8B-Instruct) plus one hosted API (gpt-4o-mini). Cross-model generalization isn't proven yet.
- Per-method calibration (e.g. the bnb/GPTQ thresholds) is an initial calibration from n=5 configs on one model — not a settled constant. Run
deltacert calibrateon your own model/workload rather than trusting the shipped default for anything production-critical. - The provider_drift result above is a same-day proxy (pinned snapshot vs. current alias), not the real weekly-cadence drift measurement, which needs two runs across real time.
d_commis a reliable within-method damage indicator but is not directly comparable across different compression methods — seeSPEC.mdfor the bnb-vs-GPTQ false-negative this caused and how it's handled.
Roadmap
- More models, more downstream tasks — firming up per-method calibration beyond n=5
- Full validation pass on the remaining 13 collectors
- Real weekly-cadence provider_drift run (beyond the same-day proxy)
- Cross-backend certification: extend
captureto TensorRT-LLM / SGLang (comparison logic is already backend-agnostic)
Citation
If you use DeltaCert, cite it as:
@software{deltacert2026,
title = {DeltaCert: Calibrated Divergence Certification for LLM Serving Systems},
author = {Shorya},
year = {2026},
url = {https://pypi.org/project/deltacert/}
}
License
Apache-2.0. See LICENSE.
Contact
- Issues and questions: open a GitHub issue
- Full validation data:
validation_results/in this repo
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file deltacert-1.1.2.tar.gz.
File metadata
- Download URL: deltacert-1.1.2.tar.gz
- Upload date:
- Size: 77.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
62fd6af313ae3c75435b3256dcc057d39ddf15198dba08a68cdda340a36e5a71
|
|
| MD5 |
8f11ba31a0229372f52b1b80b36d10b2
|
|
| BLAKE2b-256 |
f3b56adfb69063a326b8f70906529a1dbaf6cc96eedfa42ce9448a190428c9f0
|
File details
Details for the file deltacert-1.1.2-py3-none-any.whl.
File metadata
- Download URL: deltacert-1.1.2-py3-none-any.whl
- Upload date:
- Size: 68.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.13.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4da584ba03f5f91cfdfd300433f2856011bdae3fa802283229a7a2b338d0e38a
|
|
| MD5 |
42df918fc41458ec4369f9da3c4fbbe2
|
|
| BLAKE2b-256 |
9aca2a86bd43b6a8f5d3db9b6431c5078684b8881db22575b3fe1e28c7db5b02
|