Skip to main content

DeltaCert

PyPI License Python

Certify any change to your LLM serving stack — quantization, engine upgrades, batch size, model updates — with a mathematical bound, before you deploy. Unlock the cost savings (2x KV-cache batch capacity, 75% VRAM reduction) that fear of silent regressions currently blocks, and catch the two failure classes standard benchmarks and single-pass checks both miss: benchmark-blind long-generation forking, and feedback-driven collapse that passes every logit-level check and only shows up in the model's own output.

Built by Threvo Labs.

PyPI · SPEC.md · LICENSE

pip install deltacert

The catch — two ways a change can lie to you, both caught

Way 1: a benchmark says "fine" but generations fork anyway.

Config Short eval said d_COMM safe_until_token failure_after_token Verdict
Llama-3.1-8B fp16 → nf4 GSM8K 5-shot exact-match +1.0% (looks safe) 0.00 31 14 unsafe
Qwen2.5-7B fp16 → nf4 GSM8K exact-match +0.0% (looks safe) 0.00 17 18 unsafe

Same pattern, two model families, two labs, two tokenizers. GSM8K missed both. DeltaCert's trajectory certification caught both — generations fork from the fp16 reference within ~15-18 tokens on long-form coding tasks, well before a short-form benchmark would ever see it.

Way 2 — the one that made us rewrite our own default mode: a config passes every teacher-forced check and is still destroyed.

Config Single-position Trajectory (30,643 positions) Downstream reality Verdict
Qwen2.5-7B fp8 KV-cache safe (d=3.72, cosines ≥0.9998) safe (d≥3.52 at every position) GSM8K 0.88 → 0.00, 0/100 correct unsafe
Llama-3.1-8B fp8 KV-cache, identical flag safe safe GSM8K within noise genuinely safe

The identical engine flag is benign on Llama and catastrophic on Qwen. Both teacher-forced modes — single-position and full trajectory — certify the Qwen collapse safe, because the damage doesn't live in the logits; it lives in the autoregressive feedback loop, which teacher forcing structurally can't see. Catching this needed a third instrument: a free-running collector that runs the deployed decode policy on both engines and measures the actual output process, with a McNemar-exact-test guard so a benign fork's ordinary spontaneous-repetition rate can't be mistaken for caused collapse. It fires decisively on Qwen (79% excess degeneration, p≈10⁻¹⁰) and stays quiet on Llama (2.3%, p=0.63) — sensitivity on the real failure, specificity on the real clean case, same instrument, same thresholds.

Robustness-checked: excluding all 7 references with degenerate repetition, every trajectory statistic is identical (cert_trajectory_clean7.json); the fp8-KV single-position measurement was independently reproduced on a different host/stack to three decimal places.

Reproduce it yourself:

deltacert generate-cases --model meta-llama/Llama-3.1-8B-Instruct --output cases.jsonl
deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int4 \
    --checks trajectory --trajectory-cases cases.jsonl

Every number above traces to a real certificate in validation_results/; every row in the table below reproduces with one script.

The proof

Seven real flagship tests, each a full before/after comparison on a real model, real GPU, real downstream benchmark:

Change Business gain d_COMM Downstream effect Verdict
Llama-3.1-8B batch=1 → batch=64 64 concurrent requests, same GPU 6.21 GSM8K -1.0 pt ✅ Safe
vLLM 0.8.5 → vLLM 0.9.0 take the upgrade same-week, not months later 16.09 GSM8K -1.0 pt ✅ Safe
KV cache default → fp8 (vLLM native) 2x concurrent capacity 4.83 GSM8K -1.0 pt ✅ Safe
gpt-4o-mini pinned snapshot → current alias same-day provider-drift check 6.65 canary acc +0.0 pt ✅ Safe
Standard decode → speculative decode (k=5) claimed ~2x throughput 15.38 GSM8K +0.0 pt, measured 0.28x (slower) ✅ Safe on quality, not on speed
Llama-3.1-8B fp16 → nf4 (W4) +60% VRAM reduction 0.00 forks at token 14 on long generations ❌ Unsafe
Llama-3.1-8B fp16 → GPTQ int4 +75% VRAM reduction 1.16 GSM8K -8.0 pts ❌ Unsafe

Five safe, two unsafe. A tool that only ever says "safe" isn't measuring anything — the two unsafe rows above are DeltaCert doing its job.

How it works

DeltaCert compares output distributions before and after a change. d_COMM ("commutator distance," from the operator-algebraic commutator bound it's derived from) is computed from the cosine similarity c between two runs:

Δ = 4c√(1-c²)        (commutator magnitude)
d = -log(Δ/2)        (algebraic distance)
divergence_bound = 2·exp(-d)

d is an algebraic distance; 2e⁻ᵈ is the certified bound on output divergence — deterministic, minutes to compute, checkable at every token position, no eval harness or labeled data required.

"Certified" throughout this document means: measured against a calibrated threshold with a stated bound — not a guarantee of downstream quality.

  • Full derivation, clamp behavior, per-method calibration (with sample sizes disclosed), and the top-k logprobs caveat: see SPEC.md.
  • This README asserts. The spec defends. Nothing here is a proof.

Getting started

pip install deltacert

deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int8

That uses DeltaCert's shipped reference calibration (from the 7-test suite above). For a threshold tuned to your own model and workload, run the sweep yourself:

deltacert capture --model your-model --output baseline.npz
deltacert capture --model your-model --quantization int8 --output candidate.npz
deltacert calibrate --baseline baseline.npz --candidates candidate.npz \
    --names int8 --downstream-file your_evals.json

deltacert certify always tells you when it's using the shipped calibration instead of your own.

For KV-cache and any change whose damage can be feedback-driven, add a free-running check (separate subcommand — its certificate is McNemar/degeneration-based, not d_COMM-based):

deltacert free-running --model your-model --kv-cache-dtype fp8 --output cert_free_running.json

Integrations

  • vLLM plugin — official vllm.general_plugins entry point, already wired in this package. Opt-in only: complete no-op unless DELTACERT_ENFORCE=1 is set, so pip install deltacert is safe in a shared image; when enabled, a serving engine refuses to start on an uncertified change.
  • CI/CD gatepython -m deltacert.integrations.cicd_hook --cert ./cert.json exits 1 (blocks the pipeline) if not certified, 0 if certified. Works with GitHub Actions, GitLab CI, Jenkins, or any CI that checks exit codes.
  • HuggingFace auto-wiringfrom deltacert.integrations.hf_integration import auto_certify picks the right collectors for you from what's active in your config (quantization, LoRA, prefix cache) instead of calling certify_system() with raw parameters yourself.

Verifying a certificate

Optional — certificates work unsigned; signing adds tamper-evidence for sharing certs across teams or with auditors.

deltacert keygen --private-key mykey.pem --public-key mykey.pub
deltacert sign --cert cert.json --key-file mykey.pem
deltacert verify --cert cert.json --key-file mykey.pub

verify exits 0 if the signature is valid, 1 if the certificate was modified after signing or signed by a different key. Every certificate also carries a validation_status field (flagship_validated vs implemented_pending_validation) — signed as part of the payload, so a signature can never make an unvalidated collector's result look more trustworthy than it is.

Keep your private key secret — never commit it, never share it. Only the public key is meant to be distributed.

All 13 reference certificates in validation_results/ are signed with Threvo's key (deltacert-public.pem, committed in this repo). Don't take our numbers on faith — check them yourself:

deltacert verify --cert validation_results/weight_quant/cert_nf4.json --key-file deltacert-public.pem

What it certifies

DeltaCert ships all 21 collectors described in the design — the code is real and implemented, and doesn't get deleted just because a given check hasn't been run in a full end-to-end validation yet. 8 have real flagship validation results behind them so far (the tables above, including the free-running collector that catches feedback-driven failures teacher-forced checks miss); the rest are working code with the same math, not yet run through that process.

# Check CLI-drivable Status
1 weight_quant ✅ validated
2 kv_cache_quant ✅ validated
3 batch_divergence ✅ (needs vLLM) ✅ validated
4 spec_decoding ✅ (needs vLLM) ✅ validated
5 engine_swap ✅ validated
6 provider_drift ✅ validated
7 trajectory ✅ validated
8 free_running ✅ (needs vLLM) ✅ validated
9 activation_quant 🔬 implemented — validation run pending
10 prefix_cache 🔬 implemented — validation run pending
11 lora 🔬 implemented — validation run pending
12 model_swap 🔬 implemented — validation run pending
13 prompt_swap 🔬 implemented — validation run pending
14 sparse_attention Python API only 🔬 implemented — validation run pending
15 moe_token_dropping Python API only 🔬 implemented — validation run pending
16 neuron_skipping Python API only 🔬 implemented — validation run pending
17 allreduce_tp Python API only 🔬 implemented — validation run pending
18 alltoall_ep Python API only 🔬 implemented — validation run pending
19 pipeline_parallel Python API only 🔬 implemented — validation run pending
20 kv_transfer Python API only 🔬 implemented — validation run pending
21 gradient_compress Python API only 🔬 implemented — validation run pending

"Python API only" means the check needs code you supply (a custom compress_fn, attention mask, etc.) — the CLI can't conjure that for you; see import deltacert as dc; dc.certify_system(...).

Limitations (stated up front, not discovered by you later)

  • Validated on two open model families (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct) plus one hosted API (gpt-4o-mini); serving-time flagships beyond weight/KV-cache quantization are validated on Llama only.
  • Per-method calibration (e.g. the bnb/GPTQ thresholds) derives from n=6 configs per model — not a settled constant. Run deltacert calibrate on your own model/workload rather than trusting the shipped default for anything production-critical.
  • The provider_drift result above is a same-day proxy (pinned snapshot vs. current alias), not the real weekly-cadence drift measurement, which needs two runs across real time.
  • d_comm is a reliable within-method damage indicator but is not directly comparable across different compression methods — see SPEC.md for the bnb-vs-GPTQ false-negative this caused and how it's handled.
  • The feedback-driven failure class (fp8 KV-cache on Qwen) currently has one confirmed member after a pre-registered four-candidate hunt; whether other configurations populate it is still open.

Roadmap

  • More models, more downstream tasks — firming up per-method calibration beyond n=5
  • Full validation pass on the remaining 13 collectors
  • Real weekly-cadence provider_drift run (beyond the same-day proxy)
  • Cross-backend certification: extend capture to TensorRT-LLM / SGLang (comparison logic is already backend-agnostic)

Citation

If you use DeltaCert, cite it as:

@software{deltacert2026,
  title  = {DeltaCert: Calibrated Divergence Certification for LLM Serving Systems},
  author = {Shorya},
  year   = {2026},
  url    = {https://pypi.org/project/deltacert/}
}

License

Apache-2.0. See LICENSE.

Contact

  • Issues and questions: open a GitHub issue
  • Full validation data: validation_results/ in this repo

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

deltacert-1.2.0.tar.gz (93.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

deltacert-1.2.0-py3-none-any.whl (79.9 kB view details)

Uploaded Python 3

File details

Details for the file deltacert-1.2.0.tar.gz.

File metadata

  • Download URL: deltacert-1.2.0.tar.gz
  • Upload date:
  • Size: 93.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.4

File hashes

Hashes for deltacert-1.2.0.tar.gz
Algorithm Hash digest
SHA256 7ed7650312ad00339dfe78395aeaca7be4715e2e8c0ccbff98a5e4f0cfa31023
MD5 d3054f7b2db72ebb24aca9ea5b562014
BLAKE2b-256 6f7c1ad04dc4b2ec669870e6501fe046d3bb9a01fcd2941b1d9653ca8177cfcb

See more details on using hashes here.

File details

Details for the file deltacert-1.2.0-py3-none-any.whl.

File metadata

  • Download URL: deltacert-1.2.0-py3-none-any.whl
  • Upload date:
  • Size: 79.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.4

File hashes

Hashes for deltacert-1.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 9e5401c22a2833520d800fafad32da8a634bb870d618e364dc5641f9098986c6
MD5 7d11df72dbcb9c79be4a4bccfddd9cb5
BLAKE2b-256 5fc902ecf7134196064771e088139f236c50f0e981622983e65aff50e92f0b72

See more details on using hashes here.

Release history Release notifications | RSS feed

1.2.4

2 files

1.2.3

2 files

1.2.2

2 files

1.2.1

2 files

This release

1.2.0 This release

2 files

1.1.2

2 files

1.1.1

2 files

1.1.0

2 files

1.0.5

1 file

1.0.4

1 file

1.0.3

1 file

1.0.2

1 file

1.0.1

1 file

1.0.0

1 file

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page