szl-quant-bench
Honest quantization quality-curve bench for the SZL estate. Real uniform absmax quantization (round-to-nearest, signed ints, 2–16 bits) with quality metrics — cosine similarity, KL divergence over softmaxed rows, top-1 agreement, MSE — on the supplied logits. This measures logit quantization, not weight quantization, GGUF behavior, inference speed, or model quality after a weight conversion. Stdlib only; no downloads, no network.
Doctrine: empty or missing inputs return BLOCKED with a reason. The demo
fixture is synthetic and labeled as such. Nothing here claims to measure a
real model unless you pass real logits in. Run states:
MEASURED | BLOCKED | INVALID | FAILED.
Install and verify
pip install -e . pytest
python -m pytest tests/ -q
python -m szl_quant_bench.harness # deterministic fixture curve + receipt
Measured fixture curve (synthetic gaussian logits, seed 42 — NOT a real model)
| bits | cosine | KL (mean) | top-1 agreement | vs fp32 |
|---|---|---|---|---|
| 16 | 1.0000 | 0.00000 | 100.0% | 2.0x |
| 8 | 1.0000 | 0.00003 | 100.0% | 4.0x |
| 4 | 0.9882 | 0.01211 | 68.8% | 8.0x |
| 3 | 0.9420 | 0.06556 | 56.2% | 10.67x |
| 2 | 0.5715 | 0.57019 | 50.0% | 16.0x |
These numbers came out of run_curve in this repository, executed before the
initial push. Your model's curve will differ — run it on real logits.
Measuring a real model
Capture a logit batch from the exact model revision and fixed evaluation
prompt set. Store a JSON object with logits (a rectangular finite numeric
matrix) and provenance containing:
source_kind:REAL_MODEL_LOGITSmodel_id,runtime,hardware: nonempty identifiersmodel_revision: the exact 40-character lowercase commitprompt_set_sha256: the full lowercase SHA-256 of the saved prompt artifact
python -m szl_quant_bench.harness --input model-logits.json --output curve-receipt.json
The output file must not already exist. Missing provenance, malformed matrices,
non-finite values and invalid bit widths fail closed with a nonzero exit.
With no input the CLI runs only the explicitly SYNTHETIC fixture.
The Python API accepts run_curve(logits, chain=chain, provenance=provenance);
omitting provenance labels the source UNVERIFIED_CALLER_INPUT. Every
MEASURED API result carries a hash-chained receipt record. If no chain is
supplied, the API creates a fresh one-record chain for that measurement; pass
a caller-owned ReceiptChain to append the record to an existing chain.
Receipts bind the normalized float input matrix hash, complete provenance,
all quality records and the explicit measurement boundary. Provenance remains
CALLER_DECLARED: a hash cannot establish that the caller actually ran a model.
Preserve the raw logits and independent model execution evidence to support
that claim. Empty or malformed chains do not verify. KL is evaluated in log
space without silently clipping small probabilities.
Compare bit widths only on the same inputs, hardware and revision. A measured logit curve does not authorize a GGUF export decision: actual quantized model outputs and runtime evaluation are separately required.
Apache-2.0 · Doctrine v11 · SZL Holdings
Metadata
Release files for szl-quant-bench 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| szl_quant_bench-0.2.1.tar.gz | 14.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| szl_quant_bench-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 26.4 kB
Release files / szl_quant_bench-0.2.1.tar.gz
| Download URL | szl_quant_bench-0.2.1.tar.gz |
|---|---|
| Size | 14.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d2b51beb66792e955aed648e23f477d987a7b581c4ddf99abc622ea1ae630cfd
|
|
BLAKE2b-256 checksum How to use checksums |
efb06b2b0671b96e13d9a4bc7a3c0f868dc623cdf13071be6a583379cb67e5ba
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / szl_quant_bench-0.2.1-py3-none-any.whl
| Download URL | szl_quant_bench-0.2.1-py3-none-any.whl |
|---|---|
| Size | 12.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ce326a302f0d74dd575b8b6cd0c2c7a404999fd5af038c9115ee756a32153d98
|
|
BLAKE2b-256 checksum How to use checksums |
55968f6ee1ff55c9e3da142e19f70c59add143cc46be8ee2ce651db3fec4119f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log