Skip to main content

szl-quant-bench

Honest quantization quality-curve bench for the SZL estate. Real uniform absmax quantization (round-to-nearest, signed ints, 2–16 bits) with quality metrics — cosine similarity, KL divergence over softmaxed rows, top-1 agreement, MSE — on the supplied logits. This measures logit quantization, not weight quantization, GGUF behavior, inference speed, or model quality after a weight conversion. Stdlib only; no downloads, no network.

Doctrine: empty or missing inputs return BLOCKED with a reason. The demo fixture is synthetic and labeled as such. Nothing here claims to measure a real model unless you pass real logits in. Run states: MEASURED | BLOCKED | INVALID | FAILED.

Install and verify

pip install -e . pytest
python -m pytest tests/ -q
python -m szl_quant_bench.harness       # deterministic fixture curve + receipt

Measured fixture curve (synthetic gaussian logits, seed 42 — NOT a real model)

bits cosine KL (mean) top-1 agreement vs fp32
16 1.0000 0.00000 100.0% 2.0x
8 1.0000 0.00003 100.0% 4.0x
4 0.9882 0.01211 68.8% 8.0x
3 0.9420 0.06556 56.2% 10.67x
2 0.5715 0.57019 50.0% 16.0x

These numbers came out of run_curve in this repository, executed before the initial push. Your model's curve will differ — run it on real logits.

Measuring a real model

Capture a logit batch from the exact model revision and fixed evaluation prompt set. Store a JSON object with logits (a rectangular finite numeric matrix) and provenance containing:

  • source_kind: REAL_MODEL_LOGITS
  • model_id, runtime, hardware: nonempty identifiers
  • model_revision: the exact 40-character lowercase commit
  • prompt_set_sha256: the full lowercase SHA-256 of the saved prompt artifact
python -m szl_quant_bench.harness --input model-logits.json --output curve-receipt.json

The output file must not already exist. Missing provenance, malformed matrices, non-finite values and invalid bit widths fail closed with a nonzero exit. With no input the CLI runs only the explicitly SYNTHETIC fixture. The Python API accepts run_curve(logits, chain=chain, provenance=provenance); omitting provenance labels the source UNVERIFIED_CALLER_INPUT. Every MEASURED API result carries a hash-chained receipt record. If no chain is supplied, the API creates a fresh one-record chain for that measurement; pass a caller-owned ReceiptChain to append the record to an existing chain.

Receipts bind the normalized float input matrix hash, complete provenance, all quality records and the explicit measurement boundary. Provenance remains CALLER_DECLARED: a hash cannot establish that the caller actually ran a model. Preserve the raw logits and independent model execution evidence to support that claim. Empty or malformed chains do not verify. KL is evaluated in log space without silently clipping small probabilities.

Compare bit widths only on the same inputs, hardware and revision. A measured logit curve does not authorize a GGUF export decision: actual quantized model outputs and runtime evaluation are separately required.

Apache-2.0 · Doctrine v11 · SZL Holdings

Metadata

Release files for szl-quant-bench 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for szl-quant-bench 0.2.0
File Size Uploaded
szl_quant_bench-0.2.0.tar.gz 13.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for szl-quant-bench 0.2.0
File Interpreter ABI Platform
szl_quant_bench-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 25.7 kB

Release files / szl_quant_bench-0.2.0.tar.gz

Download URL szl_quant_bench-0.2.0.tar.gz
Size 13.7 kB
Tags Source
SHA-256 checksum
How to use checksums
71554d15fbdc86d23fbb859085bc519d94a37aba0dde6548ef15b68302d813c9
BLAKE2b-256 checksum
How to use checksums
f151a95cf6d2afe2d4100f5c9af16cc567f4e4899b97106138f0d0b6277c1689
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / szl_quant_bench-0.2.0-py3-none-any.whl

Download URL szl_quant_bench-0.2.0-py3-none-any.whl
Size 12.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2347bcf9ecdd135a460e5eb76557fe4e211d604eddca55acf477bf7b90703fcb
BLAKE2b-256 checksum
How to use checksums
5f81288918b11a3ea5d8c26aa00bf261a4e1804390683330a727182ecd26acd5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page