Skip to main content

quant-regress

Fail CI when quantization costs you task accuracy.

Token-level proxies — perplexity, KL divergence — routinely report that a quantized model is fine. Task accuracy does not, provided you actually measure it. This action measures a task eval at each precision and fails the build when accuracy drops past a threshold you set.

Why

On a cross-encoder claim verifier, dynamic INT8 quantization left aggregate accuracy looking survivable while destroying the capability that mattered:

precision accuracy correct
fp32 (baseline) 100.0% 20/20
int8 0.0% 0/20

Regenerate with python scripts/regen_int8_collapse.py — every figure above comes from that script, and tests/test_claims_are_reproducible.py fails the build if a number in this file stops matching it.

The pattern that matters: the quantized model still runs, still returns a valid answer, and still looks like a working model. It is simply wrong every time. That is why a token-level proxy can clear it and a task eval cannot.

Usage

As a GitHub Action:

- uses: DawnofGenX/quant-regress@v1
  with:
    eval: evals/tasks.jsonl
    model: meta-llama/Llama-3.2-1B-Instruct
    max-drop-points: 2

Or from the command line / your own CI:

pip install quant-regress

quant-regress --eval evals/tasks.jsonl \
              --model meta-llama/Llama-3.2-1B-Instruct \
              --max-drop-points 2 \
              --report quant-regress-report.json

Exit codes: 0 within tolerance, 1 accuracy regressed, 2 bad configuration.

Constraints worth knowing

  • CPU only. quantize_dynamic emits CPU-only quantized kernels; moving the model to CUDA raises at forward time. GitHub-hosted runners have no GPU, so this is not a limitation in CI — but it does mean GPU-only quantizers (AWQ/GPTQ/GGUF) are out of scope for v1.
  • The threshold unit is accuracy points, not a ratio. --max-drop-points: 2 means "fail if any candidate is more than 2 points below the baseline".
  • Comparison is exact-after-normalisation by default. Pass a scorer if you need semantic equivalence.
  • Models are downloaded at run time. A 1B model is ~2 GB; check runner disk.

Honest limits

  • Only fp32 and dynamic int8 are supported in v1.
  • The library runs generation, so it needs a task whose success is checkable in code. It is not an LLM-judge harness.
  • The bundled evals/fixtures/collapse.jsonl is model-specific. Its cases expect the answer yes, so it only demonstrates a collapse for a model that answers yes when unquantized. Pointed at an untuned tiny LM it scores 0/N at both precisions, produces no drop, and the gate correctly reports PASS. Use it to see the output format; bring your own eval for a real check.
  • What the repo's own CI pins is the exit-code contract (0 holds, 1 regressed, 2 misconfigured), not any particular model's accuracy — see tests/test_selftest_gate.py.
  • No published adoption yet. If you use it, an issue saying what you quantized and what broke is genuinely useful.

Metadata

Release files for quant-regress 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for quant-regress 0.1.1
File Size Uploaded
quant_regress-0.1.1.tar.gz 15.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for quant-regress 0.1.1
File Interpreter ABI Platform
quant_regress-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 25.8 kB

Release files / quant_regress-0.1.1.tar.gz

Download URL quant_regress-0.1.1.tar.gz
Size 15.2 kB
Tags Source
SHA-256 checksum
How to use checksums
30c22cb7b11b5cb242cab209be32665e6d61cb184993b0643ce7c38e2df6b8a8
BLAKE2b-256 checksum
How to use checksums
785f3b40e4be8c4309c854f6ff5eddc0c5f09c216603613ee081ac8f36dc13a9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / quant_regress-0.1.1-py3-none-any.whl

Download URL quant_regress-0.1.1-py3-none-any.whl
Size 10.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7dfde2b9ed68a7c0f0e2c939f6f587250365f2cdd6d1990779a695e496e11adf
BLAKE2b-256 checksum
How to use checksums
039de674bc77f7ed599d583563d1b6e08c19e839b6f8b471abdca36a45572de0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page