quant-regress
Fail CI when quantization costs you task accuracy.
Token-level proxies — perplexity, KL divergence — routinely report that a quantized model is fine. Task accuracy does not, provided you actually measure it. This action measures a task eval at each precision and fails the build when accuracy drops past a threshold you set.
Why
On a cross-encoder claim verifier, dynamic INT8 quantization left aggregate accuracy looking survivable while destroying the capability that mattered:
| precision | accuracy | correct |
|---|---|---|
| fp32 (baseline) | 100.0% | 20/20 |
| int8 | 0.0% | 0/20 |
Regenerate with python scripts/regen_int8_collapse.py — every figure above
comes from that script, and tests/test_claims_are_reproducible.py fails the
build if a number in this file stops matching it.
The pattern that matters: the quantized model still runs, still returns a valid answer, and still looks like a working model. It is simply wrong every time. That is why a token-level proxy can clear it and a task eval cannot.
Usage
As a GitHub Action:
- uses: DawnofGenX/quant-regress@v1
with:
eval: evals/tasks.jsonl
model: meta-llama/Llama-3.2-1B-Instruct
max-drop-points: 2
Or from the command line / your own CI:
pip install quant-regress
quant-regress --eval evals/tasks.jsonl \
--model meta-llama/Llama-3.2-1B-Instruct \
--max-drop-points 2 \
--report quant-regress-report.json
Exit codes: 0 within tolerance, 1 accuracy regressed, 2 bad configuration.
Constraints worth knowing
- CPU only.
quantize_dynamicemits CPU-only quantized kernels; moving the model to CUDA raises at forward time. GitHub-hosted runners have no GPU, so this is not a limitation in CI — but it does mean GPU-only quantizers (AWQ/GPTQ/GGUF) are out of scope for v1. - The threshold unit is accuracy points, not a ratio.
--max-drop-points: 2means "fail if any candidate is more than 2 points below the baseline". - Comparison is exact-after-normalisation by default. Pass a scorer if you need semantic equivalence.
- Models are downloaded at run time. A 1B model is ~2 GB; check runner disk.
Honest limits
- Only
fp32and dynamicint8are supported in v1. - The library runs generation, so it needs a task whose success is checkable in code. It is not an LLM-judge harness.
- The bundled
evals/fixtures/collapse.jsonlis model-specific. Its cases expect the answeryes, so it only demonstrates a collapse for a model that answersyeswhen unquantized. Pointed at an untuned tiny LM it scores 0/N at both precisions, produces no drop, and the gate correctly reports PASS. Use it to see the output format; bring your own eval for a real check. - What the repo's own CI pins is the exit-code contract (0 holds, 1 regressed,
2 misconfigured), not any particular model's accuracy — see
tests/test_selftest_gate.py. - No published adoption yet. If you use it, an issue saying what you quantized and what broke is genuinely useful.
Metadata
Release files for quant-regress 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| quant_regress-0.1.1.tar.gz | 15.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| quant_regress-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.8 kB
Release files / quant_regress-0.1.1.tar.gz
| Download URL | quant_regress-0.1.1.tar.gz |
|---|---|
| Size | 15.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
30c22cb7b11b5cb242cab209be32665e6d61cb184993b0643ce7c38e2df6b8a8
|
|
BLAKE2b-256 checksum How to use checksums |
785f3b40e4be8c4309c854f6ff5eddc0c5f09c216603613ee081ac8f36dc13a9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency logRelease files / quant_regress-0.1.1-py3-none-any.whl
| Download URL | quant_regress-0.1.1-py3-none-any.whl |
|---|---|
| Size | 10.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7dfde2b9ed68a7c0f0e2c939f6f587250365f2cdd6d1990779a695e496e11adf
|
|
BLAKE2b-256 checksum How to use checksums |
039de674bc77f7ed599d583563d1b6e08c19e839b6f8b471abdca36a45572de0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency log