quant-regress
Fail CI when quantization costs you task accuracy.
Token-level proxies — perplexity, KL divergence — routinely report that a quantized model is fine. Task accuracy does not, provided you actually measure it. This action measures a task eval at each precision and fails the build when accuracy drops past a threshold you set.
Why
On a cross-encoder claim verifier, dynamic INT8 quantization left aggregate accuracy looking survivable while destroying the capability that mattered:
| precision | accuracy | correct |
|---|---|---|
| fp32 (baseline) | 100.0% | 20/20 |
| int8 | 0.0% | 0/20 |
Regenerate with python scripts/regen_int8_collapse.py — every figure above
comes from that script, and tests/test_claims_are_reproducible.py fails the
build if a number in this file stops matching it.
The pattern that matters: the quantized model still runs, still returns a valid answer, and still looks like a working model. It is simply wrong every time. That is why a token-level proxy can clear it and a task eval cannot.
Usage
- uses: DawnofGenX/quant-regress@v1
with:
eval: evals/tasks.jsonl
model: meta-llama/Llama-3.2-1B-Instruct
max-drop-points: 2
Exit codes: 0 within tolerance, 1 accuracy regressed, 2 bad configuration.
Constraints worth knowing
- CPU only.
quantize_dynamicemits CPU-only quantized kernels; moving the model to CUDA raises at forward time. GitHub-hosted runners have no GPU, so this is not a limitation in CI — but it does mean GPU-only quantizers (AWQ/GPTQ/GGUF) are out of scope for v1. - The threshold unit is accuracy points, not a ratio.
--max-drop-points: 2means "fail if any candidate is more than 2 points below the baseline". - Comparison is exact-after-normalisation by default. Pass a scorer if you need semantic equivalence.
- Models are downloaded at run time. A 1B model is ~2 GB; check runner disk.
Honest limits
- Only
fp32and dynamicint8are supported in v1. - The library runs generation, so it needs a task whose success is checkable in code. It is not an LLM-judge harness.
- The bundled
evals/fixtures/collapse.jsonlis model-specific. Its cases expect the answeryes, so it only demonstrates a collapse for a model that answersyeswhen unquantized. Pointed at an untuned tiny LM it scores 0/N at both precisions, produces no drop, and the gate correctly reports PASS. Use it to see the output format; bring your own eval for a real check. - What the repo's own CI pins is the exit-code contract (0 holds, 1 regressed,
2 misconfigured), not any particular model's accuracy — see
tests/test_selftest_gate.py. - No published adoption yet. If you use it, an issue saying what you quantized and what broke is genuinely useful.
Metadata
Release files for quant-regress 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| quant_regress-0.1.0.tar.gz | 15.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| quant_regress-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.7 kB
Release files / quant_regress-0.1.0.tar.gz
| Download URL | quant_regress-0.1.0.tar.gz |
|---|---|
| Size | 15.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7b1149571bd1b366f80f17ff3e7bd860353a729bdee5e675f8b521451b6ff2bc
|
|
BLAKE2b-256 checksum How to use checksums |
05d95122760df66e4489922b64d315f29aa4c6fa0efd18ae692a701d23dde03e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|
Release files / quant_regress-0.1.0-py3-none-any.whl
| Download URL | quant_regress-0.1.0-py3-none-any.whl |
|---|---|
| Size | 10.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e54cc448f368f0e797417c513e3b225e3309273d00df29bb3ddd3469a706cd96
|
|
BLAKE2b-256 checksum How to use checksums |
f80053e39954ec1c012b1265df03f46c53c2c68b64ad5001d01cb481d8c3635d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|