Skip to main content

quant-regress

Fail CI when quantization costs you task accuracy.

Token-level proxies — perplexity, KL divergence — routinely report that a quantized model is fine. Task accuracy does not, provided you actually measure it. This action measures a task eval at each precision and fails the build when accuracy drops past a threshold you set.

Why

On a cross-encoder claim verifier, dynamic INT8 quantization left aggregate accuracy looking survivable while destroying the capability that mattered:

precision accuracy correct
fp32 (baseline) 100.0% 20/20
int8 0.0% 0/20

Regenerate with python scripts/regen_int8_collapse.py — every figure above comes from that script, and tests/test_claims_are_reproducible.py fails the build if a number in this file stops matching it.

The pattern that matters: the quantized model still runs, still returns a valid answer, and still looks like a working model. It is simply wrong every time. That is why a token-level proxy can clear it and a task eval cannot.

Usage

- uses: DawnofGenX/quant-regress@v1
  with:
    eval: evals/tasks.jsonl
    model: meta-llama/Llama-3.2-1B-Instruct
    max-drop-points: 2

Exit codes: 0 within tolerance, 1 accuracy regressed, 2 bad configuration.

Constraints worth knowing

  • CPU only. quantize_dynamic emits CPU-only quantized kernels; moving the model to CUDA raises at forward time. GitHub-hosted runners have no GPU, so this is not a limitation in CI — but it does mean GPU-only quantizers (AWQ/GPTQ/GGUF) are out of scope for v1.
  • The threshold unit is accuracy points, not a ratio. --max-drop-points: 2 means "fail if any candidate is more than 2 points below the baseline".
  • Comparison is exact-after-normalisation by default. Pass a scorer if you need semantic equivalence.
  • Models are downloaded at run time. A 1B model is ~2 GB; check runner disk.

Honest limits

  • Only fp32 and dynamic int8 are supported in v1.
  • The library runs generation, so it needs a task whose success is checkable in code. It is not an LLM-judge harness.
  • The bundled evals/fixtures/collapse.jsonl is model-specific. Its cases expect the answer yes, so it only demonstrates a collapse for a model that answers yes when unquantized. Pointed at an untuned tiny LM it scores 0/N at both precisions, produces no drop, and the gate correctly reports PASS. Use it to see the output format; bring your own eval for a real check.
  • What the repo's own CI pins is the exit-code contract (0 holds, 1 regressed, 2 misconfigured), not any particular model's accuracy — see tests/test_selftest_gate.py.
  • No published adoption yet. If you use it, an issue saying what you quantized and what broke is genuinely useful.

Metadata

Release files for quant-regress 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for quant-regress 0.1.0
File Size Uploaded
quant_regress-0.1.0.tar.gz 15.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for quant-regress 0.1.0
File Interpreter ABI Platform
quant_regress-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 25.7 kB

Release files / quant_regress-0.1.0.tar.gz

Download URL quant_regress-0.1.0.tar.gz
Size 15.2 kB
Tags Source
SHA-256 checksum
How to use checksums
7b1149571bd1b366f80f17ff3e7bd860353a729bdee5e675f8b521451b6ff2bc
BLAKE2b-256 checksum
How to use checksums
05d95122760df66e4489922b64d315f29aa4c6fa0efd18ae692a701d23dde03e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / quant_regress-0.1.0-py3-none-any.whl

Download URL quant_regress-0.1.0-py3-none-any.whl
Size 10.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e54cc448f368f0e797417c513e3b225e3309273d00df29bb3ddd3469a706cd96
BLAKE2b-256 checksum
How to use checksums
f80053e39954ec1c012b1265df03f46c53c2c68b64ad5001d01cb481d8c3635d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page