Skip to main content

quantcost

Find out what quantization actually costs you — on your machine, not someone else's.

CI PyPI Python License: MIT

pip install quantcost
quantcost run

That's it. Two to three minutes later you get speed, size, memory and quality for fp32 / int8 / int4 on your own CPU, with a plain-language verdict.

No PyTorch. No GPU. No quantizing anything yourself — the pre-quantized ONNX already published on the Hugging Face Hub gets downloaded and measured, so the whole install is ONNX Runtime, NumPy and a tokenizer, and finishes in about ten seconds.


What quantization actually costs

"Quantize it, it'll be smaller and faster" is half right. Two models, one Apple M4, measured with quantcost run at its defaults:

SmolLM2-135M-Instruct

Precision Size (MB) tok/s Speedup Peak RAM (MB) Perplexity PPL change
fp32 515 83.65 baseline 1360 23.07 baseline
int8 131 39.36 0.47x 1002 25.02 +8.5%
q4 174 57.40 0.69x 906 28.35 +22.9%

Qwen2.5-0.5B-Instruct

Precision Size (MB) tok/s Speedup Peak RAM (MB) Perplexity PPL change
fp32 1901 21.49 baseline 3282 19.26 baseline
int8 488 10.17 0.47x 3052 21.19 +10.0%
q4 750 18.37 0.85x 2303 22.28 +15.7%

Smaller: reliably, 3–4x. Less memory: yes, 25–30% off the peak. Faster: no — int8 generates tokens at 0.47x the fp32 rate on both models. That the figure lands on 0.47 twice, across a 135M and a 500M model, is what makes it look like a property of the runtime rather than an accident of one benchmark.

Why, and why it is not "quantized maths is slow"

Run the same artifacts over a 512-token prefill instead of one token at a time:

prefill (512 tok) decode (1 tok/step)
SmolLM2 int8 0.79x 0.43x
SmolLM2 q4 0.31x 0.68x
Qwen int8 0.97x 0.59x
Qwen q4 0.36x 1.02x

The int8 penalty roughly halves once there is a batch to amortise over (0.43 → 0.79, 0.59 → 0.97). That is the signature of a fixed per-step cost: ONNX Runtime unpacks the weights back to float inside each matmul, and a single token cannot amortise unpacking a whole weight matrix. fp32 decode is bandwidth-bound on reading weights; int8 decode is compute-bound on unpacking them, so it does strictly more work despite being 4x smaller.

Note also that int8 and q4 invert: q4 is the better choice for decode (0.68x / 1.02x vs int8's 0.43x / 0.59x) and much the worse for prefill (0.31x / 0.36x vs 0.79x / 0.97x). They use different kernels — MatMulNBits for q4, dynamic-quantize + MatMulInteger for int8 — with opposite strengths. Which format is right depends on whether your workload is prompt-heavy or generation-heavy.

One number that is easy to get wrong

Pin threads to your machine's performance cores, not every physical core. On this M4 (4 performance + 6 efficiency) pinning all ten cost int8 66% of its throughput — 24.1 vs 40.0 tok/s — and doubled run-to-run spread, because an ONNX Runtime parallel region ends on a barrier and one thread on an efficiency core gates the whole thing. quantcost does this by default and records how it decided in every card. It is the single easiest way to publish a wrong number, and it is how the first draft of this README got the figures above wrong.

So the honest answer to "should I quantize?" is measure it on your hardware, which is why this is a tool rather than a blog post. On this machine you quantize to fit in memory, not to go faster. On yours it may differ — that is the thing worth finding out.

See what other machines measured →

Add your machine

quantcost run
quantcost submit

submit validates your result, forks this repo, commits the card and opens the pull request for you. No GitHub CLI? It prints a prefilled link instead.

Unusual hardware is the most valuable kind: Raspberry Pi, old ThinkPads, Snapdragon laptops, bare-metal ARM. The interesting question is not who has the fastest laptop — it is where quantization pays off and where it backfires, and that only becomes visible across many real machines.

See CONTRIBUTING.md for what makes a good submission and exactly what a card contains (no hostname, no username, no paths — the fingerprint is one auditable function).

Usage

# A different model — anything with ONNX on the Hub works
quantcost run --model onnx-community/Qwen2.5-0.5B-Instruct

# Which precisions does a model actually publish?
quantcost models --model onnx-community/Qwen2.5-0.5B-Instruct

# Speed and size only; skips the perplexity pass and is much faster
quantcost run --precisions fp32,int8 --skip-perplexity

# Pin threads to compare against a specific configuration
quantcost run --threads 4

# Rebuild the leaderboard from every submitted card
quantcost leaderboard

Any Hub repo following the onnx-community / transformers.js layout works unchanged — that is thousands of models, including everything under onnx-community.

How the numbers are produced

The headline claim of this project is that its numbers are real, so the methodology is worth stating plainly.

Every precision is measured in its own subprocess. Peak RSS is a per-process high-water mark and an ONNX Runtime session does not return all of its arenas when dropped, so measuring several precisions in one process reports the union of their footprints and blames whichever ran last. Isolation is the only way the peak-RAM column means what it says.

The file cache is warmed before anything is timed. ONNX Runtime mmaps weights and faults them in lazily. Measured cold — straight after the download — an fp32 baseline came in at 17 tok/s; warm, the same machine and build measured 64. A 4x error on the number every other row is divided by is not a rounding detail, so the model file is read through once before the session is built.

Throughput is total tokens over total time, derived from the same mean latency that is reported, never the average of each run's own rate. Those are different statistics (ratio of means vs. mean of ratios) and they disagree by a few percent, which would make two columns of the same card contradict each other.

Decoding is greedy and fixed-length. No sampling, no early EOS stop, so every precision does exactly the same amount of work and throughput does not depend on the RNG.

Perplexity is scored against a corpus pinned inside the package — a fixed 65,342-byte slice of WikiText-2, hash-checked on load and recorded in every card. A leaderboard where machines score different text is not a leaderboard, and this also removes a heavy datasets dependency. Method: non-overlapping 512-token windows, summed next-token cross-entropy, exp(total_nll / total_tokens). The NumPy implementation is tested against the uniform-distribution case, where the right answer is analytically the vocabulary size, and agrees with PyTorch's cross_entropy to 4e-07 relative.

Thread count is always pinned to physical cores (SMT siblings contend for the same vector units) and recorded, because leaving it at ORT's default stores "0 — decide for me" and makes two very different runs look identical.

Instability is reported, not hidden. If run-to-run latency varies by more than 15%, the report says the machine was too busy and asks you to re-run.

What validation can and cannot do

CI rejects cards that are malformed, internally inconsistent, scored against a modified corpus, run with an unpinned thread count, or carrying identifying information. It cannot prove a number came from real silicon — nothing short of attested execution could. The defence is that every input is pinned, so a fabricated card must be self-consistent across all of them, and anyone can re-run the exact configuration recorded in the card. Cards are reviewed, not trusted.

Authoring quantized artifacts yourself

Everything above consumes pre-quantized ONNX. The original project also produces it — export, quantize, a C++ inference harness, a SIMD INT8 GEMM kernel, an Android app and a Qualcomm Snapdragon NPU path. That path needs the heavy dependencies:

pip install -e ".[quantize]"
edgellm export --help
edgellm quantize --help
edgellm bench --help      # the original torch + Optimum harness

See docs/AUTHORING.md for the full pipeline, the C++ harness and the Snapdragon notes.

Development

pip install -e ".[dev]"
ruff check . && ruff format --check . && pytest

CI asserts that importing the CLI pulls in no heavy dependency. If you need torch, import it inside the function that uses it — the fast install is what makes a one-command benchmark viable for someone who has never heard of this project.

License

MIT — see LICENSE. The bundled eval corpus is WikiText-2, CC BY-SA 4.0; see edgellm/data/SOURCE.md.

Metadata

Release files for quantcost 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for quantcost 0.2.0
File Size Uploaded
quantcost-0.2.0.tar.gz 187.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for quantcost 0.2.0
File Interpreter ABI Platform
quantcost-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 267.8 kB

Release files / quantcost-0.2.0.tar.gz

Download URL quantcost-0.2.0.tar.gz
Size 187.6 kB
Tags Source
SHA-256 checksum
How to use checksums
0a375c5f72eec4430a1dd51acf3c3baa0c5fa696e860d63cabc361d76cfc745b
BLAKE2b-256 checksum
How to use checksums
39517b44513246a2b63bc3679629e23d294ad3ef8d1d8992e5fcb3356fb3e26c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / quantcost-0.2.0-py3-none-any.whl

Download URL quantcost-0.2.0-py3-none-any.whl
Size 80.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
08fa41d851ac542171dc0e75bbe62e1da23a8173264fe2db1285527d4c6b1403
BLAKE2b-256 checksum
How to use checksums
48c1c76c5c36258619e6c26d27fc81e8537ef8747201e74df48c4d002fbbad30
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page