Skip to main content

clausius — measure the model as you run it

ci PyPI license

Benchmarks measure models; clausius measures the model as you run it — at your quantization, your cap, your thinking setting, on your prompts — with paired statistics and the truncation rate beside every number. Today that is two verbs: compare tells you whether a change broke the model, on your own prompts, with no labels, and exits non-zero so it drops into CI without glue; plan tells you how to run a model from measured cells.

Install

pip install "clausius[mlx]"        # capture + compare (Apple Silicon)
pip install clausius               # compare and analysis only — pure numpy, runs anywhere

Run it

Take 60 prompts of your own — production traffic is ideal, no labels needed — or the 60 in examples/prompts.jsonl. Capture the configuration you trust, capture the one you changed, compare:

clausius capture --model ./gemma-26b-a4b-4bit --prompts examples/prompts.jsonl --out ref.json  --max-tokens 1536
clausius capture --model ./gemma-26b-a4b-2bit --prompts examples/prompts.jsonl --out cand.json --max-tokens 1536
clausius compare ref.json cand.json --show 3
REGRESSION  (max d_z = +5.922, threshold 0.3, one-sided)
  compared 20 paired items, dropped 5 truncated
  all signals: max +5.92  p90 +11.64  mean +10.43  mean_top10 +9.00  first +2.80  gen_len +10.55

clausius compare on real captures: the 2-bit config is flagged REGRESSION with the ref's coherent reasoning shown against the 2-bit model's garbled output

That is a real run on 25 unlabeled prompts. The two checkpoints differ only in quantization; the 2-bit one independently measures 73 points lower on instruction adherence. --show 3 prints the items whose entropy moved most, with the text both configurations produced — the verdict says something broke, this says what.

First time? Try it in 30 minutes: three public checkpoints, five commands, measured output included. Then read USAGE.md: choosing prompts, setting the token cap, reading d_z and its interval, calibrating your own null, running it as a CI gate.

What the verdict means

  • The signal is the change in the model's predictive entropy over its own answers, paired prompt by prompt between the two captures (d_z).
  • The defaults are measured, not chosen: the 0.3 threshold comes from 13 configurations known to be harmless; the one-sided test from a construction that fools a two-sided one; the truncation filter from an effect that doubles once applied. src/clausius/core.py states each one and MEASUREMENTS has the evidence.
  • It is more sensitive than labels on the damage it was validated against: a −2.2pp quantization regression was flagged on 60 unlabeled prompts where a paired test on gold labels needed n=878 (F14).
  • It is a regression check, not a score: it needs a reference capture and cannot rate a configuration in isolation.

Limits

  • Sensitivity is the weaker half. Specificity is 13/13; a −5.7pp configuration is missed at any threshold that preserves that record.
  • Confidence-increasing damage would be invisible to a one-sided detector. Three mechanisms were built to produce it and none did, but it is not excluded.
  • d_z is ordinal, not proportional: 3-bit loses 25× more accuracy than 4-bit and reads 2× the d_z. "Something moved, and roughly how hard" is supportable; "you lost k points" is not.
  • The threshold is calibrated on one stack (MLX, Apple Silicon). On another framework, device, or quantizer, measure your own null first — USAGE.md has the recipe.
  • Local and self-hosted models only: hosted APIs expose no logprobs, or truncated ones, which is a different quantity.

Capture targets Apple Silicon via MLX in this release. An experimental PyTorch backend exists and is deliberately unshipped: it has been measured on mps and cpu, never calibrated on cuda or the CUDA-native quantizers, and shipping an uncalibrated threshold would contradict what this package claims about its defaults. It is archived at the tag archive/torch-backend rather than kept as a live branch, and will be revisited if a user needs it or a contribution calls for it; the bar it must clear is recorded in the research notes.

Receipts

The numbers behind the defaults, measured on consumer hardware (M5 Max, 128 GB) across five model families, with negatives kept: eight controlled failures and seven corrections to the record. Headline rows:

result where
Label-free detection works validated against five unrelated damage mechanisms whose true damage was measured independently; benign controls stay clean (a 3.3× memory reduction changes ~25% of generations textually and moves no signal) F8
More sensitive than labels the −2.2pp regression: 60 unlabeled prompts, where labels needed n=878 F14
Short benchmarks understate damage ~14× factual QA loses 1.5pp where structured generation loses 18–21pp under the same quantization F11
Offload beats downsizing an offloaded 35B at 3.40 GB scores 0.9447 on gsm8k against a natively-fitting 4B at 3.91 GB scoring 0.8426; the whole cost is latency F3–F6

Everything else — the quantization ladder and frontier chart, the claim taxonomy, what did not work, the seven corrections, prior art — is in MEASUREMENTS. The corpus and the research code that produced these results live in a separate research repository; MEASUREMENTS records each result with its n and its interval.

Ask how to run a model

clausius plan answers "thinking on or off, at what cap, and what does it cost" from measured cells — never from estimates. Each row is one configuration actually run: accuracy with a 95% interval, seconds per item, and the share of items that hit the output cap. One pick per task by a fixed rule (highest accuracy among rows that truncated at most half their items; ties go to the cheaper row). A row that truncated more than half is reported as cap-bound, because its accuracy is partly the cap's.

clausius plan                          # models on the card
clausius plan --model qwen3.5-35b-a3b --task aime-2024-25
  aime-2024-25
    strategy                        cap    n  accuracy           95% CI  s/item  trunc
    greedy                         2048   60     0.200   [0.118, 0.318]      19   0.80  cap-bound
    greedy                         8192   60     0.550   [0.425, 0.669]      47   0.67  cap-bound
    thinking                      16384   60     0.400   [0.286, 0.526]     121   0.62  cap-bound
    thinking                      32768   60     0.550   [0.425, 0.669]     204   0.47  ◀ pick

The card currently holds 37 rows across four Qwen models on gsm8k, MATH-500 L5, GPQA-Diamond, AIME, IFEval, BFCL, and LiveCodeBench, all measured on one machine (M5 Max, 128 GB) at 4-bit; the rows and their sources are in the packaged data/cards.json (--json prints them). Rows are added as cells are measured, never interpolated.

Where this is going

v0.2 adds report cards — the benchmarks official model cards report, re-measured at the configurations people actually deploy — and makes every measurement a module on one statistical core, so later measurements (agent trajectories, stopping behavior) arrive as rows on the same card rather than as new tools.

Sibling. boyle runs the model you want at the memory pressure you specify — budgeted MoE inference with speed forecasts before you download. Its predict cites this repository's measured accuracy.

Named for Rudolf Clausius, who coined the word entropy in 1865. Entropy is the signal this tool reads.

Repository layout

path what needs
src/clausius/ the tool — capture, compare, plan, CLI; data/cards.json is the measured card numpy; mlx-lm only to capture
src/clausius/measure/ the measurement engine: run a strategy over items against any OpenAI-compatible endpoint, one JSONL record per item with resume, paired statistics (exact McNemar, Wilson intervals), manifests → card rows stdlib
tests/ 80 tests, none load a model; CI installs the built wheel numpy
USAGE.md the operating manual —
docs/measurements.md the measured basis for the tool's defaults and claims —

Metadata

Release files for clausius 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for clausius 0.1.2
File Size Uploaded
clausius-0.1.2.tar.gz 42.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for clausius 0.1.2
File Interpreter ABI Platform
clausius-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 82.2 kB

Release files / clausius-0.1.2.tar.gz

Download URL clausius-0.1.2.tar.gz
Size 42.8 kB
Tags Source
SHA-256 checksum
How to use checksums
f2fb31465091d07f6c382241cf6ad69efc35dd5563cc7129556ac139bd7f9ac7
BLAKE2b-256 checksum
How to use checksums
adb0fb25f8c3eefe5633e7c3d128755a6a1b98fb20abab50c04e51c474d6a576
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release files / clausius-0.1.2-py3-none-any.whl

Download URL clausius-0.1.2-py3-none-any.whl
Size 39.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
18c738f5e7c08a457a724b9c4f5b68e78c38e9f2d2c75f2e41888e33941a466d
BLAKE2b-256 checksum
How to use checksums
519ad549b8dbef734b3f9be58d0374c1f4344ae09997619bc81a9ccbe1c0013a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page