Skip to main content

jevkit-calibrate

Calibration is Jev's central claim: the probabilities are meant to track real frequencies, so that among answers given 0.8, about 80% are right. TypeSafe measures this across groups of predictions and says plainly that it does not guarantee any individual answer.

This package checks the claim on your data, which is the part nobody else can do for you, and turns the result into a threshold you can defend.

Unofficial and unaffiliated with TypeSafe.

pip install jevkit-calibrate

Pure standard library. No numpy, no pandas, no plotting stack. A calibration check that needs a build toolchain is a calibration check that does not get run.

Measure

from jevkit_core import read_records
from jevkit_calibrate import calibrate, observations_from_records

observations = observations_from_records(read_records("labeled.jevl"))
report = calibrate(observations)

print(report.summary())
print(report.diagram())
observations: 4000
accuracy:     0.7485
mean claimed: 0.7469  (underconfident by 0.0016)
ECE:          0.0094  (over 10 bins)
MCE:          0.0167
Brier:        0.1681
log loss:     0.5043

ECE is the average gap between claimed and observed, weighted by how many observations fall in each bin. MCE is the worst gap in any one bin, which is what catches a healthy-looking ECE hiding one badly wrong region. Brier and log loss are proper scoring rules: they reward being calibrated and decisive, so a model that always says 0.5 scores badly even though it is perfectly calibrated.

Pick a threshold

TypeSafe's confidence page recommends three bands and says where you draw them depends on your data. This draws them from the data.

from jevkit_calibrate import recommend_for_accuracy

point = recommend_for_accuracy(observations, target_accuracy=0.95)
if point is None:
    print("no threshold reaches 95% on this data")
else:
    print(f"threshold {point.threshold:.2f}: "
          f"covers {point.coverage:.1%} at {point.accuracy:.1%}, "
          f"{point.errors} wrong answers acted on")

None is a real answer, not a failure. It means this question cannot be automated at that bar, and the honest move is to change the question rather than lower the threshold.

CLI

jevkit-calibrate labeled.jevl
jevkit-calibrate labeled.jevl --target-accuracy 0.95
jevkit-calibrate labeled.jevl --min-coverage 0.80
jevkit-calibrate labeled.jevl --max-ece 0.05     # CI gate
jevkit-calibrate labeled.jevl --format json

Which quantity gets calibrated

By default, the probability mass on the chosen outcome, which is what "when it says 0.8, is it right 80% of the time" means.

--use-confidence calibrates the API's confidence statistic instead. That is a different question, and the one to ask when you want to know whether your routing threshold sits in the right place. Nouls carry no confidence, so they fall back to probability either way.

License

MIT

Metadata

Release files for jevkit-calibrate 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jevkit-calibrate 0.1.0
File Size Uploaded
jevkit_calibrate-0.1.0.tar.gz 9.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jevkit-calibrate 0.1.0
File Interpreter ABI Platform
jevkit_calibrate-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 20.8 kB

Release files / jevkit_calibrate-0.1.0.tar.gz

Download URL jevkit_calibrate-0.1.0.tar.gz
Size 9.6 kB
Tags Source
SHA-256 checksum
How to use checksums
0668d15fa7258645bd7c44ec6e8a5f727f7309a6b45369830798747e26447f73
BLAKE2b-256 checksum
How to use checksums
0db356c6f945f40c5c82dc8289a6cad2982a9e9287911452d2f21a3328db55a7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.12

Release files / jevkit_calibrate-0.1.0-py3-none-any.whl

Download URL jevkit_calibrate-0.1.0-py3-none-any.whl
Size 11.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c06702ad5e819ac6d03e42e6cd7a39a3532a3fefc9aa4a95b09df3c258b1a7db
BLAKE2b-256 checksum
How to use checksums
e07d76e860bd5d35a8ed048bfc0813ae777685dad21ad2bdd7cce48bb3d78277
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.12

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page