Skip to main content

Fiducio

Tests Docs Python License

Fiducio is a model-agnostic Python library of post-hoc calibrators for 2D and 3D semantic segmentation. Give it the logits or probabilities of any segmentation model and a labelled calibration set, and a calibrator returns better-calibrated probabilities without retraining the segmentation model. It implements temperature, vector, matrix and Dirichlet scaling, their translation-invariant variants, and the class-conditional (CDC, CMSap, CMSop) calibrators introduced in our paper Rethinking Post-Hoc Calibration in Semantic Segmentation, accepted at TMLR. Whether the segmentation is preserved depends on the calibrator; see the guarantees below.

It works with PyTorch U-Net, nnU-Net, SegFormer or any model that produces logits: Fiducio only ever sees (B, C, *spatial) tensors and never needs to know the architecture.

Reliability diagrams of an over-confident model before and after calibration with TemperatureScaling and CMSop

Top-1 reliability diagrams on held-out synthetic data (regenerate with examples/readme_figure.py).

Install

For development from a checkout, use the committed uv lockfile:

uv sync --locked
uv run --no-sync pytest

This uses CPU PyTorch for development. See CONTRIBUTING.md for the pinned uv version, documentation and build commands.

For an existing uv application before the first PyPI release:

uv add "fiducio @ git+https://github.com/fiducio-ai/Fiducio.git"

Once a PyPI release is available, the standard installation is:

pip install fiducio

With pip, before the first PyPI release, install from GitHub:

pip install "git+https://github.com/fiducio-ai/Fiducio.git"

The core install depends only on numpy and torch. Optional extras: fiducio[plots] (matplotlib reliability diagrams), fiducio[docs], fiducio[dev], fiducio[all].

Quick start

import torch
from fiducio import TemperatureScaling, load_calibrator

# From a held-out, labelled calibration set:
logits = torch.randn(8, 4, 64, 64)          # (B, C, H, W) model logits
labels = torch.randint(0, 4, (8, 64, 64))   # (B, H, W) labels

calibrator = TemperatureScaling(input_type="logits").fit(logits, labels)

# Apply to new predictions (returns calibrated probabilities):
probs = calibrator.transform(torch.randn(2, 4, 64, 64))

# Save and reload (loads on CPU by default):
calibrator.save("calibrator.pt")
calibrator = load_calibrator("calibrator.pt")

Every calibrator exposes the same API: fit(predictions, targets, mask=None), transform, predict_proba, fit_transform, save, and the top-level load_calibrator.

Accepted tensor shapes

The class axis is always dimension 1.

Use case predictions targets
2D segmentation (B, C, H, W) (B, H, W)
3D segmentation (B, C, D, H, W) (B, D, H, W)
general n-D (B, C, *spatial) (B, *spatial)
per-pixel table (N, C) (N,)

Inputs may be logits or probs (set input_type). Binary segmentation is two channels (C = 2). mask and ignore_index exclude voxels from fitting.

Calibrators

Class Alias Translation-invariant Decision preservation
TemperatureScaling TS yes argmax and full order
EnsembleTemperatureScaling ETS yes argmax and full order
VectorScaling VS no —
MatrixScaling MS no —
TranslationInvariantMatrixScaling MSc yes —
DirichletCalibration DC yes —
ClassConditionalMatrixScaling CMS / CDC yes —
ArgmaxPreservingMatrixScaling CMSAP / CMSap yes argmax
OrderPreservingMatrixScaling CMSOP / CMSop yes argmax and full order

Translation-invariant means the output is unchanged when the same constant is added to every input logit of a voxel.

Every calibrator also exposes decision_function (calibrated logits) and a configurable optimizer ("adam" default, or "lbfgs").

Early stopping

Adam-fitted calibrators support the paper's recipe of early stopping on a validation set (with an optional learning-rate decay on plateaus):

from fiducio import OrderPreservingMatrixScaling

calibrator = OrderPreservingMatrixScaling(max_iter=2000, patience=20, lr_patience=10)
calibrator.fit(
    logits, labels,
    val_predictions=val_logits, val_targets=val_labels,  # held out from both
)

The iterate with the best validation NLL is kept. The default fixed budget (max_iter) is short, so it can underfit expressive calibrators; prefer a larger max_iter together with patience for CDC, CMSap and CMSop.

Paper method mapping

Paper method Class
TS TemperatureScaling
MS MatrixScaling
MSc TranslationInvariantMatrixScaling
CDC ClassConditionalMatrixScaling
CMSap ArgmaxPreservingMatrixScaling
CMSop OrderPreservingMatrixScaling

CDC, CMSap and CMSop are class-conditional: they fit one affine map per uncalibrated top class, all experts optimized jointly by a single optimizer minimizing one cross-entropy loss over every voxel at once, with L2 regularization (lambda_reg / mu_reg) applied to the affine map induced in the common logit space rather than to the raw per-expert parameters — this keeps the regularization meaningful for CMSap/CMSop, whose parameters are non-negative margins/gaps rather than raw matrix entries. Pass independent_experts=True to instead fit each expert in its own optimization loop on only the voxels routed to it.

Calibration metrics are included: negative_log_likelihood, expected_calibration_error (ECE, population-weighted), average_calibration_error (ACE, unweighted over the same bins — doesn't let a sparsely populated bin get drowned out by a large one), brier_score, reliability_curve, plus an optional reliability-diagram plot (fiducio.plots.reliability_diagram, needs fiducio[plots]).

Documentation

The paper implementation guide records method mapping, numerical-reference coverage, fitting limits and the remaining steps needed to reproduce the paper's experiments. For ensemble inputs, see examples/ensemble_pooling.py.

Results in the paper

The paper compares these calibrators on real segmentation benchmarks. The reliability diagrams below (paper appendix) show the uncalibrated ensemble pooling (p0, p̄) against the best calibrator per dataset; the exact protocol, datasets and metrics are in the paper. This repository ships the calibrators, not the fitted models or benchmark pipelines, so it does not regenerate these figures (see the implementation guide).

Reliability diagrams from the paper on three datasets

Full guide and API reference: https://fiducio-ai.github.io/Fiducio/

Citation

Fiducio accompanies the paper Rethinking Post-Hoc Calibration in Semantic Segmentation (Kirscher et al., Transactions on Machine Learning Research, 2026; OpenReview, preprint: arXiv:2607.01902). If you use Fiducio in your research, please cite it using the metadata in CITATION.cff, or:

@article{kirscher2026rethinking,
  title   = {Rethinking Post-Hoc Calibration in Semantic Segmentation},
  author  = {Kirscher, Tristan and Kahl, Kim-Celine and Kovacs, Balint and
             Rokuss, Maximilian and Maier-Hein, Klaus and Coubez, Xavier and
             Meyer, Philippe and Faisan, Sylvain},
  journal = {Transactions on Machine Learning Research},
  issn    = {2835-8856},
  year    = {2026},
  url     = {https://openreview.net/forum?id=xwNoSNxgxV}
}

Status

Fiducio is beta (0.x): usable and tested, with an API that may still change before 1.0. It is not certified for safety-critical or clinical use.

License

Apache-2.0.

Release files for fiducio 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fiducio 0.1.0
File Size Uploaded
fiducio-0.1.0.tar.gz 273.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for fiducio 0.1.0
File Interpreter ABI Platform
fiducio-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 316.6 kB

Release files / fiducio-0.1.0.tar.gz

Download URL fiducio-0.1.0.tar.gz
Size 273.8 kB
Tags Source
SHA-256 checksum
How to use checksums
e9ad428fa7603f847844bc439abd52818bf5ca20798a443885096b1936b8a983
BLAKE2b-256 checksum
How to use checksums
9650be83c47bc176943fed976f4a324482532c058060e2aaa290b9df0f12dab2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / fiducio-0.1.0-py3-none-any.whl

Download URL fiducio-0.1.0-py3-none-any.whl
Size 42.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
418092b2f5c6c70d5c11aed23aad5160f93135ec54e1c9d78a3d9cd34d69de63
BLAKE2b-256 checksum
How to use checksums
6cb3df41c68e07d385bb99ee12df3eef76fb196b8d7042dedca677e3ec345826
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page