Skip to main content

lfmm: monitoring ML models before the labels arrive

PyPI Python License: MIT

A benchmark and toolkit for performance estimation, drift detection and retraining decisions when true labels arrive late: a loan default is known up to two years after the loan is made, a medical outcome weeks later, a fraud chargeback months later.

Most monitoring methods are evaluated on artificially injected drift with every label available immediately. lfmm replays real, timestamped data as streams in which a monitor sees each label only when it would have arrived in practice, including the natural label delay of mortgage defaults, and scores monitors on four tasks.

Install

pip install lfmm                 # core: streams, tasks, estimators, detectors, retraining
pip install "lfmm[streaming]"    # + ADWIN detectors (river)
pip install "lfmm[data]"         # + downloaders and loaders for the benchmark datasets
pip install "lfmm[all]"          # everything needed to reproduce the paper

Python 3.10 or newer.

Quickstart

Any table of scored events becomes a stream: features, label, when each row was scored, and when its label becomes known.

import numpy as np
import pandas as pd
from lfmm.estimators import CBPE, DelayAdjustedCBPE, LatestCompleteCohort
from lfmm.harness import run_estimation
from lfmm.streams.base import Stream, with_label_delay

rng = np.random.default_rng(0)
n = 24_000
month = np.repeat(np.arange(24), n // 24)
X = pd.DataFrame({"x1": rng.normal(month / 12, 1), "x2": rng.normal(0, 1, n)})
y = (rng.random(n) < 1 / (1 + np.exp(-(X.x1 - X.x2 - 1)))).astype(float)
t = (np.datetime64("2022-01", "M") + month).astype("datetime64[ns]")

stream = Stream("demo", X, y, event_time=t, label_time=t)
stream = with_label_delay(stream, np.timedelta64(60, "D"))   # labels arrive 60 days late

# Train on fully labeled cohorts before 2023, freeze the model, walk forward monthly.
result = run_estimation(
    stream, train_end="2023-01-01", freq="M", metrics=("roc_auc", "prevalence"),
    estimators=[CBPE(), DelayAdjustedCBPE(), LatestCompleteCohort(freq="M")],
)
print(result.summary())                     # MAE and bias of each estimator
print(result.bootstrap_mae("prevalence"))   # with 95% block-bootstrap intervals

Monitors only ever see labels whose arrival time has passed, and the model is trained only on cohorts whose labels have (almost) all arrived.

Tasks

Task Question Scored by
Estimation (next batch) What will this batch's metric turn out to be? MAE vs the true metric, block-bootstrap intervals
Estimation (in-flight book) What is the metric of everything scored in the last N months, part of it already labeled? same
Detection Has the model got worse than at deployment? rank correlation with true degradation, AUROC, alarm precision/recall
Retrain or not Retrain now, given the labels that have arrived? total loss + λ × retrains, regret vs the hindsight-optimal schedule
from lfmm.harness import run_inflight, run_detection_events
from lfmm.retrain import build_bank, evaluate_policies

Methods included

  • Estimators: reference performance, recent arrivals, latest complete cohort, labels to date, CBPE, ATC, DoC, and delay-adjusted CBPE (ours) with its hazard variant for calendar-time shocks such as COVID forbearance.
  • Detectors: univariate KS/χ², PSI, KS on scores, domain classifier, MMD, ADWIN on scores and on arrived errors, and alarms from any estimator.
  • Retraining policies: never, always, every k periods, retrain on any detector's alarm, and the hindsight-optimal schedule (dynamic programming).

Add your own method

Subclass Estimator, Detector or Policy. An estimator is fit on a labeled reference window and then sees each batch's model scores plus a History holding only the labels that have arrived:

import numpy as np
from lfmm.estimators import Estimator, History

class MeanScore(Estimator):
    name = "mean_score"
    supports = ("prevalence",)

    def estimate(self, metric: str, proba: np.ndarray, history: History) -> float:
        return float(proba.mean())

The benchmark datasets

Stream Source Label Label delay
Freddie Mac Single-Family Loan-Level Dataset, 1999–2026 default within 24 months natural
ACS income US Census PUMS (via Folktables), 2014–2024 income > $50k simulated, 365 days
BRFSS CDC survey, 2011–2024 diagnosed diabetes simulated, 180 days
NYC taxi TLC trip records, 2019–2021 tip ≥ 25% of fare simulated, 30 days
TabReD homecredit, homesite, ecom (Rubachev et al.) default / conversion / repeat purchase simulated, 14–90 days
export LFMM_DATA_DIR=/path/to/lfmm-data
lfmm-download acs          # also: tlc, brfss, all
lfmm-download sflld        # register hand-downloaded Freddie Mac files
lfmm-download tabred       # register hand-downloaded TabReD archives

Downloads resume and write each file's source URL and SHA-256 to $LFMM_DATA_DIR/manifest.csv. Freddie Mac and TabReD need an account or accepted terms, so they are downloaded by hand; no raw data is redistributed. The eight experiments of the paper are defined in lfmm.experiments.EXPERIMENTS.

Reproducing the paper

From a clone of the repository:

uv sync --all-extras
uv run python scripts/run_estimation.py     # next-batch estimation -> results/
uv run python scripts/run_inflight.py       # in-flight book (Freddie Mac)
uv run python scripts/run_detection.py      # degradation detection
uv run python scripts/run_retrain.py        # retrain or not (slow: one model per period)
uv run python scripts/run_seeds.py          # estimation over 5 seeds
uv run python scripts/make_figures.py       # figures/
uv run python scripts/make_tables.py        # paper/tables/
uv run python scripts/make_appendix.py      # appendix tables
uv run python scripts/build_leaderboard.py  # leaderboard/
uv run pytest                               # fast tests; `uv run pytest -m slow` for the rest

Project notes: PROGRESS.md (status, design decisions, results) and DATA.md (data sources and construction).

License

MIT; see LICENSE. The datasets keep their own terms.

Metadata

Release files for lfmm 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lfmm 0.1.0
File Size Uploaded
lfmm-0.1.0.tar.gz 41.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lfmm 0.1.0
File Interpreter ABI Platform
lfmm-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 91.7 kB

Release files / lfmm-0.1.0.tar.gz

Download URL lfmm-0.1.0.tar.gz
Size 41.2 kB
Tags Source
SHA-256 checksum
How to use checksums
9ebb2419bc6a40fa66ecba506811da781ecefe3957a4295f4ee3ca8fc0950b33
BLAKE2b-256 checksum
How to use checksums
9453e115c7107c4de8d453209810cbe223a76dff924a8052825f31f0367cf933
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / lfmm-0.1.0-py3-none-any.whl

Download URL lfmm-0.1.0-py3-none-any.whl
Size 50.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6f95b8d826b0117acccb03d4bab6a747723614eb749a800a39683d1ac4768dde
BLAKE2b-256 checksum
How to use checksums
257d376960f33104aca422248353cde1e99ccbdde47d1d6f5d3ad2fb35d8d59c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page