lfmm: monitoring ML models before the labels arrive
A benchmark and toolkit for performance estimation, drift detection and retraining decisions when true labels arrive late: a loan default is known up to two years after the loan is made, a medical outcome weeks later, a fraud chargeback months later.
Most monitoring methods are evaluated on artificially injected drift with every label available immediately. lfmm replays real, timestamped data as streams in which a monitor sees each label only when it would have arrived in practice, including the natural label delay of mortgage defaults, and scores monitors on four tasks.
- Paper draft: paper/main.pdf
- Leaderboard: leaderboard/LEADERBOARD.md
Install
pip install lfmm # core: streams, tasks, estimators, detectors, retraining
pip install "lfmm[streaming]" # + ADWIN detectors (river)
pip install "lfmm[data]" # + downloaders and loaders for the benchmark datasets
pip install "lfmm[all]" # everything needed to reproduce the paper
Python 3.10 or newer.
Quickstart
Any table of scored events becomes a stream: features, label, when each row was scored, and when its label becomes known.
import numpy as np
import pandas as pd
from lfmm.estimators import CBPE, DelayAdjustedCBPE, LatestCompleteCohort
from lfmm.harness import run_estimation
from lfmm.streams.base import Stream, with_label_delay
rng = np.random.default_rng(0)
n = 24_000
month = np.repeat(np.arange(24), n // 24)
X = pd.DataFrame({"x1": rng.normal(month / 12, 1), "x2": rng.normal(0, 1, n)})
y = (rng.random(n) < 1 / (1 + np.exp(-(X.x1 - X.x2 - 1)))).astype(float)
t = (np.datetime64("2022-01", "M") + month).astype("datetime64[ns]")
stream = Stream("demo", X, y, event_time=t, label_time=t)
stream = with_label_delay(stream, np.timedelta64(60, "D")) # labels arrive 60 days late
# Train on fully labeled cohorts before 2023, freeze the model, walk forward monthly.
result = run_estimation(
stream, train_end="2023-01-01", freq="M", metrics=("roc_auc", "prevalence"),
estimators=[CBPE(), DelayAdjustedCBPE(), LatestCompleteCohort(freq="M")],
)
print(result.summary()) # MAE and bias of each estimator
print(result.bootstrap_mae("prevalence")) # with 95% block-bootstrap intervals
Monitors only ever see labels whose arrival time has passed, and the model is trained only on cohorts whose labels have (almost) all arrived.
Tasks
| Task | Question | Scored by |
|---|---|---|
| Estimation (next batch) | What will this batch's metric turn out to be? | MAE vs the true metric, block-bootstrap intervals |
| Estimation (in-flight book) | What is the metric of everything scored in the last N months, part of it already labeled? | same |
| Detection | Has the model got worse than at deployment? | rank correlation with true degradation, AUROC, alarm precision/recall |
| Retrain or not | Retrain now, given the labels that have arrived? | total loss + λ × retrains, regret vs the hindsight-optimal schedule |
from lfmm.harness import run_inflight, run_detection_events
from lfmm.retrain import build_bank, evaluate_policies
Methods included
- Estimators: reference performance, recent arrivals, latest complete cohort, labels to date, CBPE, ATC, DoC, and delay-adjusted CBPE (ours) with its hazard variant for calendar-time shocks such as COVID forbearance.
- Detectors: univariate KS/χ², PSI, KS on scores, domain classifier, MMD, ADWIN on scores and on arrived errors, and alarms from any estimator.
- Retraining policies: never, always, every k periods, retrain on any detector's alarm, and the hindsight-optimal schedule (dynamic programming).
Add your own method
Subclass Estimator, Detector or Policy. An estimator is fit on a labeled reference window and then sees each batch's model scores plus a History holding only the labels that have arrived:
import numpy as np
from lfmm.estimators import Estimator, History
class MeanScore(Estimator):
name = "mean_score"
supports = ("prevalence",)
def estimate(self, metric: str, proba: np.ndarray, history: History) -> float:
return float(proba.mean())
The benchmark datasets
| Stream | Source | Label | Label delay |
|---|---|---|---|
| Freddie Mac | Single-Family Loan-Level Dataset, 1999–2026 | default within 24 months | natural |
| ACS income | US Census PUMS (via Folktables), 2014–2024 | income > $50k | simulated, 365 days |
| BRFSS | CDC survey, 2011–2024 | diagnosed diabetes | simulated, 180 days |
| NYC taxi | TLC trip records, 2019–2021 | tip ≥ 25% of fare | simulated, 30 days |
| TabReD | homecredit, homesite, ecom (Rubachev et al.) | default / conversion / repeat purchase | simulated, 14–90 days |
export LFMM_DATA_DIR=/path/to/lfmm-data
lfmm-download acs # also: tlc, brfss, all
lfmm-download sflld # register hand-downloaded Freddie Mac files
lfmm-download tabred # register hand-downloaded TabReD archives
Downloads resume and write each file's source URL and SHA-256 to $LFMM_DATA_DIR/manifest.csv. Freddie Mac and TabReD need an account or accepted terms, so they are downloaded by hand; no raw data is redistributed. The eight experiments of the paper are defined in lfmm.experiments.EXPERIMENTS.
Reproducing the paper
From a clone of the repository:
uv sync --all-extras
uv run python scripts/run_estimation.py # next-batch estimation -> results/
uv run python scripts/run_inflight.py # in-flight book (Freddie Mac)
uv run python scripts/run_detection.py # degradation detection
uv run python scripts/run_retrain.py # retrain or not (slow: one model per period)
uv run python scripts/run_seeds.py # estimation over 5 seeds
uv run python scripts/make_figures.py # figures/
uv run python scripts/make_tables.py # paper/tables/
uv run python scripts/make_appendix.py # appendix tables
uv run python scripts/build_leaderboard.py # leaderboard/
uv run pytest # fast tests; `uv run pytest -m slow` for the rest
Project notes: PROGRESS.md (status, design decisions, results) and DATA.md (data sources and construction).
License
MIT; see LICENSE. The datasets keep their own terms.
Metadata
Release files for lfmm 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| lfmm-0.1.0.tar.gz | 41.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| lfmm-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 91.7 kB
Release files / lfmm-0.1.0.tar.gz
| Download URL | lfmm-0.1.0.tar.gz |
|---|---|
| Size | 41.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9ebb2419bc6a40fa66ecba506811da781ecefe3957a4295f4ee3ca8fc0950b33
|
|
BLAKE2b-256 checksum How to use checksums |
9453e115c7107c4de8d453209810cbe223a76dff924a8052825f31f0367cf933
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / lfmm-0.1.0-py3-none-any.whl
| Download URL | lfmm-0.1.0-py3-none-any.whl |
|---|---|
| Size | 50.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6f95b8d826b0117acccb03d4bab6a747723614eb749a800a39683d1ac4768dde
|
|
BLAKE2b-256 checksum How to use checksums |
257d376960f33104aca422248353cde1e99ccbdde47d1d6f5d3ad2fb35d8d59c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log