Skip to main content

MLVerdict

Evidence-based model selection for tabular machine learning.

PyPI Python License: MIT Downloads

Install from PyPI. Source: github.com/siva1252/ml_lib.

MLVerdict is not a black-box AutoML that dumps a leaderboard and walks away. You give it a table and a target. It diagnoses the problem, picks metrics that will not lie, trains a shortlist of models, optionally tunes the winners, and writes down why one model is the decision — including leakage signals, stability, held-out test, and a deployable artifact.

from mlverdict import Verdict

run = Verdict(enable_hpo=True).fit("train.csv", "churn")
print(run)                       # readable verdict, not a raw dict
run.artifact().save("model.joblib")

Why this exists

Typical sklearn / boosting workflow:

  1. Guess classification vs regression.
  2. Report accuracy on an imbalanced label.
  3. Train five models with default settings.
  4. Pick the highest number.
  5. Discover later that an ID column leaked, or that f1 = 0 while accuracy looked fine.

MLVerdict automates the decision path, not just training:

Step What it does
Understand Profile rows, types, imbalance, IDs, missingness
Diagnose Binary / multiclass / regression — or stop if the target is ambiguous
Plan Validation strategy + primary metric (e.g. pr_auc when classes are skewed)
Experiment Linear, trees, sklearn boosting; XGBoost / LightGBM / CatBoost if installed
HPO Bounded Optuna search on the top candidates only
Evaluate Performance + stability + generalization + latency + complexity
Decide Hard constraints first, then ranked evidence — no hardcoded favorite
Package Preprocess + model + schema + decision record, ready to save and serve

Install

Python 3.10+

pip install mlverdict

Upgrade later:

pip install -U mlverdict

Optional gradient-boosting libraries (same comparison, extra candidates):

pip install "mlverdict[all]"

Install from source (this repo):

git clone https://github.com/siva1252/ml_lib.git
cd ml_lib
pip install -e ".[dev]"

What you get with pip install

Package Role
mlverdict The engine (Verdict, Run, ModelArtifact, CLI)
pandas, numpy Tables
scikit-learn Linear / tree / HGB models, metrics, CV
optuna Bounded hyperparameter search
joblib Save / load artifacts

Two-minute start

from mlverdict import Verdict

run = Verdict(enable_hpo=True).fit("train_with_label.csv", "churn")
print(run)

That is the whole modeling loop. print(run) is the product surface: winner, why this metric, leaderboard table, held-out scores, traps such as high accuracy with zero recall, watchlist, and how to save the model.

Need a CSV first? From this repo:

python examples/quickstart.py

CLI

python -m mlverdict fit train_with_label.csv --target churn --save model.joblib
python -m mlverdict serve model.joblib

POST /predict with {"records": [{...}, ...]}.

Save, reload, predict

from mlverdict import ModelArtifact

run.artifact().save("model.joblib")
art = ModelArtifact.load("model.joblib")
art.predict(new_rows)

The joblib file is preprocess + model + schema. You do not rewrite training code to deploy.


What fit() actually does

CSV / DataFrame
    → lock a final test split (untouched until the end)
    → dataset DNA + quality + leakage signals
    → problem type
    → metric + validation plan
    → shortlist compatible models
    → cross-validation baselines
    → optional HPO on the top 2
    → multi-criteria ranking
    → score the held-out test once
    → production-readiness checks
    → trained artifact

Always in the comparison: Logistic Regression or Ridge, Random Forest, Extra Trees, HistGradientBoosting.

If you installed extras: XGBoost, LightGBM, CatBoost.

Tiny data skips overkill boosters. Identifiers such as customer_id are flagged and dropped from features.


Reading the verdict

Do not print run.best() and run.leaderboard() for humans. Those are for code.

Call For
print(run) The decision, in English
run.report() Full markdown record
run.best() Dict for scripts
run.leaderboard() DataFrame for scripts
run.artifact() Deployable pipeline

If held-out accuracy is high but precision / recall / f1 are 0, the model never predicted the rare class at the 0.5 cutoff. Ranking metrics (pr_auc, roc_auc) can still look fine. MLVerdict prints an ATTENTION block for that trap. Production packaging checks (save/reload) are not a certificate that the business metric is good.


Tests

pip install -e ".[dev]"
pytest

Coverage includes problem detection, leakage/quality, candidate selection, end-to-end fit on synthetic and sklearn datasets (breast cancer, iris, wine, diabetes), display/report formatting, and CLI save.


Project layout

ml_lib/
├── src/mlverdict/          # library (pip installable)
│   ├── api/                # Verdict, Run, CLI
│   ├── data/               # load, profile, DNA
│   ├── problem/            # classification vs regression
│   ├── metrics/            # metric plan
│   ├── models/             # catalog + sklearn / booster adapters
│   ├── experiments/        # CV + bounded HPO
│   ├── evaluation/         # stability, latency, composite score
│   ├── decision/           # winner + written record
│   ├── reporting/          # print(run) + markdown report
│   └── production/         # artifact + local serve
├── tests/                  # pytest
├── examples/               # runnable quickstart
├── pyproject.toml
├── LICENSE                 # MIT
└── README.md

Versioning PyPI vs Git

This repo is GitHub. PyPI is a separate upload.

You do GitHub pip install mlverdict
git push updates no
bump version + python -m build + twine upload no yes

Same version number cannot be uploaded twice. After this README, the published package is 0.1.1.


License

MIT. See LICENSE.

Metadata

Release files for mlverdict 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mlverdict 0.1.1
File Size Uploaded
mlverdict-0.1.1.tar.gz 62.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mlverdict 0.1.1
File Interpreter ABI Platform
mlverdict-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 131.8 kB

Release files / mlverdict-0.1.1.tar.gz

Download URL mlverdict-0.1.1.tar.gz
Size 62.3 kB
Tags Source
SHA-256 checksum
How to use checksums
618b24aa7ad7a109e6a68122615844c7f8eab4b521599d4458dc7dd17b6d5233
BLAKE2b-256 checksum
How to use checksums
fddeeb2938d21dd409844468ebc98ba7496ffe8fcab42a888135d913a979b770
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release files / mlverdict-0.1.1-py3-none-any.whl

Download URL mlverdict-0.1.1-py3-none-any.whl
Size 69.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d66c002cd360fa83545ec0980b3eb3c33e9e9e56c02047a53a5931f946361698
BLAKE2b-256 checksum
How to use checksums
e897eca458c4e554c2622b07eff9486742b103bcc5ca36174df5b7f9de6a9805
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release history Release notifications | RSS feed

0.2.1

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page