MLVerdict
Evidence-based model selection for tabular machine learning.
Install from PyPI. Source: github.com/siva1252/ml_lib.
MLVerdict is not a black-box AutoML that dumps a leaderboard and walks away. You give it a table and a target. It diagnoses the problem, picks metrics that will not lie, trains a shortlist of models, optionally tunes the winners, and writes down why one model is the decision — including leakage signals, stability, held-out test, and a deployable artifact.
from mlverdict import Verdict
run = Verdict(enable_hpo=True).fit("train.csv", "churn")
print(run) # readable verdict, not a raw dict
run.artifact().save("model.joblib")
Why this exists
Typical sklearn / boosting workflow:
- Guess classification vs regression.
- Report accuracy on an imbalanced label.
- Train five models with default settings.
- Pick the highest number.
- Discover later that an ID column leaked, or that
f1 = 0while accuracy looked fine.
MLVerdict automates the decision path, not just training:
| Step | What it does |
|---|---|
| Understand | Profile rows, types, imbalance, IDs, missingness |
| Diagnose | Binary / multiclass / regression — or stop if the target is ambiguous |
| Plan | Validation strategy + primary metric (e.g. pr_auc when classes are skewed) |
| Experiment | Linear, trees, sklearn boosting; XGBoost / LightGBM / CatBoost if installed |
| HPO | Bounded Optuna search on the top candidates only |
| Evaluate | Performance + stability + generalization + latency + complexity |
| Decide | Hard constraints first, then ranked evidence — no hardcoded favorite |
| Package | Preprocess + model + schema + decision record, ready to save and serve |
Install
Python 3.10+
pip install mlverdict
Upgrade later:
pip install -U mlverdict
Optional gradient-boosting libraries (same comparison, extra candidates):
pip install "mlverdict[all]"
Install from source (this repo):
git clone https://github.com/siva1252/ml_lib.git
cd ml_lib
pip install -e ".[dev]"
What you get with pip install
| Package | Role |
|---|---|
mlverdict |
The engine (Verdict, Run, ModelArtifact, CLI) |
| pandas, numpy | Tables |
| scikit-learn | Linear / tree / HGB models, metrics, CV |
| optuna | Bounded hyperparameter search |
| joblib | Save / load artifacts |
Two-minute start
from mlverdict import Verdict
run = Verdict(enable_hpo=True).fit("train_with_label.csv", "churn")
print(run)
That is the whole modeling loop. print(run) is the product surface: winner, why this metric, leaderboard table, held-out scores, traps such as high accuracy with zero recall, watchlist, and how to save the model.
Need a CSV first? From this repo:
python examples/quickstart.py
CLI
python -m mlverdict fit train_with_label.csv --target churn --save model.joblib
python -m mlverdict serve model.joblib
POST /predict with {"records": [{...}, ...]}.
Save, reload, predict
from mlverdict import ModelArtifact
run.artifact().save("model.joblib")
art = ModelArtifact.load("model.joblib")
art.predict(new_rows)
The joblib file is preprocess + model + schema. You do not rewrite training code to deploy.
What fit() actually does
CSV / DataFrame
→ lock a final test split (untouched until the end)
→ dataset DNA + quality + leakage signals
→ problem type
→ metric + validation plan
→ shortlist compatible models
→ cross-validation baselines
→ optional HPO on the top 2
→ multi-criteria ranking
→ score the held-out test once
→ production-readiness checks
→ trained artifact
Always in the comparison: Logistic Regression or Ridge, Random Forest, Extra Trees, HistGradientBoosting.
If you installed extras: XGBoost, LightGBM, CatBoost.
Tiny data skips overkill boosters. Identifiers such as customer_id are flagged and dropped from features.
Reading the verdict
Do not print run.best() and run.leaderboard() for humans. Those are for code.
| Call | For |
|---|---|
print(run) |
The decision, in English |
run.report() |
Full markdown record |
run.best() |
Dict for scripts |
run.leaderboard() |
DataFrame for scripts |
run.artifact() |
Deployable pipeline |
If held-out accuracy is high but precision / recall / f1 are 0, the model never predicted the rare class at the 0.5 cutoff. Ranking metrics (pr_auc, roc_auc) can still look fine. MLVerdict prints an ATTENTION block for that trap. Production packaging checks (save/reload) are not a certificate that the business metric is good.
Tests
pip install -e ".[dev]"
pytest
Coverage includes problem detection, leakage/quality, candidate selection, end-to-end fit on synthetic and sklearn datasets (breast cancer, iris, wine, diabetes), display/report formatting, and CLI save.
Project layout
ml_lib/
├── src/mlverdict/ # library (pip installable)
│ ├── api/ # Verdict, Run, CLI
│ ├── data/ # load, profile, DNA
│ ├── problem/ # classification vs regression
│ ├── metrics/ # metric plan
│ ├── models/ # catalog + sklearn / booster adapters
│ ├── experiments/ # CV + bounded HPO
│ ├── evaluation/ # stability, latency, composite score
│ ├── decision/ # winner + written record
│ ├── reporting/ # print(run) + markdown report
│ └── production/ # artifact + local serve
├── tests/ # pytest
├── examples/ # runnable quickstart
├── pyproject.toml
├── LICENSE # MIT
└── README.md
Versioning PyPI vs Git
This repo is GitHub. PyPI is a separate upload.
| You do | GitHub | pip install mlverdict |
|---|---|---|
git push |
updates | no |
bump version + python -m build + twine upload |
no | yes |
Same version number cannot be uploaded twice. After this README, the published package is 0.1.1.
License
MIT. See LICENSE.
Metadata
Release files for mlverdict 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mlverdict-0.1.1.tar.gz | 62.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mlverdict-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 131.8 kB
Release files / mlverdict-0.1.1.tar.gz
| Download URL | mlverdict-0.1.1.tar.gz |
|---|---|
| Size | 62.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
618b24aa7ad7a109e6a68122615844c7f8eab4b521599d4458dc7dd17b6d5233
|
|
BLAKE2b-256 checksum How to use checksums |
fddeeb2938d21dd409844468ebc98ba7496ffe8fcab42a888135d913a979b770
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|
Release files / mlverdict-0.1.1-py3-none-any.whl
| Download URL | mlverdict-0.1.1-py3-none-any.whl |
|---|---|
| Size | 69.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d66c002cd360fa83545ec0980b3eb3c33e9e9e56c02047a53a5931f946361698
|
|
BLAKE2b-256 checksum How to use checksums |
e897eca458c4e554c2622b07eff9486742b103bcc5ca36174df5b7f9de6a9805
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|