plainml
Machine learning in plain English.
From a spreadsheet to a trained, explained, deployable model in one command, or one click.
Try it in your browser · Documentation · PyPI
Point plainml at a CSV (or Excel, Parquet, JSON, a URL or a database) and name the column you want to
predict. It works out whether that's classification or regression, checks the data for problems, fairly
compares up to 17 models (30 with --thorough, including an optional PyTorch network), explains the
winner in plain English, and saves it with an HTML report and a model card. The model is then ready to
make predictions, serve as an API, ship in a Docker image or export.
It also forecasts, finds groups and unusual rows, ranks columns with 19 importance methods, and watches
for drift. Prefer clicking? plainml web opens a local website: drag in a file, choose what to find out,
and download whichever results you need.
Contents
Why plainml · Install · Quickstart · The website · What it can do · Models · Reports and model cards · Commands · Examples · Deploying · Trust and privacy · Python API · How it works · How plainml compares · Development
Why plainml
- No code, no setup. Blank cells, prices stored as text (
"$1,200"), dates, free-text notes and ID columns are all handled automatically, and every saved model applies exactly the same preparation to new data. - Honest scores. Models are ranked by cross-validation, the winner is checked on rows it never saw, and everything is compared against a do-nothing baseline. You get warnings for likely data leaks, imbalanced classes, overfitting, poorly calibrated probabilities, and data that simply can't predict the target.
- Explained. "Higher support_calls → 'yes' more likely (24% → 38%)", which columns matter, and why any single prediction came out the way it did.
- The whole workflow. Profile → clean → train → tune → explain → predict → watch for drift → serve → deploy, plus forecasting, clustering, anomaly detection and feature importance.
- For every level. A website, a guided mode that asks questions, one-line commands, and a Python API.
- Private by default. Everything runs on your machine with no account, and
--privatekeeps raw data values out of reports and saved files.
Install
pip install plainml
Python 3.10 or newer. Optional extras add more:
| Extra | Adds |
|---|---|
boost |
XGBoost, LightGBM and CatBoost models |
tune |
Smarter hyperparameter search with Optuna |
explain |
SHAP explanations and SHAP importance |
imbalance |
SMOTE oversampling for rare classes |
web |
plainml web, the website |
serve |
plainml serve, a REST API (FastAPI) |
onnx |
plainml export to ONNX |
forecast |
Public holidays as forecasting features |
formats / fast |
Parquet and databases / faster loading of big files (polars) |
all |
Everything above |
torch |
A PyTorch neural network (not in all, since it's large) |
mlflow |
Logging runs to MLflow, and MLflow model export (not in all) |
cluster |
K-Medoids clustering (not in all: no prebuilt wheels for recent Pythons) |
pip install "plainml[all]" # or pick: pip install "plainml[web,boost]"
30-second quickstart
The repository ships with example datasets in examples/:
plainml train examples/churn.csv --target churned
That prints a leaderboard and a plain-English summary, and saves a run folder such as
runs/20260925-125240_churn/ with the model, report.html, model_card.md and more. Then predict on any
file with the same columns:
plainml predict latest examples/churn.csv -o predictions.csv
Not sure where to start? Run plainml on its own for guided mode, or plainml web for the website.
The website
pip install "plainml[web]"
plainml web
- Upload. Drag in a file and see every column's type and a preview of the rows.
- Pick a task. Predict a column, forecast, find groups, find unusual rows, rank columns, check drift, profile or clean. The main options are up front; the rest are under More options.
- Watch it run. A progress bar and a live log show each step.
- Get the results. Headline numbers, the key findings in plain English, and the full report.
- Download what you need. Every file the run produced is listed with its own Download button, and tables can be previewed first.
The Runs page lists everything you've run (from the website or the command line), and Predict uses
any saved model on a new upload. The site works in light and dark mode and on phones. It only listens on
your machine; to share it on a network, add a token: plainml web --host 0.0.0.0 --token change-me.
Try it now at plainml-tpua.vercel.app, with nothing to install.
No server needed: plainml web --export site writes the same website as static files that run
plainml inside the visitor's browser. Their data never leaves their computer, and the site can be hosted
for free on Vercel, GitHub Pages or a Hugging Face Static Space. See the
hosting guide and the
website guide.
What it can do
| Task | Command | What you get |
|---|---|---|
| Predict a column | plainml train |
Classification (yes/no or several classes), regression, multi-label and multi-output. Models compared, the winner explained and saved |
| Forecast | plainml forecast |
Backtested forecasts with an 80% range. One per store/product (--group), with planned inputs like promotions (--inputs) and public holidays (--country) |
| Find groups | plainml cluster |
K-Means, agglomerative, Gaussian mixtures and HDBSCAN compared; each group described by what sets it apart |
| Find unusual rows | plainml anomaly |
Isolation Forest, LOF, One-Class SVM and a robust distance combined into one score, with a reason for each flagged row |
| Rank the columns | plainml importance |
Up to 19 importance methods, a consensus ranking, how many columns you actually need, and redundant groups |
| Check for drift | plainml drift |
Which columns have shifted since training (PSI per column), weighted by how much the model relies on them |
| Understand the data | plainml profile / clean |
Column types, gaps and warnings; a cleaned copy |
| Improve a model | plainml tune |
Hyperparameter search (Optuna or random search) with a fair before/after comparison |
| Use a model | plainml predict / explain / evaluate |
Predictions (also for files too big for memory), explanations of single rows, scores on new labelled data |
| Ship it | plainml serve / deploy / export |
A REST API, a Docker image, or ONNX / MLflow files |
Models
plainml models lists every model, which tasks it handles, and whether it's installed.
| Tier | Classification | Regression |
|---|---|---|
| Every run | Logistic regression, ridge, KNN, naive Bayes, SVM, decision tree, random forest, extra trees, gradient boosting, histogram gradient boosting, MLP, LDA, XGBoost†, LightGBM†, CatBoost† | Linear, ridge, lasso, elastic net, KNN, SVM, decision tree, random forest, extra trees, gradient boosting, histogram gradient boosting, MLP, Bayesian ridge, Huber, XGBoost†, LightGBM†, CatBoost† |
With --thorough |
QDA, AdaBoost, bagging, SGD, linear SVM, Gaussian process, PyTorch network‡ | AdaBoost, bagging, SGD, linear SVM, Gaussian process, Theil-Sen, RANSAC, Poisson, Gamma, Tweedie, kernel ridge, PLS, PyTorch network‡ |
| Ensembles | The top 3 averaged (every run); stacked (with --thorough) |
Same |
† with the boost extra. ‡ with the torch extra.
- Every model is compared with a do-nothing baseline (the most common class, or the average).
--quickruns only the fast models;--models rf,lightgbmor--exclude svmchoose exactly.- Naming a
--thoroughmodel runs it without the others (--models torch). - Models that can't suit the data are skipped with the reason shown. Poisson regression needs a target that's never negative, for example, and slow models are skipped on large data.
The PyTorch network (pip install "plainml[torch]") is a multi-layer perceptron for tables:
- Blocks of Linear → BatchNorm → ReLU → Dropout, narrowing layer by layer.
- Trained with AdamW (weight decay) on mini-batches, with early stopping on a validation split, keeping the best weights.
- Classification uses cross-entropy (class-weighted for rare classes). Regression uses Huber loss on standardised targets.
- It behaves like any scikit-learn model, so cross-validation, ensembles, tuning (width, depth, dropout,
learning rate, weight decay), saving,
predict,serveanddeployall work with it. - It trains on the CPU by default. Set
PLAINML_TORCH_DEVICE=cudaormpsfor a GPU.
On spreadsheet-style data, gradient boosting usually wins; the network earns its place mostly inside ensembles.
Reports and model cards
Every run writes a self-contained report.html that works offline, in light and dark mode and on a phone.
For a trained model, it has:
- the model comparison
- test-set results: confusion matrix, ROC and precision-recall curves, and a calibration curve, or predicted-vs-actual for numbers
- what drives the predictions
- the data checks
- how to reproduce the run
Forecasts, clusters, anomalies, importance and drift get reports of their own.
Each trained model also gets a model_card.md: a one-page summary of what it predicts, the data it
learned from, how well it does (per class, too), what drives it, its caveats, and how to use it. It's
meant to travel with the model.
Commands
| Command | What it does | |
|---|---|---|
| Start here | plainml train DATA -t COLUMN |
Train and compare models; save the best with a report |
plainml profile DATA |
Column types, gaps and warnings, without training | |
plainml clean DATA -o clean.csv |
Fix blanks, text numbers, dates, duplicates, capitalisation… | |
| Use a model | plainml predict MODEL DATA |
Predictions on new rows (MODEL can be latest) |
plainml evaluate MODEL DATA |
Score on new labelled data and compare with training time | |
plainml explain MODEL [DATA] [--row N] |
What the model relies on, in plain English | |
plainml drift MODEL DATA |
Has new data drifted from what the model learned on? | |
plainml report RUN |
Rebuild and open a run's HTML report | |
| Improve | plainml tune [RUN] |
Hyperparameter search on the best models |
plainml importance DATA -t COLUMN |
Rank columns with up to 19 methods and a consensus | |
plainml select DATA -t COLUMN |
Keep only the columns that matter | |
| Other problems | plainml cluster DATA |
Find groups of similar rows, and describe them |
plainml anomaly DATA |
Find unusual rows and say why they're unusual | |
plainml forecast DATA -t COLUMN |
Forecast a value over time, with a range | |
| Deploy and share | plainml web |
The website: upload, run anything, download results |
plainml serve MODEL |
A REST API with docs at /docs |
|
plainml deploy MODEL |
A ready-to-build Docker image for the API | |
plainml export MODEL |
ONNX (for C#, Java, JavaScript, C++…) or MLflow | |
| Housekeeping | plainml runs / compare / models / init |
List, compare and prune runs; list models; write a config |
Every command has examples in plainml COMMAND --help, and the command reference lists
every option.
Examples
# Regression, ranked by mean absolute error, fast models only
plainml train examples/house_prices.csv -t price --metric mae --quick
# Choose the models, cap the time, ignore a column, use SMOTE for rare classes
plainml train examples/churn.csv -t churned --models rf,lightgbm,xgboost --time-budget 5m --drop region --balance smote
# Several yes/no targets at once (multi-label)
plainml train data.csv -t is_spam,is_urgent
# Every model, including the extras and a stacked ensemble, with honest probabilities
plainml train examples/churn.csv -t churned --thorough --calibrate
# Squeeze more accuracy out of the last run
plainml tune latest --trials 50
# Which columns matter? Random forest importance, RFE, Boruta, SHAP... and a consensus
plainml importance examples/churn.csv -t churned --methods all
# Why did row 3 get its prediction?
plainml explain latest new_customers.csv --row 3
# Customer segments, suspicious payments, next month's sales
plainml cluster examples/customers.csv --drop member_id
plainml anomaly examples/transactions.csv --label is_fraud
plainml forecast examples/daily_sales.csv -t units_sold --horizon 30
# One forecast per store, using planned promotions and public holidays
plainml forecast examples/store_sales.csv -t sales --group store --inputs promo --country US --horizon 14
# Has this month's data drifted from what the model learned on?
plainml drift latest this_month.csv
# Predict a file too big for memory, 200,000 rows at a time
plainml predict latest huge.csv -o predictions.csv --chunk-size 200000
# Tidy up: keep only the 10 newest runs
plainml runs --prune --keep 10
A walkthrough of every example dataset is in examples/README.md.
Deploying
A REST API. plainml serve MODEL starts a FastAPI server:
| Endpoint | |
|---|---|
POST /predict |
Send rows as JSON ({"rows": [{...}, ...]}), get predictions and probabilities back |
POST /forecast |
For forecast models: {"horizon": 30}, optionally "groups": [...] |
GET / |
What the model is, the columns it expects, and an example request |
GET /health |
A health check |
GET /docs |
Interactive documentation where you can try requests |
Add --api-key SECRET (or set PLAINML_API_KEY) and every request except /health and /docs needs the
header X-API-Key: SECRET.
A Docker image. plainml deploy MODEL writes a folder you can build and run anywhere containers run
(Cloud Run, App Runner, Azure Container Apps, Fly.io, Kubernetes…):
plainml deploy latest -o deploy/churn
docker build -t churn deploy/churn
docker run -p 8000:8000 -e PLAINML_API_KEY=change-me churn
The folder holds:
- the model and its model card
- a slim Dockerfile that runs as a non-root user, with a health check
requirements.txt, pinned to the exact library versions the model was trained with (only the ones it needs)- a README with the commands to build, run and call it
If you installed plainml from a checkout rather than PyPI, the folder also bundles a plainml wheel, so the build doesn't need PyPI.
Files. plainml export MODEL writes ONNX (checked against the original model), for C#, Java,
JavaScript or C++. --format mlflow saves an MLflow model, and plainml train --mlflow logs runs to your
MLflow tracking server.
Hosting. The hosting guide covers putting the website online:
- In the browser, free: on Vercel, next to the docs (
vercel.jsonbuilds both), or on a Hugging Face Static Space. - On a server: on any container host, using the Dockerfile in
hosting/docker/.
Trust and privacy
- Calibration. The report shows whether "70% sure" really means right 70% of the time (a reliability
curve, the Brier score and the average gap).
--calibratefixes probabilities that are off. - Drift. A saved model remembers what its training data looked like.
plainml predictwarns when new rows look clearly different, andplainml driftshows which columns moved and how much that matters. - Private runs. With
--private(ontrain,cluster,anomalyandimportance), the report, run folder and model keep only file names, column names and summary numbers: no example rows, raw values or data paths. - Safe loading. Saved models record their library versions, and plainml explains how to fix a mismatch. Only load model files you trust: loading a joblib/pickle file can run code.
Python API
Everything the command line does is available from Python and notebooks:
import plainml
result = plainml.train("examples/churn.csv", target="churned", quick=True)
print(result.best_model, result.cv_score)
result.leaderboard # a DataFrame (renders as a table in Jupyter)
predictions = plainml.predict(result.run_dir, "new_customers.csv", proba=True)
why = plainml.explain(result.run_dir)
print(why.sentences)
groups = plainml.cluster("examples/customers.csv", drop=["member_id"])
forecast = plainml.forecast("examples/daily_sales.csv", "units_sold", horizon=30)
ranked = plainml.feature_importance("examples/churn.csv", "churned", methods=["rf", "rfe", "boruta"])
drift = plainml.check_drift(result.run_dir, "this_month.csv")
Saved models are ordinary scikit-learn objects that accept raw rows (strings, blanks and all):
import joblib, pandas as pd
model = joblib.load("runs/20260925-125240_churn/model.joblib")
model.predict(pd.read_csv("new_customers.csv"))
See the Python API guide.
Reproducible runs
Each run saves its settings, library versions and a fingerprint of the data. Rerun it exactly with:
plainml train --config runs/20260925-125240_churn/config.yaml
or start your own commented config with plainml init.
How it works
-
Profile. Each column is classified as number, category, date, free text, ID or constant. IDs and constants are set aside, and you get warnings about leaks, imbalance, heavy blanks and tiny datasets.
-
Split. 20% of rows are held out as a final test set that no model sees while models are chosen.
-
Compare. Every model runs inside the same pipeline:
- fill blanks, and flag which cells were blank
- scale numbers
- one-hot encode categories
- TF-IDF word features for text
- date parts
Each is scored with 5-fold cross-validation next to the baseline. Rare classes get weights (or SMOTE), and imbalanced yes/no problems also get a tuned decision threshold.
-
Check and explain. The winner is scored on the held-out rows, checked for calibration and overfitting, and explained with permutation importance and effect curves.
-
Save. The winner is retrained on all rows and saved with:
- its preprocessing
- the report and model card
- a summary of the training data, for drift checks
More detail, including forecasting, importance and drift, is in How it works.
How plainml compares
plainml isn't the only way to get a model without writing much code. It's built for people who want to understand a result and trust it, not squeeze out the last 0.5% of accuracy. It's free, runs on your machine, needs no account, and says what it found in plain English.
| Use it from | Runs where | Main strength | Where plainml differs | |
|---|---|---|---|---|
| plainml | Command line, guided mode, website, Python | Your machine | Explained, honest results for the whole workflow | |
| PyCaret | Python / notebooks | Your machine | Low-code workflow with many models and tasks | plainml adds a CLI, a website and plain-English findings, and needs no code |
| AutoGluon | Python | Your machine | Top accuracy through heavy stacking | plainml is lighter and faster, and explains more; AutoGluon usually scores higher |
| MLJAR AutoML | Python | Your machine | Markdown reports for every model | plainml also covers forecasting, drift, clustering and deployment, with a website |
| H2O AutoML | Python, R, web UI (Flow) | Your machine or a cluster | Scales to very large data | plainml is a single pip install with no Java, aimed at newcomers |
| Orange, KNIME | Desktop app | Your machine | Visual drag-and-drop workflows | plainml chooses the steps for you instead of you wiring them together |
| DataRobot, Dataiku, cloud AutoML (Vertex AI, SageMaker Canvas, Azure) | Web platform | Their servers or your company's | Enterprise features: governance, monitoring, teams | plainml is free and local; your data doesn't leave your machine |
Single-purpose tools also do parts of this well: ydata-profiling for profiling, Evidently for drift, Prophet and StatsForecast for forecasting, and mlxtend and BorutaPy for feature selection.
When to reach for something else: AutoGluon when accuracy is all that matters; H2O for data that doesn't fit on one machine; PyCaret if you live in notebooks and want to customise every step; an enterprise platform when a team needs approvals, monitoring and access control.
Development
git clone https://github.com/pranay-obla/plainml && cd plainml
pip install -e ".[all,dev]"
pre-commit install
pytest
Where things live
| Path | What's there |
|---|---|
plainml/cli.py, wizard.py |
The command line and guided mode |
plainml/training.py, registry.py, deep.py |
Training, the model table, the PyTorch network |
plainml/preprocessing.py, schema.py, tasks.py, metrics.py |
Column types, preparation, task detection, metrics |
plainml/profiling.py, cleaning.py, io.py |
Data checks, cleaning, reading and writing files |
plainml/evaluation.py, explaining.py, tuning.py |
Test-set scoring and calibration, explanations, tuning |
plainml/forecasting.py, clustering.py, anomaly.py |
The other kinds of problem |
plainml/importance.py, drift.py |
Feature importance and drift |
plainml/predicting.py, serve.py, deploy.py, export.py, tracking.py |
Predicting, the API, Docker, ONNX, MLflow |
plainml/report.py, card.py, runs.py |
HTML reports, model cards, run folders |
plainml/web/ |
The website: API server, job runner, and the page itself (static/) |
tests/, docs/, examples/ |
Tests, the documentation site, example datasets |
hosting/, vercel.json, .github/workflows/ |
Hosting the website and docs, CI and releases |
CI runs the tests on every Python from 3.10 to 3.14 on Linux, macOS and Windows. It also tests with every optional extra (PyTorch included), against the oldest supported library versions, and weekly against the newest releases, so a library update can't silently break plainml.
Publishing a GitHub release sends the package to PyPI (see
RELEASING.md), and the documentation in
docs/ can be hosted on Vercel or GitHub Pages.
Author
plainml grew out of curdrice, a hackathon project by @hrishitb, @pranayobla, @shriharik and @ankitthomas.
Acknowledgements
Built on scikit-learn, pandas, NumPy, Rich, Click, questionary and joblib, with optional XGBoost, LightGBM, CatBoost, PyTorch, Optuna, SHAP, imbalanced-learn, FastAPI, skl2onnx, holidays and MLflow.
Release files for plainml 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| plainml-0.1.1.tar.gz | 262.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| plainml-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 513.4 kB
Release files / plainml-0.1.1.tar.gz
| Download URL | plainml-0.1.1.tar.gz |
|---|---|
| Size | 262.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2debfb13a207994b3648bcb8c14e2dfeda1348578e09152573a4cff360b434e3
|
|
BLAKE2b-256 checksum How to use checksums |
8d70c9e67121da6db53707ee6a43486486f8503afb320a0ee11caf4e33b3d4f7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency logRelease files / plainml-0.1.1-py3-none-any.whl
| Download URL | plainml-0.1.1-py3-none-any.whl |
|---|---|
| Size | 250.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a0921d59064df57aa6a4f9444e57ad99ced766b43f0921b8cbaaf62502c0fda7
|
|
BLAKE2b-256 checksum How to use checksums |
8e76c4e45c75b6b917517009ca9ec305642e09054ccdb57deb3715a4194b13ea
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency log