Production-level automated ML pipeline — from raw data to deployed model in one line
Project description
open-mlpipe
Production-level automated ML pipeline — from raw data to deployed model in one line.
What is open-mlpipe?
open-mlpipe is an AutoML pipeline that takes your raw dataset and automatically:
- Loads and profiles your data (EDA)
- Cleans duplicates, outliers, and missing values
- Engineers features (interactions, log transforms, missingness flags)
- Compares 14+ ML models head-to-head
- Tunes hyperparameters with Optuna
- Selects features via SHAP importance
- Explains predictions with SHAP
- Saves a production-ready inference pipeline
One command. One line of code. Production model.
Quick Start
Install
pip install open-mlpipe
CLI Usage
# Zero-touch — auto-detect everything
open-mlpipe run --data dataset.csv
# With target specified
open-mlpipe run --data dataset.csv --target price
# Config-driven
open-mlpipe run --config configs/regression.yaml
Python API
from open_mlpipe import run
# One line — full pipeline
ctx = run("dataset.csv", target="price")
# Access results
print(f"Best model: {ctx.best_model_name}")
print(f"Test R²: {ctx.metrics['test_r2']:.4f}")
# Load saved model for predictions
import joblib
model = joblib.load("artifacts/model_v1.joblib")
predictions = model.predict(new_data)
YAML Config
# configs/my_pipeline.yaml
project: my-project
task: auto
data:
path: data.csv
target: price
model_selection:
candidates: [lightgbm, xgboost, random_forest, ridge]
tuning:
enabled: true
engine: optuna
n_trials: 50
open-mlpipe run --config configs/my_pipeline.yaml
Features
| Feature | Description |
|---|---|
| 14+ Models | Ridge, Lasso, ElasticNet, DecisionTree, RandomForest, ExtraTrees, XGBoost, LightGBM, GradientBoosting, HistGradientBoosting, AdaBoost, KNN, SVM, Stacking, Voting |
| Auto EDA | Column hygiene, quality assessment, distributions, statistical tests, VIF |
| Feature Engineering | Missingness flags, interaction features, log transforms, datetime extraction |
| Hyperparameter Tuning | Optuna with TPE sampler, 18 model-specific search spaces |
| SHAP Explainability | Global feature importance, local explanations |
| Production Inference | Full pipeline saved (feature engineering + model) — raw data → prediction |
| Overfitting Detection | Automatic train/test gap analysis |
| Production Metrics | R², RMSE, MAE, MAPE, F1, ROC-AUC, MCC |
Supported Tasks
- Regression — predict continuous values
- Classification — predict categories (binary/multiclass)
Installation Options
# Core (minimal)
pip install open-mlpipe
# With CatBoost
pip install open-mlpipe[catboost]
# With MLflow tracking
pip install open-mlpipe[mlflow]
# With deployment (FastAPI + Docker)
pip install open-mlpipe[deploy]
# Everything
pip install open-mlpipe[full]
Example Output
=== mlpipe Pipeline Runner ===
Project: california-housing
Data: california_housing.csv
Target: median_house_value
>> load OK (0.1s)
>> eda OK (0.2s)
>> clean OK (0.0s)
>> feature_eng OK (0.0s)
>> split OK (0.0s)
>> preprocess OK (0.0s)
>> compare OK (25.9s) — tested 14 models
>> tune OK (20.2s) — Optuna 20 trials
>> select OK (0.0s)
>> evaluate OK (0.0s)
>> explain OK (3.3s)
>> save OK (0.1s)
=== Pipeline Complete ===
Best Model: lightgbm
Test R²: 0.8474
Test MAE: 0.2924
Test MAPE: 16.77%
Model Comparison
open-mlpipe automatically compares all models and ranks them:
| Model | R² | Gap | Status |
|---|---|---|---|
| LightGBM | 0.839 | 0.072 | Best |
| XGBoost | 0.835 | 0.093 | Good |
| RandomForest | 0.771 | 0.198 | Overfit |
| Ridge | 0.709 | 0.001 | Stable |
Project Structure
open-mlpipe/
├── src/open_mlpipe/
│ ├── cli.py # CLI entry point
│ ├── config/ # Pydantic config + YAML resolver
│ ├── core/ # Pipeline runner, context, stages
│ ├── stages/ # 12 pipeline stages
│ ├── utils/ # I/O, feature engineering
│ └── deploy/ # FastAPI + Docker generation
├── tests/ # 106 unit tests + 3 integration tests
├── configs/ # 5 example YAML configs
└── pyproject.toml # Package config
Requirements
- Python >= 3.10
- See
pyproject.tomlfor full dependency list
License
MIT License — see LICENSE for details.
Contributing
Contributions welcome! Please open an issue or PR.
Author
Yash Brahmankar
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file open_mlpipe-1.0.0.tar.gz.
File metadata
- Download URL: open_mlpipe-1.0.0.tar.gz
- Upload date:
- Size: 49.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3bb18f68a6c717e6ebeb1af753ddc13d5b29577fe696d9592514959948bf11b6
|
|
| MD5 |
519ff36587c011b4359131f6e8b7c68b
|
|
| BLAKE2b-256 |
999280e9f3c32df9862101f0be2868ed9ec7a9c5eced6fa3bb37bb4d0d543e62
|
File details
Details for the file open_mlpipe-1.0.0-py3-none-any.whl.
File metadata
- Download URL: open_mlpipe-1.0.0-py3-none-any.whl
- Upload date:
- Size: 41.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
322cbb9f1e86f8b8b2033df76d3f91a75b11d41ff3d48be2e71e94b917a3a4e3
|
|
| MD5 |
777b11d86fe2304035cb68d896fb535b
|
|
| BLAKE2b-256 |
b36bb60dde1936815a9f2023dd7fd5d0065e5adbc6988b5782533df957c1f1f2
|