🚀 MLPilot
Data In → Insights Out
A production-grade Python Machine Learning Library for tabular datasets that automatically performs EDA, preprocessing, model comparison, hyperparameter tuning, explainability, and exports a deployment-ready inference pipeline.
14+ Models · Auto EDA · Optuna · SHAP · CLI · Python API
One command. One pipeline. Production-ready models.
⚡ Quick Start
Install MLPilot:
pip install mlpilotx
Run your first ML pipeline:
mlpilot run --data examples/USA_Housing.csv
Specify the target column manually:
mlpilot run --data examples/USA_Housing.csv --target price
🔥 What MLPilot Does
Instead of writing hundreds of lines of boilerplate code, MLPilot automatically performs:
| Stage | Description |
|---|---|
| 📂 Load | CSV loading & validation |
| 📊 EDA | Statistics, missing values & correlations |
| 🧹 Clean | Duplicates, sparse columns & outlier handling |
| ⚙️ Feature Engineering | Encoding, scaling & transformations |
| ✂️ Split | Train / Test split |
| 🔧 Preprocess | Scikit-learn preprocessing pipeline |
| 🏆 Compare | 14+ Machine Learning models |
| 🎯 Tune | Bayesian optimization with Optuna |
| 🔍 Explain | SHAP feature importance & visualizations |
| 📦 Export | Deployment-ready .joblib pipeline |
🔄 Pipeline Workflow
Raw Dataset
│
▼
📂 Load Dataset
│
▼
📊 Statistical EDA
│
▼
🧹 Data Cleaning
│
▼
⚙️ Feature Engineering
│
▼
✂️ Train / Test Split
│
▼
🔧 Preprocessing Pipeline
│
▼
🏆 Compare 14+ ML Models
│
▼
🎯 Optuna Hyperparameter Tuning
│
▼
🔍 SHAP Explainability Report
│
▼
📈 Model Evaluation
│
▼
📦 Export Production Pipeline (.joblib)
✨ Features
- 📂 Automatic CSV support
- 🤖 Automatic Regression & Classification detection
- 📊 Smart Exploratory Data Analysis
- 🧹 Missing value & outlier handling
- ⚙️ Feature engineering pipeline
- 🏆 Cross-validation model leaderboard
- 🎯 Bayesian hyperparameter optimization
- 🔍 SHAP explainability visualizations
- 📦 Export complete inference pipeline
- 💻 Rich CLI & Python API support
💻 CLI Usage
Automatic Training
mlpilot run --data examples/heart.csv
Regression
mlpilot run --data examples/USA_Housing.csv --target price
Classification
mlpilot run --data examples/patient_adherence_dataset.csv --target adherence
Custom Output Folder
mlpilot run --data data.csv --output outputs/
🐍 Python API
from ml_pilot import PipelineRunner
from ml_pilot.config import load_config
config = load_config()
runner = PipelineRunner(config)
context = runner.run(
data_path="examples/USA_Housing.csv",
target="price"
)
print(context.best_model_name)
print(context.metrics)
📁 Example Datasets
MLPilot includes ready-to-use datasets inside the examples/ folder.
| Dataset | Task |
|---|---|
USA_Housing.csv |
Regression |
insurance.csv |
Regression |
heart.csv |
Classification |
patient_adherence_dataset.csv |
Classification |
Student_performance_data.csv |
Classification |
Food_Delivery_Times.csv |
Regression |
Exam_Score_Prediction.csv |
Regression |
taxi_trip_pricing.csv |
Regression |
personality_synthetic_dataset.csv |
Classification |
Example:
mlpilot run --data examples/insurance.csv --target charges
📦 Generated Artifacts
Every successful run generates:
mlpilot_artifacts/
├── mlpilot_pipeline.joblib
├── leaderboard.csv
├── metrics.json
├── model_comparison.json
├── feature_importances.json
├── eda_report.json
├── run_metadata.json
├── serving_schema.json
├── predict_snippet.py
├── DEPLOY.md
└── shap/
├── shap_summary.png
├── shap_dependence_*.png
└── shap_waterfall.png
🧠 Supported Models
| Category | Models |
|---|---|
| Linear | Linear Regression, Ridge, Lasso, ElasticNet |
| Tree | Decision Tree, Random Forest, Extra Trees |
| Boosting | Gradient Boosting, HistGradientBoosting |
| Instance | KNN, SVR |
| Neural | MLP |
| Classification | Logistic Regression, SGD, Linear SVC, Passive Aggressive |
🧪 Edge Case Testing
MLPilot includes dedicated validation datasets.
tests/
└── edge_cases/
├── empty.csv
├── one_row.csv
├── all_null.csv
├── duplicate_col.csv
├── target_missing.csv
├── only_numeric.csv
└── only_categorical.csv
Run an edge-case test:
mlpilot run --data tests/edge_cases/empty.csv
📂 Project Structure
MLPilot/
├── configs/
│ └── default.yaml
├── examples/
│ ├── USA_Housing.csv
│ ├── insurance.csv
│ ├── heart.csv
│ └── ...
├── src/
│ └── ml_pilot/
│ ├── cli.py
│ ├── config/
│ ├── core/
│ ├── stages/
│ └── utils/
├── tests/
│ ├── edge_cases/
│ ├── test_load.py
│ ├── test_pipeline_smoke.py
│ └── ...
├── LICENSE
├── README.md
└── pyproject.toml
🛠 Tech Stack
- Python 3.12+
- Scikit-learn
- Pandas & NumPy
- Optuna
- SHAP
- Typer + Rich
- Joblib
- Plotly
📄 License
This project is licensed under the MIT License.
Metadata
Release files for mlpilotx 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mlpilotx-0.1.1.tar.gz | 5.9 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mlpilotx-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.9 MB
Release files / mlpilotx-0.1.1.tar.gz
| Download URL | mlpilotx-0.1.1.tar.gz |
|---|---|
| Size | 5.9 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3b408fa817810163da6f83b74140e1c32d6ff3440357ecf9c4383669ab572e60
|
|
BLAKE2b-256 checksum How to use checksums |
51948c38e7c1cf91b78e686f2edc596e12be1f461257e14898f4a4d7cb55c38d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|
Release files / mlpilotx-0.1.1-py3-none-any.whl
| Download URL | mlpilotx-0.1.1-py3-none-any.whl |
|---|---|
| Size | 45.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e4ce8775402d3bc53918221b2eafaee180acc47448a43c09281715e306fbe72c
|
|
BLAKE2b-256 checksum How to use checksums |
e6239a1d1f38bd15bf1c543102f146557c5c1f1554f36969ce4bb31595690e20
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|