MLPrimer
MLPrimer is an all-in-one Python package for rapid machine learning experimentation and deployment.
It automates the full ML workflow—from exploratory data analysis to model selection and hyperparameter tuning—behind a consistent, task-aware API.
Overview
MLPrimer orchestrates the complete ML pipeline with a single function call.
from mlprimer.bootstrap import run_pipeline
result = run_pipeline(df, target_var="target")
best_model_name, best_model = result.get_best_tuned_model()
predictions = best_model.predict(X_new)
Pipeline Stages
- Statistics – Feature profiling, normality testing, distribution analysis
- Correlation – Association testing (Pearson, Spearman, Chi-Square, ANOVA, t-tests)
- Preprocessing – Task-aware splitting, class balancing (SMOTE), target transforms
- Model Selection – 25+ models trained and ranked automatically
- Tuning – Hyperparameter optimization for the top-N models
Key Features
Data Analysis
- Comprehensive statistics: mean, median, std, skewness, normality tests
- Association testing: Pearson, Spearman, Chi-Square, Point-Biserial, ANOVA
- Automatic target type inference (binary, multiclass, regression, count)
Preprocessing
- Task-aware splitting (stratified, grouped, time-series)
- Class balancing: SMOTE, under-sampling, class weights
- Target transformations: Yeo-Johnson, Box-Cox, log1p
- Sample weighting for imbalanced or heteroscedastic targets
Model Selection (25+ Models)
- Classification: Logistic Regression, Random Forest, SVC, XGBoost, LightGBM, CatBoost, MLP
- Regression: ElasticNet, Ridge, Lasso, SVR, Random Forest, XGBoost, LightGBM, CatBoost
- Robust Models: Huber, RANSAC, Theil-Sen
- Neural Networks: MLP Classifier / Regressor
Hyperparameter Tuning
- Grid Search, Randomized Search, Successive Halving, Bayesian Optimization
- Task-aware cross-validation
- Multi-metric tracking
- Predefined parameter grids for supported models
Export & Deployment
- Persist best or top-N models with metadata
- Auto-generated inference and tuning scripts
- Export feature importance, associations, and rankings
Installation
From PyPI (recommended)
pip install mlprimer
From Source
git clone https://github.com/emrec/mlprimer.git
cd mlprimer
pip install -e .
Optional Dependencies
# Bayesian optimization
pip install mlprimer[bayesian]
# Boosting models (XGBoost, LightGBM, CatBoost)
pip install mlprimer[advanced]
# All optional features
pip install mlprimer[all]
Quick Start
Complete Pipeline
import pandas as pd
from mlprimer.bootstrap import run_pipeline
df = pd.read_csv("data.csv")
result = run_pipeline(
df,
target_var="target",
tune_top_n=5,
verbose=2
)
best_model_name, best_model = result.get_best_tuned_model()
predictions = best_model.predict(X_test)
Step-by-Step Workflow
from mlprimer.statistics import calculate_summary
from mlprimer.correlation import check_association
from mlprimer.preprocessing import infer_target_type, split_data, apply_smote_classification
from mlprimer.model_selection import run_full_model_selection
from mlprimer.tuning import tune_random
feature_summary = calculate_summary(df)
associations = check_association(df, target_var="target")
target_type = infer_target_type(df["target"])
X_train, X_test, y_train, y_test = split_data(
df.drop("target", axis=1),
df["target"],
task="classification"
)
X_train, y_train = apply_smote_classification(X_train, y_train)
results = run_full_model_selection(
X_train, y_train, X_test, y_test, target_type=target_type
)
best_model = results["best_model"]
tuned = tune_random(
best_model,
X_train,
y_train,
{"n_estimators": [100, 200], "max_depth": [10, 20]},
task="binary"
)
Exporting Models
from mlprimer.bootstrap import (
export_best_model,
export_models,
create_inference_script
)
export_best_model(result, "best_model.pkl")
export_models(result, "models/", include_top_n=5)
create_inference_script(result, "inference.py")
Core Modules
statistics.py
calculate_summary(df)check_normal(x)
correlation.py
check_association(df, target_var)analyze_associations(df, target_var)
preprocessing.py
infer_target_type(y)export_target_distribution(y, task)split_data(X, y, task, stratify=None, groups=None, time_series=False)apply_smote_classification(X, y)apply_target_transform(y, method)make_sample_weights_regression(y)
model_selection.py
run_full_model_selection(...)suggest_models(target_type)train_models(X_train, y_train, task)evaluate_models(models, X_test, y_test, task)
tuning.py
tune_gridtune_randomtune_halving_gridtune_halving_randomtune_bayesiantune_with_calibration
bootstrap.py
run_pipelineexport_resultsexport_best_modelexport_modelscreate_fine_tuning_scriptcreate_inference_scriptcreate_comparison_script
Supported Tasks
| Task | Target Type | Models | Metrics |
|---|---|---|---|
| Binary Classification | 2 classes | 15+ | F1, AUC, Precision, Recall |
| Multiclass | 3+ classes | 15+ | F1-weighted, Balanced Accuracy |
| Regression | Continuous | 14+ | MAE, RMSE, R² |
| Count | Non-negative ints | 14+ | MAE, RMSE |
Architecture
bootstrap.py
├─ statistics.py
├─ correlation.py
├─ preprocessing.py
├─ model_selection.py
└─ tuning.py
Requirements
- Python ≥ 3.8
- pandas ≥ 1.0
- numpy ≥ 1.18
- scikit-learn ≥ 0.24
- scipy ≥ 1.5
- imbalanced-learn ≥ 0.8
Optional
- xgboost
- lightgbm
- catboost
- scikit-optimize
License
MIT License
Author
Emre Can Konca
Acknowledgments
Built on top of scikit-learn, XGBoost, LightGBM, CatBoost, and the broader open-source ML ecosystem.
Release files for mlprimer 1.0.8
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mlprimer-1.0.8.tar.gz | 40.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mlprimer-1.0.8-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 78.3 kB
Release files / mlprimer-1.0.8.tar.gz
| Download URL | mlprimer-1.0.8.tar.gz |
|---|---|
| Size | 40.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cbff716f69393790d883aaf8ef814ec5b566ef63929218c545de89b9199179bb
|
|
BLAKE2b-256 checksum How to use checksums |
8ae6972d6c6c2f01eeb95915b587e6267f972c740c8b923274696e6afd36bcca
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.2
|
Release files / mlprimer-1.0.8-py3-none-any.whl
| Download URL | mlprimer-1.0.8-py3-none-any.whl |
|---|---|
| Size | 38.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
882766e820375f5b2e4e1d61ecef1bd53ff274374cfe084a192999148bfd587d
|
|
BLAKE2b-256 checksum How to use checksums |
cddfb49c3d346bfa1b3add6655f8c67690cfe199ccdf2f27a86b74ab551d5d8e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.2
|