Glass-box autonomous data science: profile, clean, model, and explain every decision.
Project description
Mudra-ML
Point it at a table of data. It cleans the data, trains several models, picks the best one, and writes a report that explains every choice it made.
What it does
Mudra-ML automates the routine part of supervised machine learning on tables: reading the file, working out what each column holds, repairing messy values, splitting the data safely, choosing which algorithms to try, tuning them, and measuring the winner on data it never saw during training. Every one of those choices follows a written rule, and every rule that fired is logged and printed in the report. Nothing is decided by a hidden model, and the library never calls the network.
Install
pip install mudra-ml
Optional extras:
pip install mudra-ml[files] # parquet and excel readers
pip install mudra-ml[boost] # xgboost, lightgbm, and catboost candidates
The base install runs fully on scikit-learn. The boosting libraries join the candidate list only when they are installed. A missing one is skipped with a note in the log, never an error.
Quickstart
from mudra_ml import Mudra
m = Mudra()
result = m.run("data.csv")
print(result.task) # what it decided to do
print(result.metrics) # how well the chosen model did
print(result.report_path) # the full explanation, on disk
That is the whole loop. run accepts a file path (csv, tsv, excel, json, or parquet) or a pandas DataFrame. If you do not name a target column, it looks for a plausible one and tells you what it picked. The report lands next to your script as markdown and HTML.
Classification
Classification predicts a label, such as whether a customer leaves or stays. Name the column you want predicted:
result = Mudra().run("churn.csv", target="churned")
print(result.metrics["f1"])
print(result.positive_label) # the class treated as the event of interest
For a two-class target the minority class is treated as the positive class, because the rarer outcome (the churn, the fraud, the diagnosis) is usually the one you care about. The report records this choice.
Regression
Regression predicts a number, such as a price:
result = Mudra().run("houses.csv", target="price")
print(result.metrics["rmse"]) # typical size of the prediction error
print(result.metrics["r2"]) # share of variation the model explains
Clustering
Clustering groups similar rows when there is nothing to predict:
result = Mudra().run("customers.csv", task="clustering")
labels = result.predict(my_dataframe) # cluster id per row
Steering the run
You can set as much or as little as you want. Anything you do not set is inferred and the report says so.
result = Mudra(random_state=7).run(
"churn.csv",
target="churned",
task="classification",
metric="f1",
constraints={"interpretable": True, "max_train_seconds": 120},
)
With interpretable set, only models a person can read directly are trained, such as logistic regression and a small decision tree.
What data it accepts
Tables. One row per example, one column per attribute, with a header row. Numeric, categorical, boolean, date, text, and id columns are all recognized and handled. Messy values that real exports produce are repaired on the way in:
- Numbers written as text, with thousands separators, currency symbols, or percent signs, are parsed back to numbers.
- Missing-value spellings that pandas does not catch (
--,?, the wordmissing) are treated as missing. Real categories such asUnknownorOtherare left alone. - Booleans written as
yes/no,true/false,t/f, or 0/1 all work. - Dates are expanded into year, month, day, and weekday features.
- Id columns carry no signal and are dropped.
Data the library cannot handle stops the run with a MudraError that names the column, states the problem, and suggests a fix. You never see a raw pandas or scikit-learn traceback.
What happens at each step
- Ingest. The file is read. Delimiter, encoding, and header row are detected.
- Profile. Each column gets a type (numeric, categorical, boolean, date, text, id) by rule.
- Goal. Target, task, and metric are taken from you or inferred, in that order of priority.
- Quality. Constant columns, duplicate rows, class imbalance, and features that look leaked from the target are flagged.
- Preprocess. Imputation, outlier clipping, encoding, and scaling are fitted on the training split only, so nothing from the test split can leak into the model.
- Recommend. A documented rule set picks a small shortlist of algorithms that fit the task, the dataset size, the dimensionality, and the sparsity. It never runs every model on every dataset.
- Tune and select. Each shortlisted model is tuned with a fixed-seed randomized search under an explicit compute budget. The best is chosen by cross-validation alone. The held-out test set is scored once, for the winner, for reporting only. When the shortlist is substantial and the budget allows, a nested cross-validation pass estimates how the whole selection procedure generalizes.
- Report. Markdown and HTML, opening with a plain-language trust summary and listing every decision and the rule that produced it.
The numbers that drive these steps are not assumed. The test split scales with the row count, the fold count scales with the training rows and the smallest class, the outlier fences respond to each column's own skewness, the encoding boundaries scale with the data, and the report records every derived value with the rule and the inputs behind it. The values that stay fixed are listed in the report with their reasons. The search itself runs inside a wall-clock budget (Mudra(time_budget_seconds=...)) enforced through a deterministic cost estimate; anything trimmed to fit is logged and reported, never silent.
Models it can train
Classification: logistic regression, decision tree, random forest, extra trees, gradient boosting, support vector classifier, k-nearest neighbors, and gaussian naive bayes. Regression: linear regression, ridge, elastic net, decision tree, random forest, extra trees, gradient boosting, support vector regressor, and k-nearest neighbors. Clustering: k-means with a swept cluster count. XGBoost, LightGBM, and CatBoost join the shortlist when the boost extra is installed.
Which of these actually run depends on your data. For example, k-nearest neighbors is only tried when the feature count is modest and the data is dense, and kernel models are only tried when the row count keeps their cost reasonable. The report states which models were shortlisted and why.
Save, load, and predict on new data
result = Mudra().run("churn.csv", target="churned")
path = result.save("churn_model") # writes churn_model.joblib + churn_model.json
loaded = Mudra.load("churn_model") # restores the exact model
preds = loaded.predict(new_dataframe) # labels
probs = loaded.predict_proba(new_dataframe) # class probabilities
A saved then loaded model gives predictions identical to the original. The .json file next to the model records the library version, python version, creation date, task, target, metric, selected model, seed, positive class, and the input schema.
New data is checked against that schema before prediction. A missing column, an unexpected column, a column whose type changed, or a category the model never saw all raise a clear MudraError instead of returning silently wrong numbers.
The result object has a stable surface you can build on: best_model, pipeline, metrics, report_path, task, target, feature_names, positive_label, and model_path.
The decision report
The report is the product as much as the model is. Open result.report_path (markdown) or the .html file next to it. It opens with a trust summary: what the pipeline did, the headline metric against the baseline, the single biggest risk to trusting the result, and a short verdict that downgrades itself on a leakage suspect, a small test set, a weak lift, or an optimistic-looking selection. After that it contains the goal and which parts of it were inferred, the full decision log by stage, the data-quality findings, every candidate model with its cross-validation score, the compute budget and anything trimmed to fit it, the winner's held-out metrics next to a naive baseline, the train-versus-test gap and the validation guard, per-class breakdowns and curves for classification, residual diagnostics for regression, feature importance with its uncertainty, and the justified defaults with their reasons.
Determinism
One seed is threaded through the split, the search, the estimators, and any sampling, and every data-derived number is a pure function of the data and the seed. Two runs on the same input produce the same derived numbers, the same model, the same metrics, and reports identical apart from one clearly labeled line: the measured wall-clock time in the compute budget section. Budget reductions are planned from a cost estimate, not the clock, so they repeat exactly too. Set the seed with Mudra(random_state=...).
Command line
mudra-ml run data.csv --target churn --metric f1 --save churn_model
mudra-ml profile data.csv
run trains and writes the report. profile prints the inferred column types, missingness, cardinality, and candidate targets without training anything.
Scope
Mudra-ML is built for tabular data that fits in memory: the everyday spreadsheet, database extract, or csv export with up to a few hundred thousand rows. It covers binary and multiclass classification, regression, and k-means clustering, end to end, with an audit trail.
Limits
Know what it does not do, and when not to use it:
- No deep learning. Images, audio, and video are out of scope.
- Text columns are reduced to simple length features. For real language understanding, use a dedicated text pipeline.
- No time-series forecasting. The random split assumes rows are exchangeable, which time series are not.
- Data larger than memory is not supported.
- Do not use the output for decisions that affect people (credit, hiring, medical, legal) without a person reviewing the report, the data quality findings, and the limits of the data. The report flags many problems, and it cannot flag them all.
License
MIT. See LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mudra_ml-0.5.1.tar.gz.
File metadata
- Download URL: mudra_ml-0.5.1.tar.gz
- Upload date:
- Size: 111.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
44c676153ed1b755837f7b70321102fa3c5847c00782e8aa495f43d1fbe6dd20
|
|
| MD5 |
702db1f915145b5fd2947563aac8bd90
|
|
| BLAKE2b-256 |
e02c3c40556f69b4b537d1cd6bc79130f215eabe278bd6c2ad7915ec2151840a
|
File details
Details for the file mudra_ml-0.5.1-py3-none-any.whl.
File metadata
- Download URL: mudra_ml-0.5.1-py3-none-any.whl
- Upload date:
- Size: 80.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ab6f43de56356e12f1a922a24df292e21ced890af691ca7382d08d686fcebd4d
|
|
| MD5 |
91ec31a8ce5f50b533eda2abc3fd7228
|
|
| BLAKE2b-256 |
990f8b4aa081768203e2f05fc9042505d6dfd443ba054b620c3e020e6ea5cdb9
|