Skip to main content

millwright

crates.io docs.rs PyPI CI downloads license

A unified ML framework for Rust — fit. predict. serve. watch.

It assembles focused Rust ML engines behind one stable data model and trait contract, so training, evaluation, export, serving, and monitoring compose.

Status: Phases 0–8 — done

Phase 0 · the spine

fit · transform · predict · Pipeline end to end over a real backend:

  • Frame / Dataset — the contiguous, row-major f64 boundary type (src/frame.rs).
  • The four traits — object-safe Transformer, Estimator, Predictor, ProbaPredictor, plus a blanket Model (src/traits.rs).
  • The first backend — a smartcore adapter (RandomForest, LinearRegression) converting Frame → DenseMatrix at the edge only (src/backends/smartcore.rs).
  • Pipeline — named steps + a final model, "step__param" addressing; pipelines nest (src/pipeline.rs).

Phase 1 · prep & select — a real, tunable, ensemble-ready workflow

  • Preprocessing (src/transform.rs, core): SimpleImputer, StandardScaler, MinMaxScaler, OneHotEncoder, Winsorize (clip outliers), PowerTransform (Yeo-Johnson), ColumnTransformer (per-subset transforms), and the supervised TargetEncoder.
  • Balancing (src/balance.rs, via imbalance-rs): Smote, RandomOverSampler as train-time Balancers — Pipeline::balance(...), applied only during fit.
  • Model selection (src/selection/, via model-selection-rs): KFold / StratifiedKFold, a Metric enum (accuracy, F1, MAE, MSE, RMSE, R²), and GridSearch / RandomSearch over a whole pipeline, tuned by path. grid! macro included.
  • Ensembles (src/ensemble.rs, core): Voting (hard/soft), Bagging, and leak-free Stacking riding the same CV engine — all Models themselves, so they compose, tune, and nest.

Phase 2 · backends & HPO — two backends, one contract

  • The second backend (src/backends/linfa.rs, via linfa, feature linfa-backend): KMeans, GaussianMixture, Dbscan (as a new Clusterer contract) and Pca (as a Transformer) — each converting Frame → ndarray at the edge, proving the boundary conversion against a whole other engine.
  • Bayesian search (src/selection/, via hyperopt-rs, feature hpo): BayesSearch runs TPE search over a SearchSpace and returns the same SearchResult as grid/random search — one search API, three strategies.

Phase 3 · insight — trust the model, not just run it

  • Evaluation reports (src/evaluate.rs, core): model.evaluate(&test) bundles task-appropriate metrics into a Report (accuracy/precision/recall/F1 or MAE/MSE/RMSE/R²).
  • Regression diagnostics (src/diagnostics.rs, via regression-diagnostics, feature diagnostics): Diagnostics::of(&data) runs OLS and exposes summary(), R², per-column VIF, residuals, and Cook's distance.
  • Explainability (src/explain.rs, via shap-rs, feature explain): model.explain(&Explainer::kernel(), &frame) gives per-row SHAP values and global importance, plus permutation_importance(...).
  • Report figures (src/viz.rs, via plotters-statistical, feature viz): viz::roc_svg(...) and viz::residuals_svg(...) render self-contained SVGs (pure-Rust backend, no system fonts).
  • Probabilities (src/logistic.rs, core): LogisticRegression is a native, probability-capable classifier — the first real ProbaPredictor.
  • Calibration (src/calibration.rs, feature calibration): PlattScaling / IsotonicRegression and reliability_curve, plus CalibratedClassifier, which wraps any ProbaPredictor and returns calibrated probabilities.
  • Anomaly detection (src/anomaly.rs, feature anomaly): Mahalanobis and KnnScore, unified behind an OutlierDetector trait.

Phase 4 · portability & Python — train once; run in Rust, Python, or any ONNX runtime

  • ONNX export (src/onnx.rs, via onnx-export-rs, feature onnx): model.export_onnx(path) for RandomForest (ONNX-ML tree ensemble) and LinearRegression; whole-pipeline export folds affine scalers into the estimator's graph as one .onnx.
  • Inference (via tract, feature onnx): InferenceModel::load(path) loads and runs any ONNX file. tract executes the linear/affine/pipeline graphs (a full round-trip); tree-ensemble ONNX-ML artifacts run in external runtimes like onnxruntime.
  • Python bindings (src/python.rs, via pyo3, feature python): a Pipeline class over the same Rust core, shipped on PyPI as an abi3 wheel.
pip install millwright
import millwright as mw
pipe = mw.Pipeline()
pipe.standard_scaler()
pipe.random_forest(n_trees=100, max_depth=8)
pipe.fit(rows, labels)          # list[list[float]], list[float]
preds = pipe.predict(rows)      # runs the Rust engine

To build from source (contributors), from a virtualenv: maturin develop --features python.

Phase 5 · operations — past where scikit-learn stops

  • Registry (src/registry.rs, feature registry): Registry::local(path) versions a model's ONNX artifact, content-addressed (identical models dedupe), with metadata + reference distribution, movable tags, and rollback.
  • Drift monitor (src/monitor.rs, via driftwatch, feature monitor): DriftMonitor::psi(reference) watches the prediction stream — observe + report give live PSI and a drift verdict.
  • Server (src/serve.rs, via axum, feature serve): Server::from_onnx exposes POST /predict (validated) over the tract runtime; with a monitor attached, every request feeds it and GET /metrics reports drift.
Server::from_onnx(reg.onnx_path("churn", "prod")?)?
    .route("/predict")
    .with_monitor(DriftMonitor::psi(&reference)?)
    .serve("0.0.0.0:8080").await?;

Phase 6 · specialized — the long tail of real workloads

Same contract, different data shapes — each gets its own trait.

  • Time series (src/backends/chronos.rs, via chronos-ts, feature timeseries): AutoArima implements a Forecaster — fit(&series) then forecast(steps).
  • Out-of-core (src/backends/incremental.rs, via incremental-rs, feature incremental): IncrementalLinear implements PartialFit + Predictor — partial_fit(&batch) learns one batch at a time.

These two crates pin ndarray 0.15 while the rest of the stack uses 0.16; Cargo links both, and the boundary conversion happens only inside these adapters — the "two ndarray worlds" the design settles, now exercised for real.

Phase 7 · synthesis — auto-sklearn, but the output actually deploys

  • AutoML (src/automl.rs, feature automl): AutoML::classifier() / regressor() searches preprocessing × model × hyperparameters under a Budget (trials or minutes), auto-ensembles the top candidates, and returns a ranked leaderboard plus the best fitted model. No new crate — it orchestrates the model-selection, ensemble, and backend machinery already built. A single-pipeline winner flows straight into export_onnx, so unlike a TPOT object the result deploys.
let result = AutoML::classifier()
    .budget(Budget::trials(40))
    .metric(Metric::F1)
    .cv(StratifiedKFold::new(5))
    .fit(&train)?;
println!("{}", result.leaderboard());
result.export_onnx("model.onnx")?;   // deployable

Phase 8 · harden → 1.0 — a framework you can bet on

Pin, prove, document — owning the one real risk of assembling young, single-author engine crates.

  • Exact-version pins (Cargo.toml): every engine — the ecosystem crates plus the smartcore and linfa families — is pinned to an exact =x.y.z, so a stray cargo update can't move a fragile engine under the stable trait contract. General infrastructure (serde, tokio, axum, …) stays on caret ranges to avoid forcing conflicts downstream.
  • Committed Cargo.lock: the whole ~300-package graph is reproducible; CI builds with --locked.
  • Golden-output tests (tests/golden.rs): lock the numeric behaviour of the engines on fixed inputs — exact for the deterministic paths (OLS, affine transforms, metric formulas), well-separated class labels for the stochastic ones. An engine bump that moves a number shows up as a diff.
  • Feature-matrix CI (.github/workflows/ci.yml): fmt, clippy -D warnings, docs, and the test suite across the feature matrix — from --no-default-features through each feature to full — plus Windows/macOS, the runnable examples, a benchmark compile-check, a cargo publish --dry-run, and a maturin wheel. The MSRV (rust-version = 1.95, dep-dictated) is enforced by cargo for consumers.
  • The tutorial (GUIDE.md + guide.html): the design brief's lifecycle, re-cast as a hands-on guide.

Ingest & EDA — the lifecycle starts where the data does

The front of the lifecycle, behind the eda feature (via polars).

  • Table (src/table.rs): a dtype-aware, polars-backed table — Table::from_csv / from_parquet read real string/categorical/datetime/null columns. It lowers to the numeric world: table.to_frame() and table.into_dataset("target") (categoricals label-encoded, nulls → NaN), so Frame stays the numeric boundary everything else already speaks.
  • Profile (src/profile.rs): Profile::of(&table) returns a typed EDA — overview, per-column numeric/categorical profiles, missingness, Pearson correlations (high-|r| pairs flagged), IQR outliers, and target relationship (class balance or feature-target correlation). It renders a self-contained to_html(path) report, lists alerts() that name the fix, and — the loop scikit-learn can't close — suggest_pipeline() drafts the preprocessing from those findings; you just add the model.
let table = Table::from_csv("customers.csv")?;
let profile = Profile::of_with_target(&table, "churned")?;
profile.to_html("eda.html")?;

let train = table.into_dataset("churned")?;
let mut pipe = profile.suggest_pipeline()      // impute · encode · scale, from the alerts
    .estimator("rf", RandomForest::new());
pipe.fit(&train)?;

Quickstart

use millwright::grid;
use millwright::prelude::*;

let pipe = Pipeline::new()
    .step("impute", SimpleImputer::median())
    .step("scale", StandardScaler::new())
    .balance(Smote::new())                 // train-time only
    .estimator("rf", RandomForest::new());

let search = GridSearch::new(pipe, grid! { "rf__max_depth" => [4, 8, 16] })
    .cv(StratifiedKFold::new(5))
    .scoring(Metric::F1)
    .fit(&train)?;

println!("best F1 = {:.3}", search.best_score());
let preds = search.predict(&test)?;

Run the end-to-end examples:

cargo run --example spine
cargo run --example explore --features "eda smartcore-backend"
cargo run --example trust --features "calibration anomaly"
cargo run --example workflow
cargo run --example backends --features "smartcore-backend linfa-backend hpo"
cargo run --example insight --features "smartcore-backend diagnostics explain viz"
cargo run --example portability --features "smartcore-backend onnx"
cargo run --example operations --features "smartcore-backend onnx registry monitor serve"
cargo run --example specialized --features "timeseries incremental"
cargo run --example automl --features "smartcore-backend automl onnx"

Building on Windows

The default toolchain is MSVC. If a Unix link.exe (e.g. from Git/Laragon) is ahead of MSVC's on PATH, linking fails with an "extra operand" error. Build from a Developer Command Prompt / PowerShell for VS 2022, or run vcvars64.bat first, so the MSVC linker is found before the shadowing one.

Roadmap

Phases 0–8 are done — the full lifecycle plus 1.0 hardening (exact-version pins, a committed lockfile, golden-output tests, and a feature-matrix CI). The design brief lays out the arc; the tutorial (GUIDE.md) is the how.

Metadata

Release files for millwright 2.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for millwright 2.2.0
File Size Uploaded
millwright-2.2.0.tar.gz 200.8 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for millwright 2.2.0
File
millwright-2.2.0-cp39-abi3-win_amd64.whl CPython 3.9 abi3 Windows x86-64 Details
millwright-2.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.9 abi3 Linux glibc 2.17+ x86-64 Details
millwright-2.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl CPython 3.9 abi3 Linux glibc 2.17+ ARM64 Details
millwright-2.2.0-cp39-abi3-macosx_11_0_arm64.whl CPython 3.9 abi3 macOS 11.0+ ARM64 Details
millwright-2.2.0-cp39-abi3-macosx_10_12_x86_64.whl CPython 3.9 abi3 macOS 10.12+ x86-64 Details

Total release size: 78.5 MB

Release files / millwright-2.2.0.tar.gz

Download URL millwright-2.2.0.tar.gz
Size 200.8 kB
Tags Source
SHA-256 checksum
How to use checksums
096b578fafc364318e9723007ebf04e041c8079b7d72cad57a803f1526d20ff5
BLAKE2b-256 checksum
How to use checksums
73120cf94b9d0306edb7383f2ac36af687d071b625e8401f7fc3c463988114b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / millwright-2.2.0-cp39-abi3-win_amd64.whl

Download URL millwright-2.2.0-cp39-abi3-win_amd64.whl
Size 14.7 MB
Tags CPython 3.9 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
a63b03578670cfac4aa9d88969cd02628e20832add9ca030f2c0047cf3b816b3
BLAKE2b-256 checksum
How to use checksums
a8d6726b8708548efd7342924a67c3bf678ca24f7a0fb578699653b2c9c2d31a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / millwright-2.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL millwright-2.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 16.7 MB
Tags CPython 3.9 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
9d24dcce56f6b2933bc3d8c54d3cc66310e36ba0ad306426bc6cd9d5f1288b6d
BLAKE2b-256 checksum
How to use checksums
010d23cf9c7edca385e1794743928871afda8edd9efc68858aa37e98d158ea63
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / millwright-2.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl

Download URL millwright-2.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Size 17.3 MB
Tags CPython 3.9 Linux glibc 2.17+ ARM64 abi3
SHA-256 checksum
How to use checksums
bf8128286b2e1b7d461d0048fa056a67d962957d39ea47d2ae16b4d661ae2989
BLAKE2b-256 checksum
How to use checksums
b46e87f93efc0a92f2d4b05250c6f5e6b6ac5a572e5c61cde8d35f471873e124
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / millwright-2.2.0-cp39-abi3-macosx_11_0_arm64.whl

Download URL millwright-2.2.0-cp39-abi3-macosx_11_0_arm64.whl
Size 14.2 MB
Tags CPython 3.9 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
89fc22618141dad905deb5c51f38c5865e2fe72e92765ed0c8cc97287388806c
BLAKE2b-256 checksum
How to use checksums
5a77447ea5bbf7cffb74840fc1c5b6a9f5b61cccec3e1fc8543f5d57a6c9f1a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / millwright-2.2.0-cp39-abi3-macosx_10_12_x86_64.whl

Download URL millwright-2.2.0-cp39-abi3-macosx_10_12_x86_64.whl
Size 15.5 MB
Tags CPython 3.9 abi3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
1a155ad93278659b85b295cf1eb2e05f8c295ddc7baa94b4c625356b8137f584
BLAKE2b-256 checksum
How to use checksums
75c9e0a115414cfebd0b5914d21161d850cbb3d0ae3b4b73dcfd677248e57cdc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release history Release notifications | RSS feed

2.3.1

6 release files

2.2.1

6 release files

This release

2.2.0 This release

6 release files

0.2.1

6 release files

0.2.0

6 release files

0.1.1

6 release files

0.1.0

6 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page