Compile a YAML recipe into a reproducible, framework-agnostic trained-model instance.
Project description
ModelFoundry
Compile a YAML recipe into a reproducible, framework-agnostic trained-model instance.
ModelFoundry consumes a materialized DataRefinery instance and compiles a single YAML model recipe into a content-addressed, atomically-promoted ModelInstance: the trained model, per-epoch metrics, hyperparameter-search trials, held-out evaluation, predictions, visualizations, and a manifest. The result object returns notebook-shaped primitives (pandas.DataFrame / numpy.ndarray / PNG bytes) and works identically inside Jupyter, Marimo, IPython, or a plain .py script — no framework imports in user code.
Reproducibility is a first-class concern: every stochastic source is seeded, the cache identity is computed from the recipe's normalized semantic form, and the same (recipe, data, seed, variant) tuple materializes to a byte-identical ModelInstance.
Status: pre-production (
0.x.yseries). APIs, CLI surface, and cache layout may change between minor versions until the1.0.0production release. Seedocs/specs/for the concept, feature, technical, and story specifications.
Installation
pip install ml-modelfoundry[pytorch]
The import name and console script are both modelfoundry; the PyPI distribution is ml-modelfoundry. The pre-production release ships an end-to-end PyTorch plugin (image classification, CIFAR-10-scale) plus a scikit-learn MLPClassifier baseline; the base install (pip install ml-modelfoundry) carries everything except the framework — a recipe selects its backend via the [pytorch] extra.
Quickstart — CIFAR-10
ModelFoundry never does data prep: splitting, cleaning, sampling, and feature engineering are DataRefinery's job. The quickstart assumes the two bundled recipes — recipes/cifar10-base.yaml (the DataRefinery dataset recipe) and recipes/cifar10_resnet20.yml (the ModelFoundry ResNet-20 recipe, bound to it).
# 1. Materialize the CIFAR-10 dataset with DataRefinery (one-time) → ./data
datarefinery materialize recipes/cifar10-base.yaml
# 2. Validate, then materialize the model with ModelFoundry → ./models
modelfoundry validate recipes/cifar10_resnet20.yml
modelfoundry materialize recipes/cifar10_resnet20.yml
materialize runs the full pipeline — hyperparameter optimization → training → held-out evaluation → output-expectation checks → persistence → report — and atomically promotes the result into the content-addressed cache. Re-running the same recipe finds the existing instance; pass --overwrite to recompute.
Then consume the materialized instance — from a script, a notebook, or the CLI:
from datarefinery import DataRefinery
from modelfoundry import ModelFoundry
data = DataRefinery.from_recipe("recipes/cifar10-base.yaml").materialize()
model = ModelFoundry.from_recipe("recipes/cifar10_resnet20.yml", data=data).materialize()
model.evaluation["test"] # dict[str, value] — held-out metrics for the test split
model.metrics # alias for .evaluation: {split: {metric: value}}
model.confusion_matrix # dict[str, np.ndarray] — per-split confusion matrices
model.predictions # pandas.DataFrame — per-record predictions + class probabilities
model.figures # dict[str, bytes] — reporting-visualization PNGs, keyed by name
model.predict(X) # np.ndarray — predicted labels for new inputs
Swap the model — three baselines, one workflow
In ModelFoundry the recipe is the model definition. Changing the classification model — from a chance baseline, to a scikit-learn MLP, to a PyTorch CNN — is a declarative edit to two lines of YAML (plugin + Architecture). The Python you write to train and evaluate is identical across all of them:
from modelfoundry import ModelFoundry
for recipe in ("recipes/cifar10_random.yml", # chance floor — the `random` plugin
"recipes/cifar10_mlp.yml", # scikit-learn MLP baseline
"recipes/cifar10_cnn.yml"): # PyTorch simple_cnn
mi = ModelFoundry.from_recipe(recipe, data="./data").materialize()
print(recipe, mi.evaluation["test"]["accuracy"])
The three recipes share the same DataRefinery binding, Training, and Evaluation blocks — only the head changes (full annotated recipes live in recipes/):
# cifar10_random.yml — the chance floor
plugin: random
Architecture: {type: dummy_classifier, num_classes: 10, strategy: stratified}
Loss: {op: cross_entropy}
Optimizer: {op: "none"} # a chance baseline has no optimizer
# (omits Optimization + Visualizations — a fixed baseline has neither)
# cifar10_mlp.yml — flattened-pixel scikit-learn MLP
plugin: sklearn
Architecture: {type: mlp_classifier, num_classes: 10, hidden_layer_sizes: [256, 128], max_iter: 50}
Loss: {op: cross_entropy}
Optimizer: {op: adam, learning_rate: 0.001} # drives the MLPClassifier solver
# (omits Optimization + Visualizations — the baseline plugin implements neither)
# cifar10_cnn.yml — PyTorch simple_cnn
plugin: pytorch
Architecture: {type: simple_cnn, num_classes: 10, in_channels: 3}
Loss: {op: cross_entropy}
Optimizer: {op: adamw, learning_rate: 0.001}
Training: {max_epochs: 5, ...} # a deliberately small base budget
The
sklearnandrandombaselines reuse the PyTorch feature path, so all three currently need the[pytorch]extra (pip install ml-modelfoundry[pytorch]).
Capacity is latent until you scale the budget
Run all four and the result tells a bigger story than "use a CNN" (CPU, deterministic, on the 1,700-image CIFAR-10 subset):
| Model | recipe | test accuracy |
|---|---|---|
| Random (chance) | cifar10_random.yml |
0.095 |
| PyTorch CNN — 5 epochs | cifar10_cnn.yml |
0.275 |
| scikit-learn MLP | cifar10_mlp.yml |
0.352 |
| PyTorch CNN — 40 epochs | cifar10_cnn.yml --variant well_trained |
0.403 |
The more-expressive CNN loses to the flattened-pixel MLP at a small training budget, and only overtakes it once the budget is scaled up — the same capacity-vs-budget dynamic that separates a legacy model from a modern over-parameterized one. Scaling the budget is itself a one-line recipe change, expressed as a variant:
variants:
well_trained:
Training: {max_epochs: 40}
modelfoundry materialize recipes/cifar10_cnn.yml --variant well_trained
Every run is content-addressed and reproducible, so each comparison is cached and byte-stable — re-running finds the existing instance instead of recomputing.
This is a teaching illustration, not a benchmark. On the 1,700-image subset the per-epoch trajectory is noisy and non-monotonic — the minimal recipes use no LR schedule or early stopping, so a single run can dip or spike between budgets (a swept study even shows
resnet20degrading past its peak). The endpoint contrast above is real and reproducible, but a robust capacity-vs-budget crossover needs more data and a proper training regime. Seescripts/experiments/for the full sweep and that finding.
Library API
ModelFoundry.from_recipe(...) binds a recipe to a materialized DataRefinery instance; the verbs (validate / materialize / status / inspect / report / clean / check) are thin methods over that binding, co-equal with the CLI.
from modelfoundry import ModelFoundry, ModelInstance
mf = ModelFoundry.from_recipe("model.yml", data=data)
report = mf.validate() # FR-2 static checks; report.passed is a bool
instance = mf.materialize() # train + optimize + evaluate; returns a ModelInstance
# A reloaded instance predicts identically (byte-stable round-trip):
reloaded = ModelInstance.load(instance.path)
data may be a pre-bound DataRefineryInstance (as above) or a path to the DataRefinery cache root, in which case the recipe's Data: block is resolved against it.
CLI
modelfoundry check # environment + plugin health
modelfoundry validate <recipe> # static FR-2 recipe checks
modelfoundry materialize <recipe> [--overwrite] # train + optimize + evaluate
modelfoundry status <recipe> # is it materialized? show the manifest
modelfoundry report <instance-dir> # re-render the instance report
modelfoundry inspect <instance-dir> --view training_curves
modelfoundry clean --older-than 7d # cache management
modelfoundry init <recipe-out> --data <datarefinery-recipe> # scaffold a recipe
Shared options apply to every verb: --cache-root / --data-cache-root (defaults ./models and ./data), --log-level, --log-target (JSON-lines operational logs), --plugin-path, and -v / -q.
Notebook-substrate-neutral
The same surface works identically in a Jupyter cell, a Marimo cell, an IPython REPL, or a plain .py script — the ModelInstance returns plain pandas / numpy / PNG-bytes primitives, so user code imports no framework:
from IPython.display import Image
mi = ModelFoundry.from_recipe("model.yml", data=data).materialize()
Image(mi.figures["training_curves"]) # render the reporting PNG
mi.predictions.head() # a DataFrame, renders natively in any host
Choosing an accelerator
Hardware acceleration is auto-detected by default — the PyTorch plugin picks Metal (Apple Silicon) → CUDA → CPU in that order. To pin a specific device (e.g. for CPU-speed benchmarking on a GPU-equipped machine, or to debug a non-deterministic op), set Training.device in the recipe:
Training:
max_epochs: 10
batch_size: 32
device: cpu # one of: auto (default) | cpu | cuda | mps
device participates in the recipe's canonical hash, so the same recipe run with device: cpu and device: mps materializes into two distinct ModelInstance cache entries — no silent cross-device collision. Use the variants: block to keep both side-by-side without maintaining two recipe files:
variants:
cpu_bench:
Training: {device: cpu}
modelfoundry materialize model.yml --variant cpu_bench
Documentation
docs/specs/concept.md— why the project existsdocs/specs/features.md— what it does (CR / FR / UR / TR requirements)docs/specs/tech-spec.md— how it is builtdocs/specs/project-essentials.md— must-know invariants (cache identity, determinism, loose coupling)docs/specs/stories.md— the implementation plan
License
Apache-2.0. Copyright (c) 2026 Pointmatic.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ml_modelfoundry-0.10.2.tar.gz.
File metadata
- Download URL: ml_modelfoundry-0.10.2.tar.gz
- Upload date:
- Size: 818.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
aa50cada3ccd8731ff2a9581868e8bd5572e4869e9d1d65c8ab0de2c81cc1f7f
|
|
| MD5 |
f9070ed4fb40ba66bbb76b9c6ff44bb9
|
|
| BLAKE2b-256 |
e122b18b5b353b661d7fdab38befe97a9169e8b10b154b4feed25c91d013048a
|
Provenance
The following attestation bundles were made for ml_modelfoundry-0.10.2.tar.gz:
Publisher:
publish.yml on pointmatic/modelfoundry
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
ml_modelfoundry-0.10.2.tar.gz -
Subject digest:
aa50cada3ccd8731ff2a9581868e8bd5572e4869e9d1d65c8ab0de2c81cc1f7f - Sigstore transparency entry: 1858240302
- Sigstore integration time:
-
Permalink:
pointmatic/modelfoundry@a094f862ef036d483b263236cf72f6c5979c3856 -
Branch / Tag:
refs/tags/v0.10.2 - Owner: https://github.com/pointmatic
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a094f862ef036d483b263236cf72f6c5979c3856 -
Trigger Event:
push
-
Statement type:
File details
Details for the file ml_modelfoundry-0.10.2-py3-none-any.whl.
File metadata
- Download URL: ml_modelfoundry-0.10.2-py3-none-any.whl
- Upload date:
- Size: 147.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bc4f7ed93ae696bc731153ec9540e8525e5f03f2452e19d207d8a89a3ae079c2
|
|
| MD5 |
7026863e77b92382d15394f43409bc4f
|
|
| BLAKE2b-256 |
055430c6e58f14b17a6b3e8d975b81ea6689508c91a6d0edd0175cc61589c1da
|
Provenance
The following attestation bundles were made for ml_modelfoundry-0.10.2-py3-none-any.whl:
Publisher:
publish.yml on pointmatic/modelfoundry
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
ml_modelfoundry-0.10.2-py3-none-any.whl -
Subject digest:
bc4f7ed93ae696bc731153ec9540e8525e5f03f2452e19d207d8a89a3ae079c2 - Sigstore transparency entry: 1858240364
- Sigstore integration time:
-
Permalink:
pointmatic/modelfoundry@a094f862ef036d483b263236cf72f6c5979c3856 -
Branch / Tag:
refs/tags/v0.10.2 - Owner: https://github.com/pointmatic
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a094f862ef036d483b263236cf72f6c5979c3856 -
Trigger Event:
push
-
Statement type: