Skip to main content

Compile a YAML recipe into a reproducible, framework-agnostic trained-model instance.

Project description

ModelFoundry

License: Apache 2.0

Compile a YAML recipe into a reproducible, framework-agnostic trained-model instance.

ModelFoundry consumes a materialized DataRefinery instance and compiles a single YAML model recipe into a content-addressed, atomically-promoted ModelInstance: the trained model, per-epoch metrics, hyperparameter-search trials, held-out evaluation, predictions, visualizations, and a manifest. The result object returns notebook-shaped primitives (pandas.DataFrame / numpy.ndarray / PNG bytes) and works identically inside Jupyter, Marimo, IPython, or a plain .py script — no framework imports in user code.

Reproducibility is a first-class concern: every stochastic source is seeded, the cache identity is computed from the recipe's normalized semantic form, and the same (recipe, data, seed, variant) tuple materializes to a byte-identical ModelInstance.

Status: pre-production (0.x.y series). APIs, CLI surface, and cache layout may change between minor versions until the 1.0.0 production release. See docs/specs/ for the concept, feature, technical, and story specifications.

Installation

pip install ml-modelfoundry[pytorch]

The import name and console script are both modelfoundry; the PyPI distribution is ml-modelfoundry. The pre-production release ships an end-to-end PyTorch plugin (image classification, CIFAR-10-scale) plus a scikit-learn MLPClassifier baseline; the base install (pip install ml-modelfoundry) carries everything except the framework — a recipe selects its backend via the [pytorch] extra.

Quickstart — CIFAR-10

ModelFoundry never does data prep: splitting, cleaning, sampling, and feature engineering are DataRefinery's job. The quickstart assumes the two bundled recipes — recipes/cifar10-base.yaml (the DataRefinery dataset recipe) and recipes/cifar10_resnet20.yml (the ModelFoundry ResNet-20 recipe, bound to it).

# 1. Materialize the CIFAR-10 dataset with DataRefinery (one-time) → ./data
datarefinery materialize recipes/cifar10-base.yaml

# 2. Validate, then materialize the model with ModelFoundry → ./models
modelfoundry validate    recipes/cifar10_resnet20.yml
modelfoundry materialize recipes/cifar10_resnet20.yml

materialize runs the full pipeline — hyperparameter optimization → training → held-out evaluation → output-expectation checks → persistence → report — and atomically promotes the result into the content-addressed cache. Re-running the same recipe finds the existing instance; pass --overwrite to recompute.

Then consume the materialized instance — from a script, a notebook, or the CLI:

from datarefinery import DataRefinery
from modelfoundry import ModelFoundry

data = DataRefinery.from_recipe("recipes/cifar10-base.yaml").materialize()
model = ModelFoundry.from_recipe("recipes/cifar10_resnet20.yml", data=data).materialize()

model.evaluation["test"]   # dict[str, value] — held-out metrics for the test split
model.metrics              # alias for .evaluation: {split: {metric: value}}
model.confusion_matrix     # dict[str, np.ndarray] — per-split confusion matrices
model.predictions          # pandas.DataFrame — per-record predictions + class probabilities
model.figures              # dict[str, bytes] — reporting-visualization PNGs, keyed by name
model.predict(X)           # np.ndarray — predicted labels for new inputs

Library API

ModelFoundry.from_recipe(...) binds a recipe to a materialized DataRefinery instance; the verbs (validate / materialize / status / inspect / report / clean / check) are thin methods over that binding, co-equal with the CLI.

from modelfoundry import ModelFoundry, ModelInstance

mf = ModelFoundry.from_recipe("model.yml", data=data)

report = mf.validate()              # FR-2 static checks; report.passed is a bool
instance = mf.materialize()         # train + optimize + evaluate; returns a ModelInstance

# A reloaded instance predicts identically (byte-stable round-trip):
reloaded = ModelInstance.load(instance.path)

data may be a pre-bound DataRefineryInstance (as above) or a path to the DataRefinery cache root, in which case the recipe's Data: block is resolved against it.

CLI

modelfoundry check                              # environment + plugin health
modelfoundry validate    <recipe>               # static FR-2 recipe checks
modelfoundry materialize <recipe> [--overwrite] # train + optimize + evaluate
modelfoundry status      <recipe>               # is it materialized? show the manifest
modelfoundry report      <instance-dir>         # re-render the instance report
modelfoundry inspect     <instance-dir> --view training_curves
modelfoundry clean       --older-than 7d        # cache management
modelfoundry init        <recipe-out> --data <datarefinery-recipe>   # scaffold a recipe

Shared options apply to every verb: --cache-root / --data-cache-root (defaults ./models and ./data), --log-level, --log-target (JSON-lines operational logs), --plugin-path, and -v / -q.

Notebook-substrate-neutral

The same surface works identically in a Jupyter cell, a Marimo cell, an IPython REPL, or a plain .py script — the ModelInstance returns plain pandas / numpy / PNG-bytes primitives, so user code imports no framework:

from IPython.display import Image

mi = ModelFoundry.from_recipe("model.yml", data=data).materialize()
Image(mi.figures["training_curves"])   # render the reporting PNG
mi.predictions.head()                  # a DataFrame, renders natively in any host

Choosing an accelerator

Hardware acceleration is auto-detected by default — the PyTorch plugin picks Metal (Apple Silicon) → CUDA → CPU in that order. To pin a specific device (e.g. for CPU-speed benchmarking on a GPU-equipped machine, or to debug a non-deterministic op), set Training.device in the recipe:

Training:
  max_epochs: 10
  batch_size: 32
  device: cpu              # one of: auto (default) | cpu | cuda | mps

device participates in the recipe's canonical hash, so the same recipe run with device: cpu and device: mps materializes into two distinct ModelInstance cache entries — no silent cross-device collision. Use the variants: block to keep both side-by-side without maintaining two recipe files:

variants:
  cpu_bench:
    Training: {device: cpu}
modelfoundry materialize model.yml --variant cpu_bench

Documentation

License

Apache-2.0. Copyright (c) 2026 Pointmatic.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ml_modelfoundry-0.8.2.tar.gz (511.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ml_modelfoundry-0.8.2-py3-none-any.whl (140.3 kB view details)

Uploaded Python 3

File details

Details for the file ml_modelfoundry-0.8.2.tar.gz.

File metadata

  • Download URL: ml_modelfoundry-0.8.2.tar.gz
  • Upload date:
  • Size: 511.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for ml_modelfoundry-0.8.2.tar.gz
Algorithm Hash digest
SHA256 c98cf5d591d3f879dcfef0dc99a722359e978f87e9e51d093568db82adae1e4a
MD5 094350adc14f340e6c9d5c9dd96351fa
BLAKE2b-256 d9a84b519c136b85c62fc415962496d20259333aa94b9363ab4914dba03a422f

See more details on using hashes here.

Provenance

The following attestation bundles were made for ml_modelfoundry-0.8.2.tar.gz:

Publisher: publish.yml on pointmatic/modelfoundry

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ml_modelfoundry-0.8.2-py3-none-any.whl.

File metadata

  • Download URL: ml_modelfoundry-0.8.2-py3-none-any.whl
  • Upload date:
  • Size: 140.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for ml_modelfoundry-0.8.2-py3-none-any.whl
Algorithm Hash digest
SHA256 12a6b04fa02a1dff9216742f594d0d0baa0c8f6699ca378178b5ba956d272292
MD5 cc77b499b3755bf33c268feb8f17247b
BLAKE2b-256 ff340eba45e1023905e60abc02ca59e60dca57d306e2cc0db6a845b8cd63c337

See more details on using hashes here.

Provenance

The following attestation bundles were made for ml_modelfoundry-0.8.2-py3-none-any.whl:

Publisher: publish.yml on pointmatic/modelfoundry

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page