Skip to main content

maatml

PyPI License Python CI

MaatML fine-tunes small, task-specific models across text, vision, and vision-language, and takes them from experimentation to production through a single declarative model.yml: prepare → train → evaluate → export → serve. Licensed under Apache-2.0.

What makes it different: correctness is checked outside the model by validators. The same validator gates your synthetic data and your evaluation, and guards your live inference: maatml serve runs it on every response, reporting the result by default and rejecting outputs that fail under --enforce. So a MaatML model ships with a contract, not just weights. That validator-gated data → eval → serving loop, now across modalities, is what general fine-tuning tools leave out.

Site: maatml.pages.dev · PyPI: maatml · Source: github.com/moralfish/maatml

Installation

python -m venv .venv
source .venv/bin/activate

# Library + CLI (no torch)
pip install maatml

# Training / evaluation stack
pip install "maatml[ml]"

# Optional extras
pip install "maatml[ml,cuda]"    # QLoRA on NVIDIA CUDA (bitsandbytes)
pip install "maatml[ml,pref]"    # DPO / ORPO (TRL)
pip install "maatml[ml,vision]"  # torchvision + ONNX (examples/vision)
pip install "maatml[vllm]"       # Linux-only vLLM serving (examples/vision-vlm)
pip install "maatml[teacher]"    # OpenAI-compatible teacher for datagen
pip install "maatml[docs]"       # mkdocs site

Then:

maatml --help
maatml scaffold ~/models/my-task --architecture causal_sft --name my-task
maatml validate ~/models/my-task

For contributing to this repository (editable install), see CONTRIBUTING.md.

Example models

Six reference models share the identical folder layout and CLI, from a one-command support-ticket triage to a vLLM-servable vision-language model:

Model Task Architecture Base
Support Ticket Triage triage → JSON causal_sft (LoRA) Qwen3-0.6B
Vision VLM describe a scene image vlm_sft (vLLM-servable) SmolVLM-256M-Instruct
Vision scene + detect + pose vision_multitask MobileNetV3-Large
Vision Describer caption from vision JSON seq2seq flan-t5-small
JCL Validator jcl_validation classifier (4-head) ModernBERT-base
Spool Interpreter spool_interpretation seq2seq flan-t5-base

Any directory with a valid model.yml works the same way: install maatml from PyPI and point the CLI at the folder. Scaffold a new model folder with maatml scaffold.

Where MaatML fits

MaatML builds on Hugging Face transformers / peft / trl and does the one thing those building blocks leave to you: it wraps them in an opinionated, validator-gated lifecycle for small task-specific models you can train on a laptop and deploy to the edge or vLLM.

  • Complements general fine-tuning tools (Axolotl, LLaMA-Factory, Unsloth, TRL) rather than competing on scale. Reach for those for large models, multi-node training, RL, or broad model coverage.
  • Runs its own fixed lifecycle (prepare → train → evaluate → export → verify / serve), as one command: maatml run, which skips steps that are already fresh and stops non-zero at the first failure. It is not a general-purpose workflow scheduler: no triggers, no arbitrary shell/Python steps, no remote executors. Drop maatml train into MLflow / Prefect / Metaflow when you need that.
  • Its niche: local-first, multimodal, structured-output models with correctness gated outside the model, from data generation through serving.

Requirements

  • Python 3.10+ (developed against 3.13)
  • OS macOS, Linux (Windows untested)
  • Disk / memory ~3 GB for the ML stack; 16 GB unified memory is the design target for local training

CLI overview

Most commands take a model folder (containing model.yml) as their first argument. Outputs land under <model-folder>/output/ (gitignored). Run maatml <command> --help for the full flag list. Errors in your input (a missing file, an unparseable model.yml, an unregistered plugin) print one line; maatml --debug <command> prints the traceback.

Command What it does
run The whole lifecycle in one command: prepare, train, evaluate (gated), export, verify. Skips steps that are already fresh
prepare Build train/val/test splits from the seed corpus
train Fine-tune the model (--smoke, --resume auto|PATH, --set K=V)
sweep Offline grid HPO over --param K=a,b
evaluate Score a checkpoint; --gate exits non-zero on a gate miss. The token budget defaults to packaging.max_input_tokens
export Deployable bundle + manifest.json (--format, --parity)
verify Recompute sha256 of an export against its manifest.json
serve JSON inference API; --enforce (422), --max-retries, --auth-token, --capture
datagen Validator-gated seed generation (--teacher, --allow-ungated)
distill Validator-gated teacher labels over a prompt pool (--replay offline)
mint Preference pairs (chosen/rejected) from validator-scored candidates
ingest Import external samples (--map field=col, --sanitize tag)
runs List recorded training runs (--compare tabulates their metrics)
plan Show which lifecycle steps are stale (alias for run --dry-run)
plugins List discovered trainers, validators, and metrics
doctor Check the environment, plugins, and a model folder; exits 1 on problems
scaffold Create a new model folder (--architecture, --plugin, --force)
validate Check model.yml and paths (--no-plugins skips plugin code)

Multi-GPU (CUDA): accelerate launch -m maatml.cli train <model-dir>/ or torchrun --nproc_per_node=N -m maatml.cli train <model-dir>/.

QLoRA (CUDA + [cuda]): set training.quantization.load_in_4bit: true in model.yml. Preference data: dataset.format: preference_jsonl with {prompt, chosen, rejected} rows; scaffold with --architecture dpo.

Export defaults to a safetensors bundle + manifest.json. GGUF/MLX need external tooling (llama.cpp convert / mlx_lm). Pin base-model revisions with training.model_revision.

Docs: maatml.pages.dev · Roadmap: ROADMAP.md · In-repo docs: docs/ (pip install "maatml[docs]" then mkdocs serve).

End-to-end example (Support Ticket Triage)

The quickest model to run: a LoRA fine-tune of Qwen3-0.6B that turns a raw support ticket into {priority, category, team, summary} JSON, gated by a schema validator plus a category → team routing contract enforced outside the model. Every reference model now registers a validator and declares evaluation.gates.

git clone https://github.com/moralfish/maatml.git
cd maatml
pip install "maatml[ml]"

maatml prepare  examples/support-ticket-triage/
maatml train    examples/support-ticket-triage/ --smoke   # fast pipeline check
maatml train    examples/support-ticket-triage/
maatml evaluate examples/support-ticket-triage/ --gate    # enforce eval gates
maatml serve    examples/support-ticket-triage/           # JSON inference API

For a multimodal walkthrough (image → description, servable by vLLM) see examples/vision-vlm/. For an advanced model with a custom column-aware tokenizer, see examples/jcl-validator/. Its tokenizer is built once up front:

python examples/jcl-validator/scripts/build_seeds.py --target 10000 \
  --out examples/jcl-validator/datasets/samples/tokenizer_corpus.jsonl
python examples/jcl-validator/scripts/build_tokenizer.py

Batch scripts

# Deterministic seed corpora (no API calls)
python examples/jcl-validator/scripts/build_seeds.py
python examples/spool-interpreter/scripts/build_seeds.py

# Train / evaluate example models
python scripts/train_all.py --smoke
python scripts/train_all.py
python scripts/evaluate_all.py

Apple Silicon / MPS notes

  • Default precision is bf16 autocast with fp32 master weights.
  • Trainers set eval_steps: 9999 to disable mid-training eval on MPS (unified memory does not release val-set tensors between eval and training).
  • grad_checkpointing defaults to false; dataloader_num_workers=0 everywhere (multi-worker + MPS can deadlock via fork pickling).
  • PYTORCH_ENABLE_MPS_FALLBACK=1 is set by the CLI for unsupported ops.

Repository layout

examples/                   # reference task models (plugins + data)
  jcl-validator/
    model.yml               # single source of truth
    jcl_plugin/             # validator, metrics, predictor, tokenizer, …
    datasets/               # schemas, prompt specs, seed samples
    scripts/                # seed + tokenizer builders
    output/                 # gitignored prepared / checkpoints / eval
  spool-interpreter/        # same layout (uses core seq2seq)
  support-ticket-triage/

src/maatml/                 # core framework (architectures, CLI, harnesses)
scripts/                    # batch train/eval/validate
tests/                      # core unit tests

Trust boundary

A model folder is executable code, not just data. Any plugins: entry in model.yml is imported as Python the moment the folder is loaded, and every command that reads model.yml, including maatml validate and maatml plan, loads the folder. Running any maatml command against a folder therefore runs that folder's code with your privileges. Only run maatml on model folders you trust, the same way you would only run a script you trust. Use maatml validate --no-plugins to check the schema and paths without importing plugin code.

Development

See CONTRIBUTING.md for setup, PR expectations, DCO sign-off, and versioning policy. AI coding agents: AGENTS.md.

Community: CODE_OF_CONDUCT.md · Security: SECURITY.md · Changes: CHANGELOG.md

Licensing

  • maatml is licensed under the Apache License 2.0.
  • This repository does not redistribute base-model weights, only Hugging Face Hub IDs. The reference bases (ModernBERT, flan-t5, and related Apache-2.0 models such as Qwen3 when used) are Apache-2.0; your fine-tuned checkpoints inherit the base model's license terms.
  • Seed corpora are fully synthetic, produced by deterministic builders under examples/*/scripts/, with no proprietary mainframe dumps shipped.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

maatml-0.8.0.tar.gz (173.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

maatml-0.8.0-py3-none-any.whl (152.7 kB view details)

Uploaded Python 3

File details

Details for the file maatml-0.8.0.tar.gz.

File metadata

  • Download URL: maatml-0.8.0.tar.gz
  • Upload date:
  • Size: 173.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for maatml-0.8.0.tar.gz
Algorithm Hash digest
SHA256 acf4951ae58385797cc2e762cbca6a02f9ef5295b2098fd0cf43223bc38f686e
MD5 7ddc5c4d1336f3cf4e26b4200801d356
BLAKE2b-256 c1319c9e80e591d9593e581ff9264aa82ed924225e43157852b3ce9164e345d7

See more details on using hashes here.

Provenance

The following attestation bundles were made for maatml-0.8.0.tar.gz:

Publisher: publish.yml on moralfish/maatml

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file maatml-0.8.0-py3-none-any.whl.

File metadata

  • Download URL: maatml-0.8.0-py3-none-any.whl
  • Upload date:
  • Size: 152.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for maatml-0.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 d932790661da0e2f012a6d33e1b6b9f96f80eb9f0f88e791b3edefda21dfa785
MD5 15108d64ea2f21b9e52f1c459802fb60
BLAKE2b-256 4c232b45cee2bea00091ddb9bef5bcef7b84efc3b54455053202c1001cf94bb9

See more details on using hashes here.

Provenance

The following attestation bundles were made for maatml-0.8.0-py3-none-any.whl:

Publisher: publish.yml on moralfish/maatml

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.13.0

2 files

0.12.0

2 files

0.11.2

2 files

0.11.1

2 files

0.11.0

2 files

0.10.0

2 files

0.9.1

2 files

0.9.0

2 files

This release

0.8.0 This release

2 files

0.7.0

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page