mojolearn
Machine learning that trains and predicts bitwise identically across Apple, NVIDIA and AMD GPUs, for certified configurations.
Give mojolearn the same code, data, hyperparameters and seed on two certified
machines and you get the same bits on both. Not close, not within a tolerance.
The same bits. The byte-level language model record holds on an NVIDIA RTX
4090 (sm_89), an AMD MI325X (gfx942) and an Apple M4
(record);
the forest record of September 12 holds on an NVIDIA H100 (sm_90a), an AMD
MI300X (gfx942) and an Apple M4
(brief). A model trained
on AMD and the same model trained on NVIDIA are byte for byte the same model,
and either one makes exactly the same predictions. This is identical mode, and it is the
default. The claim is proven by stage-level identity cards and separating
sabotage tests, never inferred from a final-output hash, and it holds only for
the configurations recorded in the support matrix.
RF/ET offer inference_engine="sequential" (existing host prediction) and
experimental inference_engine="parallel_groves" (shared GPU prediction).
Both retain GPU training; see the inference algorithms and numerical contract.
Since 0.8.0 the library has no NumPy runtime dependency and returns
mojolearn.Array objects. Existing NumPy inputs remain supported; callers can
use numpy.asarray(result) for a zero-copy view. See the
NumPy-free contract and
qualification roadmap.
Random forest and ExtraTrees fit now export model bytes directly into owned Array buffers, avoiding per-node Python objects; see the forest ownership contract.
Training performance priority
Optimize GPU training for large real datasets. Training-speed claims and performance-driven default changes require measurements on representative large workloads, the two real datasets of ENGINEERING_RULES.md section 9 (NYC taxi and Istella-S LETOR) at 1 million rows or more, with held-out quality and memory pressure recorded. Size is not a universal row cutoff: feature count, classes, bins, tree depth and device memory also determine the workload. Small synthetic fixtures remain useful for correctness, smoke tests and isolated diagnostics; they do not establish a large-data speed gain or justify a speed default. See the tree roadmap and GPU measurement plan.
What bitwise identity means, and why it is not the default
Floating-point addition is not associative, so the order in which a GPU sums
numbers changes the answer. Vendors choose that order differently, and they
differ again in FMA contraction, denormal handling, tie-breaking, and how
exp, log and the other elementary functions are spelled. Two GPUs given
the same job return two slightly different answers, and the difference does
not stay small. One rounding can flip a tree learner's winning split and
every node beneath it. It can redraw UMAP's neighbor graph and the embedding
built from it. Inside a training loop it perturbs a gradient, then the
optimizer state, then every step after that, so two machines running the same
job walk away with two different models.
Mojo and MAX compile one source for Metal, CUDA and HIP, which is what makes the code portable. Portability is inherited. Identity is not, and none of the above is fixed by recompiling. mojolearn supplies the part that does not come for free.
- An inventory, at the algorithm level, of every operation that can move model bits.
- A frozen numerical profile covering reduction order, partitioning, FMA policy, flush-to-zero seams, transcendental spellings and tie rules.
- Portable replacements for order-dependent reductions and for closed vendor libraries whose internals cannot be pinned.
- A per-estimator choice of
fast, same-devicedeterministic, or cross-deviceidentical, with an explicit refusal when the promise cannot be met. - Optional stage hashing and three-vendor certificates that test the promise instead of asserting it.
The kernels are mojolearn's own, in Mojo. Identical mode does not delegate to PyTorch, to MAX's matrix-multiplication kernels, or to vendor BLAS and solver libraries, because owning the arithmetic and the reduction order is the whole mechanism.
What it is for
Exact model bytes make a computation auditable. Replay a certified workload on different supported hardware, compare the recorded stage traces, and you can say where two runs first diverged with no tolerance to argue about. That is the basis for audits, regression tests and model change control, and it matters most in finance, healthcare, legal services and government, where a review can require a computation to be reproduced and its changes accounted for.
It also lets a job move. Train on rented NVIDIA capacity, continue on AMD from the checkpoint, and the run stays on the same trajectory rather than a nearby one. Hardware stops being a confounding variable in a mixed fleet.
There is a second reason to be here, independent of the contract. CatBoost, XGBoost, LightGBM and cuML have no Metal backend, so GPU tree training and GPU classical learning have not run on Apple silicon at all. One Mojo source builds for Metal, CUDA and HIP, which puts them on the laptop as well as the datacenter.
The reference has to be created and replayed under the same numerical profile.
Identity does not certify a run performed in fast mode, in another
framework, or on a device that has not passed the same checks.
The evidence behind the claim
- Neural inference and training. Mamba and transformer forward computations agree bit for bit across the three vendors on their recorded fixtures, as do gradients, optimizer updates and checkpoint bytes in fixed-shape transformer training. A two-block, 34,944-parameter byte-level language model trained on real text ran 128 steps with byte-identical parameters, gradients, optimizer state and loss on Apple Metal, NVIDIA CUDA and AMD HIP; held-out loss fell from 5.5413 to 2.8436 on all three (three-vendor record). Checkpoint continuation between NVIDIA and AMD, in both directions, preserves the uninterrupted training trajectory (cross-vendor record). Metal checkpoint resume remains open.
- Trees and classical learning. Three layers of evidence, each scoped to its commit. First, stage-level three-vendor identity cards for gradient boosting, random forests, Extra Trees, k-means, DBSCAN, k-NN, PCA, truncated SVD, OLS, ridge, logistic regression, FP32 matrix multiplication, isolation forest and ARIMA filtering, recorded at August 2026 commits (identity-path ledger); model state and recorded training stages match there, not only predictions. Second, a three-vendor prediction diff at the current default, September 12, for random forest, Extra Trees and isolation forest (45 of 45 cells equal on an Apple M4, an NVIDIA H100 and an AMD MI300X), SVC (three fit hashes match) and GBDT (36 of 36 cells on the 0.8.2 line) (AMD confirmations brief, CHANGELOG). Third, September 13, every public lane at the 0.8.4 default: 28 estimators on nine hostile fixtures, 252 cells, identical on an Apple M4, an NVIDIA H100 and an AMD MI325X, plus 189 cells of predictions on rows the model never saw and 72 cells of saved model bytes for the forest and GBDT lanes, all identical across the three (record). Fourth, the same day, CPUs with no GPU at all: random forests, Extra Trees and the four gradient boosting variants trained on each of the three GPUs save the same model bytes, and a CPU-only binding reproduces every prediction digest of all 24 recordings on seven CPUs (Intel Xeon, AMD EPYC, Azure Cobalt Neoverse-N2, Apple M1), with a sabotage build refused on each (fixtures, workflow run 34782452584).
- UMAP. Neighbor selection and iterative updates match across the three vendors on named fixtures.
Two other modes sit beside identical, selectable at runtime on the three
tree estimators only (see "Which families offer which tiers" below):
| mode | contract |
|---|---|
fast |
Optimize for throughput; repeated fits need not return identical bits. |
deterministic |
The same build, input, and device return the same bits on repeated runs. It makes no cross-vendor promise. |
identical |
Certified configurations return the same bits across Metal, CUDA, and HIP. |
Bitwise identity carries implementation and execution costs. Measured against
cuML, cuBLAS and PyTorch, identical mode is competitive on some measured
tree workloads and substantially slower on many classical, matrix and neural
workloads. Those measurements reflect both the numerical constraints and
optimization gaps in the current kernels. A paper collecting the numbers is
in preparation outside this repository; the raw records behind them live
under bench/results/.
identical is the default, in the published 0.8.4 wheels and in this
source. For the tree estimators you opt out of it, not into it, by
setting MOJOLEARN_NUMERIC_MODE=fast or deterministic in the environment
before import, or by calling mojolearn.set_numeric_mode(...) in code.
Which families offer which tiers
One rule: the tree lanes ship three tiers, everything else ships identical
only (DEVIATION 2490, 0.8.0).
| family | bindings | tiers |
|---|---|---|
| Trees: gradient boosting, random forest, extra trees | gbdt, rf, trees |
fast, deterministic, identical |
| Everything else: k-means, k-NN, PCA, truncated SVD, linear models, SVC, SVR, isolation forest, kernel density, clustering, UMAP, GP, ARIMA, preprocessing, and the whole neural surface | all others | identical only |
Asking an identical-only family for a lower tier raises a named error rather
than resolving to something weaker.
Cross-vendor bitwise identity is the product, and it is the default. A fast
tier only earns its place where it has a measured win over the opponent's own
CPU, and that is trees on Apple silicon: tree fitting calls no BLAS, so the
opponent gets nothing from Accelerate's AMX coprocessor, and extra trees
measured 1.25-1.61x scikit-learn on all ten cores at covtype 581k. The
classical families have a BLAS call in the inner loop, and on an M4 Accelerate
reaches 1438 GFLOP/s of fp32 GEMM on four performance cores against roughly
4000 for the ten-core GPU, with one CPU thread already taking 88 of the 120
GB/s the two share. A fast kernel there wins about 2.5x at best over a CPU
scikit-learn gets for free, for the price of the reproducibility guarantee.
SVC and SVR could beat libsvm's single thread, but two families with a
fast tier that are not "trees" is a rule you would have to look up, and one
rule beats two wins. The neural lanes gate every fused kernel on the identical
contract, so their lower tiers were slower than the default anyway.
What every other family offers instead is the part no vendor sells: cuML is CUDA and Linux only and does not run on Apple silicon at all, and cross-vendor bitwise identity is available nowhere else.
Who this is for
- People who need a reproducibility contract, same bits on repeated runs or
across vendors, and will pay for it in time. The cost is small on some
measured tree workloads and large elsewhere; read the records under
bench/results/before deciding. - People on Apple silicon who want GPU gradient boosting, random forests, Extra Trees, clustering, nearest neighbors, decompositions and linear models without leaving the machine.
- Not yet people training real neural networks. The certified trainers are fixed small shapes, an MLP and the two-block byte LM above. Larger models, other shapes and other optimizers are outside the evidence, and the byte-LM native trainer is not in any published wheel.
Install
python3 -m venv .venv
source .venv/bin/activate
pip install mojolearn
Version 0.8.4 is published on PyPI as an alpha API release, a macOS arm64
wheel and one Linux x86-64 wheel that now carries CUDA sm_89, CUDA sm_90 and
HIP gfx942 together, the tree bindings in all three numeric modes and every other binding in identical only, plus the
identical-mode byte-LM trainer extension per architecture. NVIDIA Linux is no
longer source-build-only. For 0.8.3 the installed Linux wheel passed its
identical qualification jobs on HIP gfx942 and CUDA sm_90a (all 29 smoke lanes
with equal hashes on both) and an installed SVC fit check on an H100; sm_89 was
not qualified installed, and the fast and deterministic qualification jobs do
not run on the 0.8 release line. The wheels expose public linalg, umap, training,
Mamba and Transformer APIs, including UMAP transform and CSR support. Newer
Python API exposure does not inherit every numerical certificate. See
CHANGELOG.md and the
support matrix for exact artifacts and limits.
There is no CPU fallback for the estimators. The byte LM is the exception, and
it has two CPU surfaces, both needing no GPU at all. LanguageModelInference
runs the forward pass; see
docs/BYTE_LM_CPU_INFERENCE.md for the CPUs it
is certified on. LanguageModelHostTrainer runs one training step, forward,
backward and the AdamW update, and reproduces the recorded GPU bytes of the
retained three-vendor capture for all 128 of its steps, the gradient and the
loss and the post-step parameters and both Adam moments alike; see
docs/BYTE_LM_CPU_TRAINING.md, which also states
what it does not claim. Both are one model profile at one batch shape, and
identity is claimed per shape because the weight gradients contract over the
token count.
Run the diagnostic command before depending on a new machine:
mojolearn doctor
The exact wheel, architecture, Python, and evidence boundaries live in SUPPORT_MATRIX.md. Source builds may support hardware outside the architectures packaged in a released wheel; that is not the same as released-wheel support.
Project status
Stability and release cadence
mojolearn went from 0.1.0 on 2026-08-23 to 0.8.4 on 2026-09-13, eleven PyPI
releases in under three weeks (0.1.0, 0.2.0, 0.3.0, 0.3.1, 0.5.0, 0.6.0,
0.7.0, 0.8.0, 0.8.1, 0.8.2, 0.8.3; 0.3.2, 0.4.0 and 0.6.1 are recorded in CHANGELOG.md but
were not published to PyPI). One release was yanked. 0.3.0, published
2026-08-30 as the first release with a Linux wheel, had been compiled for the
build machine's CPU and
carried unconditional AVX-512 instructions in its host code, so every numeric
mode died with SIGILL on any x86-64 host without AVX-512. It is yanked on PyPI
with the reason "SIGILL on x86-64 without AVX-512; use 0.3.1". 0.3.1 pinned
the Linux baseline to x86-64-v3 and added a gate on the shipped binary; the
defect and both gates are documented in
packaging/linux/isa_baseline_linux.py and packaging/wheel_ci.py.
The Python API is beta and will change between minor versions. The stable
surface is the set of numerical profiles (fast, deterministic,
identical) and the certified configurations recorded in
SUPPORT_MATRIX.md: a profile version changes only
through an explicit decision, and a numerical change must either prove itself
bit-inert or introduce a new profile version. For production or archival work
pin both the package version and the numeric profile, in code or through
MOJOLEARN_NUMERIC_MODE. A certificate names a commit, a configuration (the
fixture, the numeric profile, the parameters) and the devices it ran on, and
never more. A newer version, a different shape or an unrun vendor column is
not covered by it.
Maintenance and bus factor
The project has one maintainer today. Three things limit what that means for a reader.
The identity cards and legs cited in this README are recorded under
bench/results/, each naming its commit, device, toolchain, mode and
limitations. The per-release install qualification logs are retained outside
the tree and summarized per release in CHANGELOG.md. Each
recorded card is reproducible from the commands in the docs
(verification, conformance bundles,
release runbook). Historical cards and investigations
under bench/results/ and archive/ are evidence, not current guidance;
SUPPORT_MATRIX.md is updated only from recorded evidence.
Contributions are governed by CONTRIBUTING.md and
GOVERNANCE.md. A contributor needs one GPU of any vendor and
marks the vendor columns they did not run cross-vendor-pending; closing a
cross-vendor claim is a maintainer job. Any change that can move identical
bits must show that it is bit-inert, supply a separating fixture and a
profile-version decision, or add a named refusal. External pull requests get
an admission report and a hosted CPU report; there is no GPU automation and
no automatic merge. Governance uses lazy consensus with a seven-day objection
window, maintainership is explicitly transferable, a sole maintainer records
nominations in a public issue, and the succession steps for a sole maintainer
(nominate two successors, transfer access, document release and certification
steps, rotate credentials, publish open blockers) are written down. The code
is Apache-2.0.
You can verify a certificate without trusting the maintainer. On any
supported GPU, MOJOLEARN_NUMERIC_MODE=identical python -m mojolearn verify
runs a pinned fixture, captures its stage-level identity card and compares it
with the reference card shipped in the installation; python -m mojolearn check-fixture checks the fixture's input hashes without a GPU. Recorded
cards carry stage tags, dtypes, element counts and raw-bit hashes and are
compared with tools/identity_trace_diff.py, the one comparator the
repository uses. python -m mojolearn conformance exports and validates
bundles so another implementation can compare itself without running Mojo,
and tools/verify_umap_qualification.py rechecks retained release evidence
against a wheel without GPU work. One local run establishes one build on one
device; a cross-vendor claim needs every named leg, and the identity cards
for each leg are under bench/results/. The per-release install
qualification logs are kept outside the tree and summarized in
CHANGELOG.md.
Quick start
import numpy as np
import mojolearn
rng = np.random.default_rng(0)
X = rng.random((100_000, 20), dtype=np.float32)
y = (X[:, 0] + X[:, 1] > 1.0).astype(np.float32)
model = mojolearn.GradientBoosting(
loss="Logloss", n_estimators=200, max_depth=6,
numeric_mode="deterministic",
)
model.fit(X, y)
print(model.predict_proba(X[:5]))
print(model.numeric_mode_used(), mojolearn.vendor())
Choose a process default with mojolearn.set_numeric_mode("identical"), or
set the starting default before import:
MOJOLEARN_NUMERIC_MODE=identical python train.py
More than one tier may be loaded in one process through per-estimator
numeric_mode= arguments.
Public API
Classical estimators include:
- Gradient boosting, random forests, and Extra Trees
- K-means, nearest-neighbor estimators, DBSCAN, hierarchical and spectral clustering
- PCA, truncated SVD, linear and logistic regression, ridge, lasso, and elastic net
- SVC, SVR, kernel density, isolation forest, and Gaussian-process regression
- Exponential smoothing and batched ARIMA
- UMAP embeddings with dense Euclidean input and 2D/3D spectral initialization; version 0.6.0 adds unseen-sample transformation and CSR graph storage
Additional modules provide scoring metrics, FP32 matrix multiplication, optimizer/training primitives, and reference-pinned Mamba and transformer blocks. These surfaces do not all have the same validation depth; consult the support matrix before treating an experimental surface as release-qualified.
UMAP in the 0.5.0 API supports fitting and embedding the supplied samples:
X = np.array([0, 1, 2.2, 4, 6.5, 10, 14.5, 20], dtype=np.float32)[:, None]
embedding = mojolearn.UMAP(
n_neighbors=3, n_components=2, n_epochs=4, random_state=19,
numeric_mode="identical",
).fit_transform(X)
The 0.5.0 implementation stores a dense graph and does not support
transform. In 0.6.0, public fitting stores
the graph in CSR form, using O(n_samples × n_neighbors) graph space, and
transform(X_new) embeds unseen samples against a frozen fitted model. Input
remains a dense Euclidean array; CSR describes internal graph storage.
Exact neighbor search still performs quadratic pair comparisons.
Source checks for the integrated fit/transform API passed all three numeric modes on Apple, NVIDIA and AMD. The named IDENTICAL held-out embeddings match across all three vendors. The macOS 0.6.0 candidate also passed clean installed fit/transform and quality checks. See the version-specific evidence.
Transformation retains private training data and embedding copies. Changing parameters or numeric mode requires refitting, and changing query batching can change results. Supervised targets, alternate metrics and alternate initialization remain unsupported.
The APIs intentionally resemble scikit-learn, but mojolearn is not a drop-in replacement. Where an algorithm has a settled convention for a default, that convention is followed. Unsupported parameters raise explicitly rather than being silently ignored.
The exact scope of the claim
Fix a source commit, a supported configuration, a seed and byte-identical
input. On any two certified machines, every recorded training stage has the
same bits, and either model produces exactly the same predictions. This is a
claim about the trained model, not byte-for-byte equality of archive
metadata. If a configuration cannot meet the contract, the library raises a
named error instead of silently returning a possibly different model; a
refusal is reported as a refusal, never counted as a pass. The claim is
float32 only. Float64 input is converted with a copy for most estimators
(python/mojolearn/_buffer.py) and refused by name on the linalg
(python/mojolearn/_linalg_impl.py) and Mamba (python/mojolearn/_mamba_impl.py)
surfaces. The cross-vendor identity cells behind the September 12 diff are
recorded at fixtures of up to 20,000 rows by 16 columns
(tools/identity_break.py), not at the 1M-row speed workloads.
Cross-vendor identity is a profile, not a statement that every GPU operation is universally identical. A profile fixes relevant reduction order, partitioning, FMA policy, flush-to-zero seams, transcendental spellings, and tie rules. A numerical change must either prove bit-inertness against the profile or introduce a new profile version. Additional devices must pass the same identity checks; the guarantee covers only devices and configurations that have.
The project distinguishes four artifact classes:
source check -> Python binding -> built native artifact -> installed wheel
Evidence for one class does not automatically validate the next. Current
certificates, configurations, and outstanding vendor legs are listed in
SUPPORT_MATRIX.md. Historical cards and investigations
under bench/results/ and archive/ are evidence, not current guidance.
Limitations
What will get in your way first:
- GPU hardware is required for every estimator and for training. There is no CPU fallback for them, and the library refuses rather than silently running elsewhere. The one exception is byte LM inference, which runs on certified CPUs through its own explicitly named class.
- mojolearn is not a drop-in replacement for scikit-learn, CatBoost or cuML. Parameter coverage is intentionally smaller than any of them, and unsupported parameters raise.
- Source builds need the Mojo toolchain through pixi, and one build targets one GPU architecture. NVIDIA Linux is source-build-only today.
- The support matrix is honest about gaps. Several public surfaces still have vendor legs or independent-reference checks pending, and an unrun column is pending, never inferred.
And the standing limits of the contract itself:
- Released-wheel support is narrower than source-build support.
fastdeliberately makes no repeatability promise, and is built only for the three tree families (DEVIATION 2490).deterministicdoes not promise agreement between different devices.identicalcovers certified profiles and fixtures, not arbitrary untested shapes or future toolchains.- Some recent Python and neural-operator surfaces still have vendor legs or independent-reference checks pending.
- Parameter coverage is intentionally smaller than scikit-learn, CatBoost, or cuML.
- The experimental k-NN selector remains behind an explicit build flag; normal wheel builds retain the existing dispatch.
mojolearn is beta software. Pin the package version and numerical profile for production or archival work.
Development
Start with docs/START_HERE.md. The shortest full local
check is pixi run probe.
A numerical test counts as evidence only after a separating arm demonstrates that it fails when the relevant rule is broken. Contributors need one supported GPU; maintainers close cross-vendor certification columns.
Current priorities are in ROADMAP.md. See also verification, release, engineering rules, contributing, governance, and notices.
Citation
Every line of Mojo in this repository was written for it. The library implements published machine-learning algorithms, and where a specific published formulation is followed closely enough that a reader would want the reference, the source file names it. The numerical contract that is the project's distinguishing result has no counterpart anywhere.
To cite mojolearn, use CITATION.cff. The concept DOI is 10.5281/zenodo.22068632.
Release files for mojolearn 0.8.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mojolearn-0.8.5-py3-none-manylinux_2_35_x86_64.whl | Python 3 | none | Linux glibc 2.35+ x86-64 | Details |
| mojolearn-0.8.5-py3-none-macosx_11_0_arm64.whl | Python 3 | none | macOS 11.0+ ARM64 | Details |
Total release size: 76.8 MB
Release files / mojolearn-0.8.5-py3-none-manylinux_2_35_x86_64.whl
| Download URL | mojolearn-0.8.5-py3-none-manylinux_2_35_x86_64.whl |
|---|---|
| Size | 56.3 MB |
| Tags | Linux glibc 2.35+ x86-64 Python 3 |
|
SHA-256 checksum How to use checksums |
9c415b87a3cbbb4777cc34093d882e738bab4264a1dba92a95b515b3a41e62b3
|
|
BLAKE2b-256 checksum How to use checksums |
99a8e8093ce4bcfb9b0fc6062b0958f9ef21efdacf0d51d5f74ce9e325c09ee7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency logRelease files / mojolearn-0.8.5-py3-none-macosx_11_0_arm64.whl
| Download URL | mojolearn-0.8.5-py3-none-macosx_11_0_arm64.whl |
|---|---|
| Size | 20.4 MB |
| Tags | Python 3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
6d8332d8c77716b767e260b128f722e0e06cbd3c470f0b2b7e2281ef9237deea
|
|
BLAKE2b-256 checksum How to use checksums |
1b4502fd50ddca3be3bd288f0a813ff199244bb9f89ee34b2af0abb5575144d3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.
Transparency log