mixedlm-rs fits Gaussian linear mixed-effects models using lme4's profiled
ML/REML formulation and a compiled Rust core. It provides a statsmodels-style
formula and array API, with no R installation required.
Use it for repeated measurements within subjects, patients within clinics, or other data with one grouping factor, including correlated random intercepts and slopes. The implementation is designed to make models with many groups fast while checking convergence explicitly.
This is beta software. It supports a subset of statsmodels.MixedLM;
crossed effects, multiple levels of nesting and GLMMs are outside its scope.
See compatibility
before migrating an existing analysis.
Install
python -m pip install mixedlm-rs
Python 3.10 or later is required. NumPy, SciPy, pandas and Patsy are installed as dependencies. Neither statsmodels nor R is needed for normal use.
Binary wheels are available on PyPI:
| Platform | Architectures |
|---|---|
| Windows | x86-64 |
| macOS | Intel x86-64, Apple Silicon ARM64 |
| Linux, glibc 2.17+ | x86-64, ARM64 |
The wheels use Python's stable ABI (cp310-abi3). On a matching platform,
installation needs no Rust compiler. Building from source requires Rust 1.83+
and a compatible native build toolchain; pip installs the Maturin build backend.
There are no published wheels for 32-bit systems or musl-based Linux.
Currently verified: the project's CI matrix covers Python 3.10–3.14 on Windows, macOS and Linux, with separate checks for the declared dependency floors and minimum Rust version. Release checks install each platform's wheel outside the source tree and test a wheel rebuilt from the source archive. See the CI runs for the status of a particular commit. Python versions beyond 3.14 are not yet part of that test matrix.
Quick start
This example simulates 18 subjects measured on 10 days. Each subject has their own baseline and rate of change. It runs after installation, with no data file to fetch.
import numpy as np
import pandas as pd
import mixedlm_rs as mlm
rng = np.random.default_rng(0)
n_subjects, n_days = 18, 10
subject = np.repeat(np.arange(n_subjects), n_days)
day = np.tile(np.arange(n_days), n_subjects)
intercept = rng.normal(250, 25, n_subjects)[subject]
slope = rng.normal(10, 6, n_subjects)[subject]
reaction = intercept + slope * day + rng.normal(0, 25, subject.size)
data = pd.DataFrame({"reaction": reaction, "day": day, "subject": subject})
model = mlm.mixedlm(
"reaction ~ day",
data,
groups="subject",
re_formula="~day",
)
result = model.fit() # REML by default
print(result.summary())
print(f"fixed effects: {result.fe_params}")
print(f"converged: {result.converged}")
print(f"singular: {result.singular}")
reaction ~ day specifies the fixed effects. groups="subject" identifies
independent clusters; re_formula="~day" gives each subject a random intercept
and slope, with their covariance estimated. Omit re_formula for a random
intercept only. Group sizes do not have to be equal.
To use maximum likelihood instead, call .fit(reml=False) on a new model.
Use ML when comparing likelihoods of models with different fixed-effect
specifications fitted to the same observations; their REML criteria are not
directly comparable.
Inspect estimates and predict
Continuing the example:
print(result.fe_params_labelled) # Fixed effects as a named pandas Series
print(result.bse_fe) # Fixed-effect standard errors
print(result.cov_re) # Random-effect covariance matrix
print(result.scale) # Residual variance
print(result.diagnostics) # Optimizer and numerical diagnostics
new_data = pd.DataFrame({"day": [0, 5, 10]})
print(result.predict(new_data))
predict() returns the population-level, fixed-effects prediction. It does
not add fitted subject effects, so this example needs no subject column.
result.fittedvalues includes the fitted random effects for the training
observations. result.random_effects contains the estimated effects by group.
The main numerical attributes, including fe_params and cov_re, are NumPy
arrays even for formula fits. Use fe_params_labelled or params_labelled
when you need names.
Use arrays directly
The equivalent model can be constructed without a formula:
X = np.column_stack([np.ones(len(data)), data["day"].to_numpy()])
array_result = mlm.MixedLM(
endog=data["reaction"].to_numpy(),
exog=X,
groups=data["subject"].to_numpy(),
exog_re=X,
).fit()
Include an intercept column explicitly in array inputs when the model needs
one. exog defines fixed effects; exog_re defines random effects within each
group.
Save and load a fit
result.save("fit.pkl")
restored = mlm.MixedLMResults.load("fit.pkl")
print(restored.predict(new_data))
The default save retains the training data needed to rebuild supported formula designs. Only load files you trust: this uses Python pickle. Custom formula functions and reduced-data saves have additional persistence limitations.
Moving from statsmodels
For a supported model, start by changing the formula API import:
- import statsmodels.formula.api as smf
+ import mixedlm_rs as smf
The mixedlm(...), MixedLM(...) and MixedLM.from_formula(...) entry points
follow familiar statsmodels conventions. Existing code still needs review:
some return types differ, optimizer arguments do not select the same
algorithms, and several methods are unimplemented. The
compatibility guide
lists the supported contract, warnings and exceptions.
Convergence and inference
Inspect both result.converged and result.singular:
| Attribute | Meaning |
|---|---|
converged |
The fit passed the implementation's numerical convergence checks. |
singular |
The estimated random-effect covariance is at or near a singular boundary. |
A singular fit can be a valid converged optimum, for example when a random
intercept variance is zero. Conversely, returning estimates does not establish
convergence. This package can also return converged=False with a warning.
The optimizer checks local stationarity and probes away from boundary
solutions. These checks do not prove that it found a global optimum or that
the model is identifiable. Examine warnings and result.diagnostics before
using an uncertain fit.
Reported fixed-effect tests use normal/Wald inference. There are no Satterthwaite or Kenward–Roger small-sample corrections. Fixed-effect covariance is conditional on the fitted variance parameters; joint covariance across fixed and variance parameters is not available. Wald intervals for variance parameters are especially unreliable near zero.
Numerical validation
The primary reference is lme4, with checks against published fits and
recorded R results for sleepstudy, Dyestuff and Dyestuff2, under ML and
REML. These cover correlated random slopes and a zero-variance boundary fit.
The reference datasets are GPL-2 test fixtures kept in the Git repository;
they are not shipped in the wheels or source archive. See
dataset provenance.
Agreement is assessed at the precision and tolerances of the recorded references. Additional tests compare against statsmodels, check derivatives with finite differences, and probe the objective independently of the analytic gradient. A convergence warning from another package alone does not establish that its estimates are wrong.
The recorded comparison on 120 randomised fixtures reports:
| Outcome | Count |
|---|---|
| statsmodels did not converge | 26 |
| mixedlm-rs found a strictly better optimum | 42 |
| same optimum | 52 |
| mixedlm-rs found a worse optimum | 0 |
“Better” and “worse” refer to the fitted likelihood criterion on the same model and data, using the experiment's absolute deviance tolerance. They do not mean better prediction or establish which model is scientifically appropriate.
A separate 400-case adversarial sweep recorded 99 statsmodels
nonconvergences, 131 better optima for mixedlm-rs, 167 ties and 3 worse
optima. Those losses have absolute deviance gaps of approximately
2.74e-6 to 5.71e-5. One case also failed the sweep's independent stationarity
check. These are finite test suites, not estimates of failure rates in general use.
Both experiments are recorded in
bench/baseline.json
at commit 55c20ac. The
correctness report
provides the reference fits, thresholds, generators and known difficult cases.
Performance
The recorded benchmark below uses synthetic data with a random intercept and slope, fitted through each package's formula API on the same machine. Timings include model construction and fitting; they exclude data generation and printing a summary.
| Observations | Groups | statsmodels | mixedlm-rs | Speedup |
|---|---|---|---|---|
| 2,000 | 100 | 0.325 s | 0.0082 s | 40x |
| 10,000 | 500 | 1.655 s | 0.0091 s | 183x |
| 20,000 | 1,000 | 2.543 s | 0.0111 s | 229x |
| 40,000 | 5,000 | 11.019 s | 0.0167 s | 661x |
| 100,000 | 20,000 | 41.873 s | 0.0384 s | 1092x |
| 200,000 | 50,000 | 102.791 s | 0.0732 s | 1404x |
| 500,264 | 125,066 | Not measured | 0.218 s | — |
Measurement conditions: Windows 11, AMD Ryzen 5 5600G, 12 logical CPUs, Python 3.14.3 and statsmodels 0.15.0. mixedlm-rs reports the minimum of three runs. statsmodels also uses three runs through 40,000 observations, but only one run for the 100,000- and 200,000-observation cases. Thread counts were not pinned: Rayon used the available logical CPUs, and OMP/OpenBLAS thread settings were unset. These are wall-clock comparisons with those defaults, not comparisons at a fixed CPU budget. Speedups use unrounded timings.
Both fits reported convergence in every timed comparison, and their estimates and likelihoods passed the benchmark's agreement checks. The workload favors many small groups; speedups depend on model shape, data, dependencies and hardware, and should not be assumed for every analysis.
The measurements were recorded on 2026-09-10 at commit b34a88a.
bench/performance.json
contains all repetitions, dependency versions, thread settings and the
extension hash. Reproduce the comparison with
bench/performance.py.
The largest case uses a group count motivated by statsmodels issue #9097. Its reported 41-minute timing is not a measurement made here. Our synthetic case is not a reproduction of their dataset and should not be read as a measured speedup over that report.
Historical R/lme4 and pymer4 timings are presented separately in the benchmark report. They were not rerun for this measurement. The pymer4 path also performs lmerTest inference and result extraction, so its elapsed time covers different work from a bare model fit.
How it works
The Rust core evaluates a profiled ML/REML criterion using penalised least squares. With one grouping factor, the random-effect system separates into small per-group blocks. The implementation precomputes group statistics, factorises those blocks in Rust, and uses Rayon for parallel work.
An analytic gradient reuses the factorisation work. The default optimizer is SciPy's L-BFGS-B, calling the Rust objective and gradient. In the recorded implementation-stage comparison, replacing this project's finite differences with its analytic gradient reduced objective evaluations from 44–64 to 11–13. That is a comparison between stages of this implementation, not a claim that analytic gradients are new or that statsmodels lacks them.
See the design and gradient derivation for the equations, implementation details and validation.
Scope and limitations
| Area | Current boundary |
|---|---|
| Grouping | One grouping factor. No crossed subject/item effects or separate effects at multiple nesting levels. |
| Model family | Gaussian linear mixed models only; no binomial or Poisson GLMMs. |
| Covariance structures | No variance-component formulas, free masks or covariance penalties. |
| Additional fitting methods | No fit_regularized, profile_re, bootstrap or get_distribution. |
| Fixed-effect design | Rank-deficient or numerically rank-deficient designs are rejected; redundant columns are not dropped automatically. |
| Inference | Normal/Wald inference, with the limitations described above. |
| Optimization | Local convergence checks; no guarantee of a global optimum. method="rust" is experimental. |
These restrictions matter for designs such as crossed participant/item experiments or students within classrooms within schools. Combining IDs into one grouping variable does not reproduce separate random effects at each level. Read the full limitations and API compatibility guide for argument handling, result types, numerical edge cases and persistence.
Development
Building a checkout requires Rust 1.83+ and Python 3.10+. Use a virtual environment, then install the package and test tools:
git clone https://github.com/Fatin-Ishraq/mixedlm-rs.git
cd mixedlm-rs
python -m pip install ".[test,lint]"
python -m pytest tests/ -q
cargo test --lib --locked
python -m ruff check .
The Git checkout includes the numerical reference fixtures and benchmark scripts omitted from published distributions. CI also runs type checking, Rust linting, dependency-floor checks and artifact validation.
For bug reports, include a minimal example, package and dependency versions,
platform, warnings and result.diagnostics when a fit is available. Use
GitHub issues.
See the changelog
for release history.
License and acknowledgements
mixedlm-rs is MIT licensed. Bundled dependencies retain their own licenses; see third-party notices and license texts.
The statistical formulation follows lme4 and the work of Bates, Mächler, Bolker and Walker, Fitting Linear Mixed-Effects Models Using lme4 (2015). The Python API follows statsmodels. This is an independent implementation; neither upstream project maintains or endorses it.
Metadata
Release files for mixedlm-rs 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mixedlm_rs-0.1.1.tar.gz | 266.1 kB | Details |
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| mixedlm_rs-0.1.1-cp310-abi3-win_amd64.whl | CPython 3.10 | abi3 | Windows x86-64 | Details |
| mixedlm_rs-0.1.1-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.10 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| mixedlm_rs-0.1.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl | CPython 3.10 | abi3 | Linux glibc 2.17+ ARM64 | Details |
| mixedlm_rs-0.1.1-cp310-abi3-macosx_11_0_arm64.whl | CPython 3.10 | abi3 | macOS 11.0+ ARM64 | Details |
| mixedlm_rs-0.1.1-cp310-abi3-macosx_10_12_x86_64.whl | CPython 3.10 | abi3 | macOS 10.12+ x86-64 | Details |
Total release size: 2.0 MB
Release files / mixedlm_rs-0.1.1.tar.gz
| Download URL | mixedlm_rs-0.1.1.tar.gz |
|---|---|
| Size | 266.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
967bf928608610c658d5429e148d2aa37a5df57b655e3e63f961b4a725644312
|
|
BLAKE2b-256 checksum How to use checksums |
d67187e6ce691b3a013c22f495013daa8ac4a883d63d533988a434d0203a6a88
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / mixedlm_rs-0.1.1-cp310-abi3-win_amd64.whl
| Download URL | mixedlm_rs-0.1.1-cp310-abi3-win_amd64.whl |
|---|---|
| Size | 253.0 kB |
| Tags | CPython 3.10 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
fcc4b13307bd0cfe7e21d69bd44f2c92f6179487fd43115b46df29f808433207
|
|
BLAKE2b-256 checksum How to use checksums |
5e21484775940100f0ca4a769c1af771d6506ed5f245430b47a66a0359c85f9c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / mixedlm_rs-0.1.1-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | mixedlm_rs-0.1.1-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 388.3 kB |
| Tags | CPython 3.10 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
03e96ac3bcd4fb64b55713d2337502ed88b34826eaa08be4f044ef9c60ddf800
|
|
BLAKE2b-256 checksum How to use checksums |
643649325f6d3c7c0a67a055cc884bcbddd0c8591685994df18f0bf390cfa500
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / mixedlm_rs-0.1.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
| Download URL | mixedlm_rs-0.1.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl |
|---|---|
| Size | 385.9 kB |
| Tags | CPython 3.10 Linux glibc 2.17+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
74686d07d4cfcc97b05560b88db6b7cc4a7885deeacd09b150db2c7f1efbca83
|
|
BLAKE2b-256 checksum How to use checksums |
54b89d3c86d3cc6f74d9a1aafdcda0b0fc16197535e357611b766aadadf74fcc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / mixedlm_rs-0.1.1-cp310-abi3-macosx_11_0_arm64.whl
| Download URL | mixedlm_rs-0.1.1-cp310-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 348.7 kB |
| Tags | CPython 3.10 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
45d0420e13fdcabfde442ba7562d4f261c8f2e91b14e0c750d633fec3325913c
|
|
BLAKE2b-256 checksum How to use checksums |
de45b58abe56b29682a7f30d7f5e864d08af2b1bd42ce514eb9435d579c59fe7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / mixedlm_rs-0.1.1-cp310-abi3-macosx_10_12_x86_64.whl
| Download URL | mixedlm_rs-0.1.1-cp310-abi3-macosx_10_12_x86_64.whl |
|---|---|
| Size | 358.1 kB |
| Tags | CPython 3.10 abi3 macOS 10.12+ x86-64 |
|
SHA-256 checksum How to use checksums |
1544d11ea306f209e2d24cedef59317f83a64ed2c353456b8de74dad32dcee68
|
|
BLAKE2b-256 checksum How to use checksums |
5e6484a6135bd3fcb48e421bd46297e2b4d267cbbb285f40d41adf3660956609
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency log