memopro
Run work larger than your device's memory, with unchanged results, within a memory ceiling you set.
한국어 · English
memopro is an open-source library for work that needs more memory than the machine has. It is verified on 8-16 GB laptops and Macs and small GPUs; the design does not depend on the memory size (design, in Korean). The core is written in Rust and the interface in Python.
- Lossless: no quantization or approximation. Generation and inference outputs are bit-identical to a plain run; training losses are bit-identical whatever the budget.
- Guaranteed ceiling: memory use stays within the budget you set. A budget that cannot work is refused before anything runs.
- No disk writes: model weights are re-read from their original files and data created in memory is compressed losslessly. No swap or cache files are created.
import memopro
r = memopro.finetune("Qwen/Qwen2.5-7B-Instruct", texts, budget="1.5GiB") # 16-bit LoRA on an 8 GB Mac
r.adapter.save_pretrained("my-lora") # a standard PEFT adapter
print(memopro.generate(r.model, "Hello!", draft="Qwen/Qwen2.5-1.5B-Instruct"))
"Memory" means hardware memory (GPU and RAM), not agent or conversation memory.
Results
All measured against criteria fixed in advance. Raw data and environments are in docs/research/data/; the scripts are in experiments/.
Language models (MacBook Air M1, 8 GB)
| Task | Result |
|---|---|
| Qwen2.5-7B bf16 LoRA training (14.2 GiB of weights) | Completes with a 1.5 GiB budget; losses bit-identical across budgets; 7.3 tokens/s |
| Qwen2.5-3B bf16 LoRA, same setup as mlx-tune | mlx-tune runs out of memory before step 1; memopro trains at 20.0 tokens/s with a 1 GiB budget |
| Qwen2.5-7B lossless generation (int4 draft + row-invariant verification) | Same output as plain generation, 8.95 → 2.37 s/token (3.8x) |
| Qwen2.5-3B bf16 inference (CPU) at 1/4 of the memory it needs | Same output, about 5x faster than OS paging |
| Qwen2.5-7B bf16 LoRA training on a Colab T4 (15 GB) | Plain training runs out of GPU memory; memopro completes with a 4 GiB budget, losses identical across budgets |
| Qwen2.5-3B bf16 generation vs. other tools | memopro 1.09 s/token (2.1 GB process); llama.cpp CPU 14.75 s/token (3.3 GB); llama.cpp Metal runs out of memory |
Ordinary programs (Linux, limit = 1/2 of the memory they need)
| Task | Result |
|---|---|
| Unmodified NumPy image processing | Same result, 3.3x faster than OS swap at the same limit |
| Unmodified scikit-learn classification | Same result (a plain run is killed for lack of memory) |
One-line setup memopro.enable / memopro run (MacBook Air M1, whole-process ceiling)
| Task | Result |
|---|---|
| Unmodified NumPy image processing, ceiling = 1/2 of what it needs | Identical results, peak within ceiling + 1.8 MiB, 2.4x the time |
| 1 GiB array, 4 passes, ceiling = 1/2 and 3/4 of what it needs | Identical results; extra time predicted before running 3.81 s vs 3.72 s measured (3/4: 1.73 vs 1.86 s) |
| Cost of leaving it on when memory is ample | 1.02-1.08x |
On Linux (Colab) the results were identical as well, the ceiling held (0 MiB over) and the estimate was within 10%. Keeping it on when memory was ample cost 1.20x on Linux.
Vision models (Colab T4, inference with streamed weights)
| Model | Result |
|---|---|
| ResNet-152 (CNN) | Output bit-identical to a plain run at 60 and 120 MiB budgets, 0.13 -> 0.14 s |
| DINOv2-giant (1.1B parameters, 4.5 GB) | Identical output at 1.1 and 2.2 GB budgets, 1.57 -> 1.65 s |
Dataframe workloads with heavy random access, and data that is still larger than the limit after compression, do not yet run at a practical speed at 1/2 (Limitations).
Installation
Install from PyPI (alpha: APIs may change). Wheels are built for macOS arm64 and Linux x86_64/aarch64; elsewhere pip builds the source distribution, which needs a Rust toolchain.
pip install "memopro[llm]"
The latest development version: pip install "memopro[llm] @ git+https://github.com/imhyensuk/memopro" (needs a Rust toolchain).
| Extra | For |
|---|---|
memopro[torch] |
PyTorch integration (load, train_session, hibernation) |
memopro[hf] |
Loading Hugging Face models |
memopro[llm] |
finetune, generate (includes PEFT) |
memopro[notebook] |
Jupyter magics |
Usage
Full guide: docs/guide
0. One line
import memopro
s = memopro.enable() # measure this machine: ceiling = in use + free, minus headroom
s = memopro.enable(budget="8GB") # or set the whole-process ceiling (8GB of a 16GB machine)
print(s.estimate(sample, total="12GB", passes=3)) # before running: predicted slowdown
... # ordinary NumPy / PyTorch code
print(s.measured()) # afterwards: time memopro took, real slowdown
- Measures the hardware (OS, memory, GPU), sets one ceiling for the whole process and turns on what this platform supports; print the returned session to see what was applied or skipped.
- Every part keeps to the ceiling together: each runtime and pager shrinks its share by the memory the process holds outside it (macOS physical footprint, Linux RSS).
- NumPy (Linux, macOS): arrays of 16 MiB or more are paged within the ceiling, losslessly compressed, without code changes.
- PyTorch: once the process passes 75% of the ceiling, activations saved for backward move into runtime buffers; the bytes come back unchanged, so gradients do not change (checked on CPU and MPS).
finetune,generate,load,train_sessionandmemopro.rtuse the ceiling unless told otherwise. Hugging Facefrom_pretrainedloads likememopro.loadonly when the model does not fit as stored.memopro.disable()(orwith memopro.enable(...):) undoes it.- Limits: memory nothing can move (the interpreter, libraries, small objects, model weights the code loaded itself) counts too; when it alone passes the ceiling, the pager records overruns and runtimes raise
BudgetExceeded. Kernel I/O on paged arrays (ndarray.tofile,np.fromfile) may raiseOSError;np.save/np.loadare routed around it. Predictions assume passes in a fixed order.
1. Train and generate with LLMs larger than memory
import memopro
r = memopro.finetune(
"Qwen/Qwen2.5-3B-Instruct", texts,
budget="1GiB", # memory ceiling for the weights
seq_len=512, rank=8, # LoRA on q/k/v/o
)
print(r.losses, r.tokens / r.seconds)
text = memopro.generate(r.model, "Summarize: ...", draft="Qwen/Qwen2.5-1.5B-Instruct")
- 16-bit weights stream layer by layer from the original safetensors files, handed to the Apple GPU without copies.
- Activation memory is planned inside the budget; a budget that cannot hold one step is refused before training starts.
- With
draft, a small int4 draft model speculates; verification follows the same computation path as plain generation, so the output does not change. - Verified on Apple silicon (MPS), NVIDIA CUDA (Colab T4) and CPU.
2. Run local models (vision and language models, inference)
Any model downloaded in Hugging Face format (safetensors) can be loaded and run as is, even if it is larger than memory.
import transformers
import memopro.rt.torch as rtt
m = rtt.stream_model("facebook/dinov2-giant", budget="1GB", device="cuda", # "mps", "cpu"
model_class=transformers.Dinov2Model)
features = m(pixel_values=images).last_hidden_state # same output as the model loaded normally
- Weights stay in their files and are read within the budget as each layer computes. Download the model first (
huggingface-cli download ...). - For language models use
memopro.generate, for LoRA trainingmemopro.finetune(section 1). For models that fit,memopro.loadpicks a form within the budget (quantized if needed). - Training your own PyTorch models:
memopro.enable()moves saved activations within the ceiling. Streaming the weights themselves is supported for Hugging Face models only.
3. Large arrays within a budget
from memopro.rt import Runtime
rt = Runtime(budget="2GB")
weights = rt.load_npy("big.npy") # from a file: dropped and re-read when needed (hash-checked)
work = rt.array((50_000, 4_096), "float32") # new buffer: compressed losslessly when needed
with work.view(write=True) as a: # in memory only while used, as a zero-copy NumPy view
a[:] = 1.0
print(rt.stats())
Objects you already hold (tensors, modules, optimizers, KV caches, dicts of NumPy arrays) can be handed to the runtime as well.
h = rt.adopt(state) # managed by the runtime while idle (compressed losslessly when needed)
with h: # used as usual inside the block
step(state)
4. Run scripts without changing them
memopro run --budget 6GB train.py # loads from_pretrained models that do not fit within the budget
memopro run --transparent 1GB analysis.py # Linux, macOS: pages large NumPy arrays with compression, no disk writes
memopro run --dry-run train.py # show what it would do
5. Load and train within a budget
model, tok = memopro.load("Qwen/Qwen2.5-7B-Instruct", tokenizer=True, quality="high")
with memopro.train_session(model, optimizer) as s: # micro-batching and checkpointing, exact methods first
for batch in loader:
s.step(batch, lambda mb: model(**mb).loss)
memopro doctor # available memory and budget per device, RAM, disk
memopro check Qwen/Qwen2.5-7B-Instruct --batch-size 4 --seq-len 512 # predict memory before running
qualitybounds the loss allowed automatically:"lossless"<"high"<"balanced"(default) <"low".- When nothing fits,
BudgetExceededlists settings that would actually load.
Budget forms
The same forms work in every API, the CLI (--budget), memopro.toml and MEMOPRO_BUDGET.
| Form | Meaning |
|---|---|
"auto" |
measured available memory minus 10% headroom (default) |
"6GB", 0.5, "50%" |
a cap, or a fraction of the measured value |
"-2GB" |
leave 2 GB of the measured value free |
"2GB..6GB" |
at most 6 GB; do not run if 2 GB cannot be had |
"6GB!" |
exactly 6 GB regardless of measurement (swap risk accepted) |
{"device": "80%", "host": "-2GB"} |
per memory pool |
Architecture
Python API finetune · generate · load · train_session · run · adopt
─────────────────────────────────────────────────────────
Access layer detect environment → budget → choose a configuration → apply → report measurements
(existing techniques such as quantization, offloading and checkpointing are wrapped)
─────────────────────────────────────────────────────────
Rust runtime per buffer, chosen by measured cost:
keep · compress losslessly · re-read the source (hash-checked) · recompute
+ prefetching, ceiling guarantee, slowdown prediction
─────────────────────────────────────────────────────────
Platform zero-copy Apple GPU buffers · asynchronous CUDA copies · Linux userfaultfd
| Component | Location |
|---|---|
Rust core (crates.io memopro) |
crates/memopro |
| C ABI | crates/memopro-c |
Linux allocation interposer (LD_PRELOAD) |
crates/memopro-preload |
| Python package | python/memopro |
Supported platforms
| Platform | Status |
|---|---|
| macOS, Apple silicon (MPS) | Primary platform; LLM training and generation verified; transparent paging (signals) |
| Linux, NVIDIA GPU (CUDA) | Generation, LoRA training, vision inference (ResNet, DINOv2), enable ceiling and estimate verified on a Colab T4 |
| Linux, CPU | Verified in CI; transparent paging (userfaultfd) |
| Windows | Basic features checked in CI |
Python 3.11+, PyTorch 2.4+.
Limitations
- Speed: memory is saved at the cost of time. 7B training on an 8 GB Mac takes about 18 s per 129-token step. Models that fit in memory run faster with existing tools: on a Colab T4, Unsloth (fp16) trained 3B and 7B 16-bit LoRA 10-33x faster than memopro (stored bf16). memopro is for models larger than the GPU or device memory, which do not run otherwise.
- Transparent paging: workloads with heavy random access (sorting, group-by) and data still larger than the limit after compression do not run at a practical speed. Linux and macOS (not Windows).
- Model coverage: the LLM path is verified mainly on the Qwen2.5 family (1.5B-7B), vision on ResNet-152 and DINOv2. Diffusion and speech models are not tested.
- Alpha: APIs may change.
Development
python3 -m venv .venv && .venv/bin/pip install "maturin>=1.9,<2" pytest ruff
VIRTUAL_ENV=$PWD/.venv .venv/bin/maturin develop --release
.venv/bin/pytest -q
cargo test -p memopro
Experiment scripts are in experiments/, raw results and environments in docs/research/data/, design documents in docs/design/ (in Korean).
Citation
If you use memopro in research, please cite it with "Cite this repository" (CITATION.cff).
License
MIT or Apache-2.0, at your option.
Metadata
Release files for memopro 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| memopro-0.1.0.tar.gz | 229.0 kB | Details |
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| memopro-0.1.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.11 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| memopro-0.1.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl | CPython 3.11 | abi3 | Linux glibc 2.17+ ARM64 | Details |
| memopro-0.1.0-cp311-abi3-macosx_11_0_arm64.whl | CPython 3.11 | abi3 | macOS 11.0+ ARM64 | Details |
Total release size: 3.2 MB
Release files / memopro-0.1.0.tar.gz
| Download URL | memopro-0.1.0.tar.gz |
|---|---|
| Size | 229.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
03e24a369c22419f78c695e251320bcd85aa4f47fa633f890b786f3b25f21e23
|
|
BLAKE2b-256 checksum How to use checksums |
63cedcb2bf324c174f3ea54db867c790c1dee9a970af2724db4918132cde76c1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / memopro-0.1.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | memopro-0.1.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 1.1 MB |
| Tags | CPython 3.11 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
d787d263fd214c08350d4fd2360ebce1bcc6c8b4671b6c9b22f1bf21f21b753c
|
|
BLAKE2b-256 checksum How to use checksums |
1c024583614c6293b5619697bb17b11fcb2202e88b95a6fce186f085c5572ec1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / memopro-0.1.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
| Download URL | memopro-0.1.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl |
|---|---|
| Size | 1.0 MB |
| Tags | CPython 3.11 Linux glibc 2.17+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
9c13a8a44082bf114e4a02afdccae1c1a1220b6f88729c03eb4ab566475958e5
|
|
BLAKE2b-256 checksum How to use checksums |
436b50c84d9afc9e2357f31912b34f24c97b5207d411f980b251acb63f4d060c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / memopro-0.1.0-cp311-abi3-macosx_11_0_arm64.whl
| Download URL | memopro-0.1.0-cp311-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 869.5 kB |
| Tags | CPython 3.11 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
111832a2c03e269e9363dbc8b144bb7a608e874fb2ca277236c6e3d560520710
|
|
BLAKE2b-256 checksum How to use checksums |
b48b5dbd8d43a6c02d146aeda65201195b370d7c56c377fe24fc5fcab9a13d09
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log