mila-llm
The LLM stack you can read and understand.
Mila is a C++23/CUDA runtime for open large language models — inference and training, built from explicit neural-network components. Device and precision are compile-time decisions, every forward pass is explicit, and there is no hidden execution engine. This package is its Python projection.
pip install mila-llm
import mila
mila.initialize("warning")
# Once: fetch a published model into the local store (~6.8 GB).
mila.ModelStore().pull("gemma-4-12b-it-fp4", mila.default_hub_owner())
tokenizer = mila.BpeTokenizer.from_store("gemma-4-12b-it-fp4")
model = mila.GemmaModel.from_store("gemma-4-12b-it-fp4", 4096)
reason = model.generate(tokenizer.encode(prompt), print)
A model is named, not pathed. from_store reads the local store's record, which is
what knows the weights are already FP4 — so nothing pairs a weights path with a
tokenizer path, and nothing has to be told what the bytes are. Pull and load are
separate verbs: a load never reaches the network, so an uninstalled name is an
error rather than a surprise download.
generate hands each token to the callback as it is produced and returns why it
stopped — stop, length, context_limit or cancelled — which the tokens
themselves cannot tell you.
The GIL is released around generation, so the callback runs on a live interpreter
and StopController cancels a decode loop already in flight.
Requirements
An NVIDIA GPU. The CUDA runtime libraries arrive as dependencies
(nvidia-cublas, nvidia-curand, and nvidia-cuda-runtime on Linux, where the
extension links cudart dynamically) — no CUDA Toolkit installation is required.
An installed Toolkit is used as a fallback if those are absent.
gemma-4-12b-it-fp4 wants roughly 12 GB of VRAM at a 4096 context; a Llama 3.2 3B
is the smaller first run.
What it exposes
| Symbol | Members |
|---|---|
mila.initialize |
log_level = trace | info | warning | error |
mila.BpeTokenizer |
from_store(name), load_llama32, load_gemma, load_qwen, encode, decode, token_to_string, is_valid_token, vocab_size, bos_token_id, eos_token_id, pad_token_id |
mila.GemmaModel |
from_store(name, context_length, device_index=0), load(path, context_length, device_index=0, quantization="fp4"), generate(prompt_tokens, on_token, ...), get_config |
mila.LlamaModel |
from_store(name, context_length, device_index=0), load(path, context_length, device_index=0, quantization="bf16"), generate(prompt_tokens, on_token, ...), get_config |
mila.QwenModel |
from_store(name, context_length, device_index=0), load(path, context_length, device_index=0, quantization="fp4"), generate(prompt_tokens, on_token, ...), get_config |
mila.qwen_format_prompt |
(history, enable_thinking=False, reasoning_effort=3, tools_json="") — the runtime's own Qwen 3.8 template |
mila.qwen_parse_tool_call |
(response) → {call id, name, arguments} or None |
mila.qwen_protocol_tokens |
Qwen's control tokens, for a caller that streams |
mila.ModelStore |
root, list, locate, remove, usage, install, pull, list_hub_models |
mila.StopController |
request_stop, stop_requested |
The store is shared with Mila's chat harness and inference server, so a model installed by any of them is loadable by all of them.
What it does not
Stated because the limits are documentation, not an omission from it.
- No weights are bundled. They are fetched on request into a local store, over a transport this package supplies from the standard library — the wheel carries no HTTP client of its own.
- A load never downloads.
pullandfrom_storeare separate calls on purpose. - No GPT-2. It exists in the C++ library and is not bound.
- A published model's quantization is fixed — its bytes are already FP4 or FP8. Choosing a quantization applies only to unquantized weights loaded by path.
- No training, no batching, and a model instance is not thread-safe: serialize calls through a single worker thread.
- Text in, text out. No embeddings, logits, or hidden-state access.
Links
- Documentation: https://mila.toddt.me
- Source: https://github.com/toddthomson/Mila
- Models: https://huggingface.co/mila-llm — what
pullfetches from
Metadata
Release files for mila-llm 0.20.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| mila_llm-0.20.0-cp313-cp313-win_amd64.whl | CPython 3.13 | CPython 3.13 | Windows x86-64 | Details |
| mila_llm-0.20.0-cp313-cp313-manylinux_2_38_x86_64.whl | CPython 3.13 | CPython 3.13 | Linux glibc 2.38+ x86-64 | Details |
| mila_llm-0.20.0-cp312-cp312-win_amd64.whl | CPython 3.12 | CPython 3.12 | Windows x86-64 | Details |
| mila_llm-0.20.0-cp312-cp312-manylinux_2_38_x86_64.whl | CPython 3.12 | CPython 3.12 | Linux glibc 2.38+ x86-64 | Details |
Total release size: 38.5 MB
Release files / mila_llm-0.20.0-cp313-cp313-win_amd64.whl
| Download URL | mila_llm-0.20.0-cp313-cp313-win_amd64.whl |
|---|---|
| Size | 8.9 MB |
| Tags | CPython 3.13 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
a5587abb644255764b524c812841853363fcd7bb378a7c9a2fd774c58f953272
|
|
BLAKE2b-256 checksum How to use checksums |
977d2148acd34d4287fd7d7c355f096f23736771bcac19db369dc147c5848651
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|
Release files / mila_llm-0.20.0-cp313-cp313-manylinux_2_38_x86_64.whl
| Download URL | mila_llm-0.20.0-cp313-cp313-manylinux_2_38_x86_64.whl |
|---|---|
| Size | 10.3 MB |
| Tags | CPython 3.13 Linux glibc 2.38+ x86-64 |
|
SHA-256 checksum How to use checksums |
006296661dca986d0eced9010dcb22974a885f1165651ee7dc6310ced8381d5d
|
|
BLAKE2b-256 checksum How to use checksums |
00a14d6167548447bc4d456cc0ce7eaa12e5ad6dd5e412d161f66a75d2708a09
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|
Release files / mila_llm-0.20.0-cp312-cp312-win_amd64.whl
| Download URL | mila_llm-0.20.0-cp312-cp312-win_amd64.whl |
|---|---|
| Size | 8.9 MB |
| Tags | CPython 3.12 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
9c175c87ab36383cf220c2a5b06056e68c5c93f11a7c922a65b87b21d3ee1103
|
|
BLAKE2b-256 checksum How to use checksums |
0097846337a0178e21a8cdc3af9dd3bb2282809cc7806537914d9f114eb73190
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|
Release files / mila_llm-0.20.0-cp312-cp312-manylinux_2_38_x86_64.whl
| Download URL | mila_llm-0.20.0-cp312-cp312-manylinux_2_38_x86_64.whl |
|---|---|
| Size | 10.3 MB |
| Tags | CPython 3.12 Linux glibc 2.38+ x86-64 |
|
SHA-256 checksum How to use checksums |
285a10fb6d85e7faff78cd04f539e497a57dbf545bcf236740ff846b99fb277d
|
|
BLAKE2b-256 checksum How to use checksums |
028af81af70451da443f0ef763c63a9d90db0946e543d172aeb51b8ad9b330de
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|