Skip to main content

mila-llm

The LLM stack you can read and understand.

Mila is a C++23/CUDA runtime for open large language models — inference and training, built from explicit neural-network components. Device and precision are compile-time decisions, every forward pass is explicit, and there is no hidden execution engine. This package is its Python projection.

pip install mila-llm
import mila

mila.initialize("warning")

# Once: fetch a published model into the local store (~6.8 GB).
mila.ModelStore().pull("gemma-4-12b-it-fp4", mila.default_hub_owner())

tokenizer = mila.BpeTokenizer.from_store("gemma-4-12b-it-fp4")
model = mila.GemmaModel.from_store("gemma-4-12b-it-fp4", 4096)

reason = model.generate(tokenizer.encode(prompt), print)

A model is named, not pathed. from_store reads the local store's record, which is what knows the weights are already FP4 — so nothing pairs a weights path with a tokenizer path, and nothing has to be told what the bytes are. Pull and load are separate verbs: a load never reaches the network, so an uninstalled name is an error rather than a surprise download.

generate hands each token to the callback as it is produced and returns why it stopped — stop, length, context_limit or cancelled — which the tokens themselves cannot tell you.

The GIL is released around generation, so the callback runs on a live interpreter and StopController cancels a decode loop already in flight.

Requirements

An NVIDIA GPU. The CUDA runtime libraries arrive as dependencies (nvidia-cublas, nvidia-curand, and nvidia-cuda-runtime on Linux, where the extension links cudart dynamically) — no CUDA Toolkit installation is required. An installed Toolkit is used as a fallback if those are absent.

gemma-4-12b-it-fp4 wants roughly 12 GB of VRAM at a 4096 context; a Llama 3.2 3B is the smaller first run.

What it exposes

Symbol Members
mila.initialize log_level = trace | info | warning | error
mila.BpeTokenizer from_store(name), load_llama32, load_gemma, load_qwen, encode, decode, token_to_string, is_valid_token, vocab_size, bos_token_id, eos_token_id, pad_token_id
mila.GemmaModel from_store(name, context_length, device_index=0), load(path, context_length, device_index=0, quantization="fp4"), generate(prompt_tokens, on_token, ...), get_config
mila.LlamaModel from_store(name, context_length, device_index=0), load(path, context_length, device_index=0, quantization="bf16"), generate(prompt_tokens, on_token, ...), get_config
mila.QwenModel from_store(name, context_length, device_index=0), load(path, context_length, device_index=0, quantization="fp4"), generate(prompt_tokens, on_token, ...), get_config
mila.qwen_format_prompt (history, enable_thinking=False, reasoning_effort=3, tools_json="") — the runtime's own Qwen 3.8 template
mila.qwen_parse_tool_call (response) → {call id, name, arguments} or None
mila.qwen_protocol_tokens Qwen's control tokens, for a caller that streams
mila.ModelStore root, list, locate, remove, usage, install, pull, list_hub_models
mila.StopController request_stop, stop_requested

The store is shared with Mila's chat harness and inference server, so a model installed by any of them is loadable by all of them.

What it does not

Stated because the limits are documentation, not an omission from it.

  • No weights are bundled. They are fetched on request into a local store, over a transport this package supplies from the standard library — the wheel carries no HTTP client of its own.
  • A load never downloads. pull and from_store are separate calls on purpose.
  • No GPT-2. It exists in the C++ library and is not bound.
  • A published model's quantization is fixed — its bytes are already FP4 or FP8. Choosing a quantization applies only to unquantized weights loaded by path.
  • No training, no batching, and a model instance is not thread-safe: serialize calls through a single worker thread.
  • Text in, text out. No embeddings, logits, or hidden-state access.

Metadata

Release files for mila-llm 0.20.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for mila-llm 0.20.0
File
mila_llm-0.20.0-cp313-cp313-win_amd64.whl CPython 3.13 CPython 3.13 Windows x86-64 Details
mila_llm-0.20.0-cp313-cp313-manylinux_2_38_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.38+ x86-64 Details
mila_llm-0.20.0-cp312-cp312-win_amd64.whl CPython 3.12 CPython 3.12 Windows x86-64 Details
mila_llm-0.20.0-cp312-cp312-manylinux_2_38_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.38+ x86-64 Details

Total release size: 38.5 MB

Release files / mila_llm-0.20.0-cp313-cp313-win_amd64.whl

Download URL mila_llm-0.20.0-cp313-cp313-win_amd64.whl
Size 8.9 MB
Tags CPython 3.13 Windows x86-64
SHA-256 checksum
How to use checksums
a5587abb644255764b524c812841853363fcd7bb378a7c9a2fd774c58f953272
BLAKE2b-256 checksum
How to use checksums
977d2148acd34d4287fd7d7c355f096f23736771bcac19db369dc147c5848651
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / mila_llm-0.20.0-cp313-cp313-manylinux_2_38_x86_64.whl

Download URL mila_llm-0.20.0-cp313-cp313-manylinux_2_38_x86_64.whl
Size 10.3 MB
Tags CPython 3.13 Linux glibc 2.38+ x86-64
SHA-256 checksum
How to use checksums
006296661dca986d0eced9010dcb22974a885f1165651ee7dc6310ced8381d5d
BLAKE2b-256 checksum
How to use checksums
00a14d6167548447bc4d456cc0ce7eaa12e5ad6dd5e412d161f66a75d2708a09
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / mila_llm-0.20.0-cp312-cp312-win_amd64.whl

Download URL mila_llm-0.20.0-cp312-cp312-win_amd64.whl
Size 8.9 MB
Tags CPython 3.12 Windows x86-64
SHA-256 checksum
How to use checksums
9c175c87ab36383cf220c2a5b06056e68c5c93f11a7c922a65b87b21d3ee1103
BLAKE2b-256 checksum
How to use checksums
0097846337a0178e21a8cdc3af9dd3bb2282809cc7806537914d9f114eb73190
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / mila_llm-0.20.0-cp312-cp312-manylinux_2_38_x86_64.whl

Download URL mila_llm-0.20.0-cp312-cp312-manylinux_2_38_x86_64.whl
Size 10.3 MB
Tags CPython 3.12 Linux glibc 2.38+ x86-64
SHA-256 checksum
How to use checksums
285a10fb6d85e7faff78cd04f539e497a57dbf545bcf236740ff846b99fb277d
BLAKE2b-256 checksum
How to use checksums
028af81af70451da443f0ef763c63a9d90db0946e543d172aeb51b8ad9b330de
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page