Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

mila-llm

The LLM stack you can read and understand.

Mila is a C++23/CUDA runtime for open large language models — inference and training, built from explicit neural-network components. Device and precision are compile-time decisions, every forward pass is explicit, and there is no hidden execution engine. This package is its Python projection.

pip install mila-llm
import mila

mila.initialize("warning")

# Once: fetch a published model into the local store (~6.3 GB).
mila.ModelStore().pull("gemma-4-12b-it-fp4", mila.default_hub_owner())

tokenizer = mila.BpeTokenizer.from_store("gemma-4-12b-it-fp4")
model = mila.GemmaModel.from_store("gemma-4-12b-it-fp4", 4096)

reason = model.generate(tokenizer.encode(prompt), print)

A model is named, not pathed. from_store reads the local store's record, which is what knows the weights are already FP4 — so nothing pairs a weights path with a tokenizer path, and nothing has to be told what the bytes are. Pull and load are separate verbs: a load never reaches the network, so an uninstalled name is an error rather than a surprise download.

generate hands each token to the callback as it is produced and returns why it stopped — stop, length, context_limit or cancelled — which the tokens themselves cannot tell you.

The GIL is released around generation, so the callback runs on a live interpreter and StopController cancels a decode loop already in flight.

Requirements

An NVIDIA GPU. The CUDA runtime libraries arrive as dependencies (nvidia-cublas, nvidia-curand, and nvidia-cuda-runtime on Linux, where the extension links cudart dynamically) — no CUDA Toolkit installation is required. An installed Toolkit is used as a fallback if those are absent.

gemma-4-12b-it-fp4 wants roughly 12 GB of VRAM at a 4096 context; a Llama 3.2 3B is the smaller first run.

What it exposes

Symbol Members
mila.initialize log_level = trace | info | warning | error
mila.BpeTokenizer from_store(name), load_llama32, load_gemma, load_qwen, encode, decode, token_to_string, is_valid_token, vocab_size, bos_token_id, eos_token_id, pad_token_id
mila.GemmaModel from_store(name, context_length, device_index=0), from_pretrained(path, context_length, device_index=0, quantization="fp4"), generate(prompt_tokens, on_token, ...), get_config
mila.LlamaModel from_store(name, context_length, device_index=0), from_pretrained(path, context_length, device_index=0, quantization="bf16"), generate(prompt_tokens, on_token, ...), get_config
mila.QwenModel from_store(name, context_length, device_index=0), from_pretrained(path, context_length, device_index=0, quantization="fp4"), generate(prompt_tokens, on_token, ...), get_config
mila.qwen_format_prompt (history, enable_thinking=False, reasoning_effort=3, tools_json="") — the runtime's own Qwen 3.8 template
mila.qwen_parse_tool_call (response) → {call id, name, arguments} or None
mila.qwen_protocol_tokens Qwen's control tokens, for a caller that streams
mila.ModelStore root, list, locate, remove, usage, install, pull, list_hub_models
mila.StopController request_stop, stop_requested

The store is shared with Mila's chat harness and inference server, so a model installed by any of them is loadable by all of them.

What it does not

Stated because the limits are documentation, not an omission from it.

  • No weights are bundled. They are fetched on request into a local store, over a transport this package supplies from the standard library — the wheel carries no HTTP client of its own.
  • A load never downloads. pull and from_store are separate calls on purpose.
  • No GPT-2. It exists in the C++ library and is not bound.
  • A published model's quantization is fixed — its bytes are already FP4 or FP8. Choosing a quantization applies only to unquantized weights loaded by path.
  • No training, no batching, and a model instance is not thread-safe: serialize calls through a single worker thread.
  • Text in, text out. No embeddings, logits, or hidden-state access.

Links

Metadata

Release files for mila-llm 0.20.0b3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for mila-llm 0.20.0b3
File
mila_llm-0.20.0b3-cp313-cp313-win_amd64.whl CPython 3.13 CPython 3.13 Windows x86-64 Details
mila_llm-0.20.0b3-cp313-cp313-manylinux_2_38_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.38+ x86-64 Details
mila_llm-0.20.0b3-cp312-cp312-win_amd64.whl CPython 3.12 CPython 3.12 Windows x86-64 Details
mila_llm-0.20.0b3-cp312-cp312-manylinux_2_38_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.38+ x86-64 Details

Total release size: 36.3 MB

Release files / mila_llm-0.20.0b3-cp313-cp313-win_amd64.whl

Download URL mila_llm-0.20.0b3-cp313-cp313-win_amd64.whl
Size 8.5 MB
Tags CPython 3.13 Windows x86-64
SHA-256 checksum
How to use checksums
f3b9b85486f7d0e991baf55dd0fcbf3d74f73f7625dce6d777bd0bb483bc99e9
BLAKE2b-256 checksum
How to use checksums
af226db21e2eedb95c82e3f12baa593678dfb76ed9bbce10e84fd154f9fc8f13
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / mila_llm-0.20.0b3-cp313-cp313-manylinux_2_38_x86_64.whl

Download URL mila_llm-0.20.0b3-cp313-cp313-manylinux_2_38_x86_64.whl
Size 9.7 MB
Tags CPython 3.13 Linux glibc 2.38+ x86-64
SHA-256 checksum
How to use checksums
78ba43b808347178a8c7c0c009ee4354befd2068e5895c637727d77864af4933
BLAKE2b-256 checksum
How to use checksums
cb0367d64056ac0e06bfce6bb361d10c5c4af631d69e30907bcdf92b26a9a356
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / mila_llm-0.20.0b3-cp312-cp312-win_amd64.whl

Download URL mila_llm-0.20.0b3-cp312-cp312-win_amd64.whl
Size 8.5 MB
Tags CPython 3.12 Windows x86-64
SHA-256 checksum
How to use checksums
4bf6e3bd27f7f263f01eca136f683f58b3660365b6cb03cc9deeb90e213f3e37
BLAKE2b-256 checksum
How to use checksums
c621b84f93a08160d0c0e386a65c870de1c818a06ed4b925bb22368b1635f015
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / mila_llm-0.20.0b3-cp312-cp312-manylinux_2_38_x86_64.whl

Download URL mila_llm-0.20.0b3-cp312-cp312-manylinux_2_38_x86_64.whl
Size 9.7 MB
Tags CPython 3.12 Linux glibc 2.38+ x86-64
SHA-256 checksum
How to use checksums
2343e8dc315a4d116f18f96d12750e71f432c29ec1605e31fdf5233c0be5392b
BLAKE2b-256 checksum
How to use checksums
90c4f15e36c2b729e352db900cbe5420f507a79de3827467e15d71448a6eec3d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page