transformers-mblt
Run Mobilint pre-quantized generative AI models on Mobilint NPUs through Hugging Face
Transformers. transformers-mblt supplies the
Mobilint configuration, model, cache, processor, and generation classes behind the standard
transformers Auto classes and pipeline(...). It covers LLMs, VLMs, speech recognition,
image captioning, masked language models, and EAGLE-3 speculative decoding.
Models run on Mobilint ARIES and
REGULUS boards. Supported target-device identifiers are
aries-rb, regulus-ra, regulus-rb, regulus-ra-usb, and regulus-rb-usb.
Version 0.0.0 is the first standalone release. It was extracted from mblt-model-zoo 2.10.0.
Installation
pip install transformers-mblt
The following are installed as required dependencies:
transformers[serving]>=4.54.0,<5.18.0mblt-npu-python, the shared Mobilint NPU backendmobilint-qb-runtime
NPU execution requires a supported Mobilint NPU driver and device. Qwen3-ASR additionally needs the
upstream qwen-asr package:
pip install "transformers-mblt[qwen-asr]"
Quick start
Mobilint models are published on the Mobilint Hugging Face organization.
Call transformers_mblt.register() once, and the standard transformers Auto classes and pipeline(...) load them
from the installed package without running Hub remote code:
import transformers_mblt
from transformers import AutoTokenizer, TextStreamer, pipeline
transformers_mblt.register()
model_id = "mobilint/Llama-3.2-1B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
pipe = pipeline(
"text-generation",
model=model_id,
tokenizer=tokenizer,
streamer=TextStreamer(tokenizer=tokenizer, skip_prompt=False),
model_kwargs={"core_mode": "single"},
)
messages = [{"role": "user", "content": "What is an NPU?"}]
pipe(messages, max_new_tokens=128)
pipe.model.dispose()
Loading through Hub remote code
Each mobilint/* repository also ships proxy_*.py remote code, loaded with trust_remote_code=True. The proxies
import transformers_mblt first and fall back to mblt_model_zoo.hf_transformers, so this path works with either
package installed.
NPU placement is controlled with keyword arguments. The main ones are mxq_path, dev_no,
core_mode, target_cores, target_clusters, target_device, revision, embedding_weight,
and npu_prefill_chunk_size. Multi-backend models accept the same arguments with vision_,
text_, encoder_, decoder_, base_, or draft_ prefixes. The
API reference describes each argument.
Use list_tasks() and list_models() to discover the supported tasks and published models:
from transformers_mblt import list_models, list_tasks
print(list_tasks())
print(list_models("text-generation"))
Supported architectures
| Task | Architectures |
|---|---|
text-generation |
Llama, Qwen2, Qwen3, EXAONE 3.5, EXAONE 4.0, Cohere2, EAGLE-3 (Llama, Qwen2, Qwen3) |
image-text-to-text |
Qwen2-VL, Qwen3-VL, Aya Vision (SigLIP + Cohere2) |
automatic-speech-recognition |
Whisper, Qwen3-ASR |
image-to-text |
BLIP |
fill-mask |
BERT |
Command line
The transformers-mblt command provides these subcommands:
listshows the published models for each task.tpsmeasures tokens per second.- Upstream Transformers commands such as
chat,serve,run,download,env, andversionare passed through with the Mobilint models registered.
transformers-mblt list --task text-generation
transformers-mblt chat mobilint/Llama-3.2-1B-Instruct --trust-remote-code
transformers-mblt tps measure --model mobilint/Llama-3.2-1B-Instruct --prefill 512 --decode 128 --repeat 10
transformers-mblt tps sweep --model mobilint/Llama-3.2-1B-Instruct \
--prefill-range 128:512:128 --cache-lengths 1024,2048,4096 --decode-window 128 --json tps.json
python -m transformers_mblt.cli is equivalent to transformers-mblt. Run
transformers-mblt tps measure --help for the EAGLE-3, VLM, and per-backend core-mode options.
Using with mblt-model-zoo
This package replaces mblt_model_zoo.hf_transformers, and mblt-model-zoo is moving to depend on it. The
modules have the same contents, and only the import prefix differs:
| mblt-model-zoo | transformers-mblt |
|---|---|
mblt_model_zoo.hf_transformers.models.<arch> |
transformers_mblt.models.<arch> |
mblt_model_zoo.hf_transformers.utils |
transformers_mblt.utils |
mblt-model-zoo tps ... / mblt-model-zoo chat ... |
transformers-mblt tps ... / transformers-mblt chat ... |
Hub proxy modules first import transformers_mblt. If that package is missing, they fall back to
mblt_model_zoo.hf_transformers, so the same Hub repositories work with either package.
Development
Use uv to manage the development environment:
git clone https://github.com/mobilint/transformers-mblt.git
cd transformers-mblt
uv venv --python 3.12
source .venv/bin/activate
uv pip install -e . --group dev
pre-commit install
scripts/test_transformers_matrix.py uses uv to rebuild the environment for each supported
transformers release line and run the test phases against it.
Documentation and tests
- The API reference covers model loading, NPU keyword parameters, the Qwen3-VL release contract, the EAGLE-3 policies, and the TPS CLI.
- The test guide explains the quick, full-matrix, and core-mode sweep test runs.
- The benchmark guide covers the text-generation, VLM, and ASR benchmark scripts.
Support and issues
For installation, model, or runtime support, visit the Mobilint forum or contact tech-support@mobilint.com. Report reproducible package issues in the transformers-mblt issue tracker.
License
Distributed under the BSD 3-Clause License.
Metadata
Release files for transformers-mblt 0.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| transformers_mblt-0.0.0.tar.gz | 280.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| transformers_mblt-0.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 605.6 kB
Release files / transformers_mblt-0.0.0.tar.gz
| Download URL | transformers_mblt-0.0.0.tar.gz |
|---|---|
| Size | 280.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4df3387d4b6688542796dbd1f675234af9fed7a0855f963785bdbe300cc18b78
|
|
BLAKE2b-256 checksum How to use checksums |
952af2e445bc122dfde6d7b3a71fc6e69bbd3bdde60ed53cfadaa0ac356a3f2a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / transformers_mblt-0.0.0-py3-none-any.whl
| Download URL | transformers_mblt-0.0.0-py3-none-any.whl |
|---|---|
| Size | 325.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1f7133b9f2b2827c9e02d5625a5be93d3bde2c47e380a336cfdf2b2cfeddf306
|
|
BLAKE2b-256 checksum How to use checksums |
b2b761570daebe19d1fa562476e31d91bd8ff26acb496217909886614c3a26ab
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log