Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

 loom-py

import loom

model = loom.Model.from_pretrained("loom-ai-org/lfm2-350m-monolithic-loom")
print(model.generate("The capital of France is", max_new_tokens=14))
# ':\nA) Paris\nB) Lyon\nC) Marseille\nD'

Text in, text out

generate tokenizes with the vocabulary the GGUF embeds, runs the driver, and detokenizes what comes back. The same steps are available separately when you want them:

model.tokenize("The capital of France is")   # [1, 1098, 5706, 803, 4481, 856]
model.detokenize([1, 1098, 5706])            # '<|startoftext|>The capital'
model.tokenizer                               # <loom.Tokenizer 'gpt2' size=64400>

The four vocabulary families a loom GGUF can carry — byte-level BPE, SentencePiece, WordPiece and byte-level — are dispatched on the file's own tokenizer.ggml.model, so this is one call whichever one a model uses.

For a speech model there is nothing to encode; detokenizing the driver's output is the other half of the same thing:

transcript = model.detokenize(model.infer(waveform=audio, audio_samples=len(audio)))

A TTS model has no generate, and that is a real limitation rather than a missing feature. Matcha, VITS, Kokoro and StyleTTS2 consume phoneme ids that a phonemiser produces outside the engine, so their GGUFs embed no vocabulary at all — model.tokenizer is None for them and they take ids directly:

audio = model.infer(tokens=[16, 40, 22, 30, 12, 3], n_steps=4, seed=1234)

Why there is so little API

A loom GGUF carries its own graph topologies and its own driver script alongside its weights, so this package contains no per-architecture code at all. Loading a model registers whatever topologies the file declares and attaches a KV cache to the ones that say they need it; running one calls the driver the file shipped with. A model this library has never heard of works the day loom-exporter can produce it.

That is also why infer takes **kwargs: its arguments are the driver's arguments, and which ones a model takes is a property of the model. model.driver_source prints the Lua that will run, whose header comment documents its inputs — that is the authority.

model = loom.Model.from_file("granite_speech_mil.gguf")
model.architecture          # 'granite-speech'
model.topologies            # ['encoder', 'embed', 'decoder', 'lm_head']
model.hparam("samples_per_chunk")   # 192000
print(model.driver_source)  # what infer() will run, and what it accepts

Supported models

Seventeen, published at huggingface.co/loom-ai-org and loadable by id with from_pretrained (needs the [hub] extra). This package has no per-architecture code, so the list is a property of loom-exporter, not of anything here.

Language models

Model Exported from
loom-ai-org/qwen3-0.6b-base-loom Qwen/Qwen3-0.6B-Base
loom-ai-org/lfm2-350m-monolithic-loom LiquidAI/LFM2-350M
loom-ai-org/lfm2-350m-modular-loom LiquidAI/LFM2-350M
loom-ai-org/smollm2-360m-instruct-loom HuggingFaceTB/SmolLM2-360M-Instruct
loom-ai-org/gemma-3-270m-it-loom google/gemma-3-270m-it

These are the models generate works on.

Speech recognition

Model Exported from
loom-ai-org/whisper-small-loom openai/whisper-small
loom-ai-org/conformer-ctc-small-loom nvidia/stt_en_conformer_ctc_small
loom-ai-org/parakeet-tdt-0.6b-loom nvidia/parakeet-tdt-0.6b-v3
loom-ai-org/parakeet-rnnt-0.6b-loom nvidia/parakeet-rnnt-0.6b
loom-ai-org/gigaam-v3-rnnt-loom ai-sage/GigaAM-v3
loom-ai-org/qwen3-asr-0.6b-loom Qwen/Qwen3-ASR-0.6B
loom-ai-org/granite-speech-4.0-1b-loom ibm-granite/granite-4.0-1b-speech

model.detokenize(model.infer(waveform=audio, audio_samples=len(audio))) — the mel frontend is inside the graph, so a raw waveform is the input.

Speech synthesis

Model Exported from
loom-ai-org/kokoro-82m-loom hexgrad/Kokoro-82M
loom-ai-org/matcha-tts-ljspeech-loom Matcha-TTS (LJSpeech checkpoint)
loom-ai-org/supertonic-2-loom Supertone/supertonic-2
loom-ai-org/vits-piper-en-gb-miro-loom OpenVoiceOS/pipertts_en-GB_miro
loom-ai-org/styletts2-ljspeech-loom yl4579/StyleTTS2-LJSpeech

Supertonic is the one with a text door — model.tokenize works on it and is None for the other four, which take phoneme ids (see above). model.driver_source is the authority on what each accepts.

The three repos

loom.cpp the engine, vendored here as a submodule
loom-exporter produces the GGUFs this runs
loom-py this one

Installing

pip install loom-py-rt           # once published -- `loom-py` on PyPI clashes with `loompy`,
                                  # and `loom-engine` normalizes to the already-taken `loomengine`
pip install loom-py-rt[hub]      # + from_pretrained()

From a checkout — note --recursive, since the engine is a submodule:

git clone --recursive https://github.com/loom-ai-org/loom-py
cd loom-py && pip install -e .

No runtime dependencies. Arrays cross the boundary as plain sequences of floats, so numpy is something you may use rather than something this package makes you install — list, array.array, numpy arrays and torch tensors all work.

Testing

pytest tests/ci      # the Python layer: coercion and error paths. No model. What CI runs.
pytest tests/gate    # a real exported GGUF, end to end.
export LOOM_TEST_MODEL=~/loom-fixtures/matcha_mil.gguf
export LOOM_TEST_MODEL_INPUTS='{"tokens":[16,40,22,30,12,3],"n_steps":4,"seed":1234}'
pytest tests/gate -q

The gate suite is written against no particular architecture on purpose: it asserts the shape of what a loom model is, which is the whole of what this package knows. A test that expected one model's inputs would be this package learning about a model, which is the thing the design exists to avoid.

Roadmap

Shared with loom.cpp, because three of the four are the engine's and this package inherits them by having no per-architecture code of its own.

1. GPUs and NPUs. The engine talks to a single ggml_backend_t and uses no ggml_backend_sched, which is what has to change before a second device can hold part of a graph. Nothing in this package should need to change with it — device selection is a property of how the engine is built.

2. Wheels for more platforms. Linux x86-64 today; next macOS on Intel, macOS on Apple Silicon and Linux on ARM. This is the item most visible from here, since it is what pip install loom-py-rt can resolve to.

3. More models — P5 in the ledger, ordered by coverage per unit of effort: BERT token classifiers (the smallest possible template, and the first non-audio task) → codec decoders → CNN+CTC and SANM encoders → the remaining TTS families → text encoder-decoders → small classifiers → music. Each lands here for free: a model this package has never heard of works the day the exporter can produce it.

4. The follow-ups the docs already nameBACKLOG.md is the ledger for all three repos and the authority. The ones that would show up in this API: a permissively-licensed phonemiser, which is what would give the four phoneme-input TTS models a tokenize; and the KvCache memory redesign and quantized KV cache, which decide how large a model this can run on a given machine.

Licence

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

loom_py_rt-1.0.0rc1.tar.gz (1.6 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

loom_py_rt-1.0.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

loom_py_rt-1.0.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

loom_py_rt-1.0.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

loom_py_rt-1.0.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

loom_py_rt-1.0.0rc1-cp39-cp39-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.9manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

File details

Details for the file loom_py_rt-1.0.0rc1.tar.gz.

File metadata

  • Download URL: loom_py_rt-1.0.0rc1.tar.gz
  • Upload date:
  • Size: 1.6 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.2

File hashes

Hashes for loom_py_rt-1.0.0rc1.tar.gz
Algorithm Hash digest
SHA256 9c043ec0fa123cf65007a0ef853579897cd0d9f7114dc8f954c5d37e63e53b1a
MD5 492cf65f1efc36228eedb6a34c624d22
BLAKE2b-256 3ea178d15a46245d70982553a9bd6e32389d79a254bc9e4c1f238c58b13a1d75

See more details on using hashes here.

File details

Details for the file loom_py_rt-1.0.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for loom_py_rt-1.0.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 eb31b5a6af9b1675ee3a673adc11abf54587562f91f3763fcf775019b0363d18
MD5 ddb79188703fdc815f208cf8d8ed53a6
BLAKE2b-256 211fa702ea98ff63402612acc13ab935acc13673bc48d4780ed471d34f8f7d1d

See more details on using hashes here.

File details

Details for the file loom_py_rt-1.0.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for loom_py_rt-1.0.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 8cf3cb6060face7d64b5de9d490612cc12a229919596862ab5bec4fd8b2142b9
MD5 78eca546880daf9b70dda717db1763be
BLAKE2b-256 de57bdf6ae6f05c8683e560962b17a475555199876f091b4697f454258735d9f

See more details on using hashes here.

File details

Details for the file loom_py_rt-1.0.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for loom_py_rt-1.0.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 6eb369785c1cf624b939cfc6f2984d1e942d8e2a3aee58512055379140d95add
MD5 3f69284dba49d2693488784439f1482e
BLAKE2b-256 97b8a49a8b08edaaa5a2a6a6db996fa895d294a1a1dde2add23fd17db7053928

See more details on using hashes here.

File details

Details for the file loom_py_rt-1.0.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for loom_py_rt-1.0.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 41c77288a759db5e4b5c037cc14f46f3366bdd9e11ee691637d3dada046dcc8b
MD5 0651a99daa734b917eccf5dd61ab1dee
BLAKE2b-256 7618e1d4e31c81535d11bad893128d2fae53de9e5cb156428e8e989c6d2ab5e8

See more details on using hashes here.

File details

Details for the file loom_py_rt-1.0.0rc1-cp39-cp39-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for loom_py_rt-1.0.0rc1-cp39-cp39-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 79191a771b13b16dc266375b3f34c3fafb0de8ad9a579c2b4a5e3a3254e2e81e
MD5 062ec64457dc243733a039e22e7ee751
BLAKE2b-256 0a404f108058994e9c32943bc7d29ba9f941378d90673636ca39313300e31744

See more details on using hashes here.

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page