This release is a pre-release and may not be stable for production use.
loom-py
import loom
model = loom.Model.from_pretrained("loom-ai-org/lfm2-350m-monolithic-loom")
print(model.generate("The capital of France is", max_new_tokens=14))
# ':\nA) Paris\nB) Lyon\nC) Marseille\nD'
Text in, text out
generate tokenizes with the vocabulary the GGUF embeds, runs the driver, and detokenizes what comes
back. The same steps are available separately when you want them:
model.tokenize("The capital of France is") # [1, 1098, 5706, 803, 4481, 856]
model.detokenize([1, 1098, 5706]) # '<|startoftext|>The capital'
model.tokenizer # <loom.Tokenizer 'gpt2' size=64400>
The four vocabulary families a loom GGUF can carry — byte-level BPE, SentencePiece, WordPiece and
byte-level — are dispatched on the file's own tokenizer.ggml.model, so this is one call whichever
one a model uses.
For a speech model there is nothing to encode; detokenizing the driver's output is the other half of the same thing:
transcript = model.detokenize(model.infer(waveform=audio, audio_samples=len(audio)))
A TTS model has no generate, and that is a real limitation rather than a missing feature. Matcha,
VITS, Kokoro and StyleTTS2 consume phoneme ids that a phonemiser produces outside the engine, so
their GGUFs embed no vocabulary at all — model.tokenizer is None for them and they take ids
directly:
audio = model.infer(tokens=[16, 40, 22, 30, 12, 3], n_steps=4, seed=1234)
Why there is so little API
A loom GGUF carries its own graph topologies and its own driver script alongside its weights, so this package contains no per-architecture code at all. Loading a model registers whatever topologies the file declares and attaches a KV cache to the ones that say they need it; running one calls the driver the file shipped with. A model this library has never heard of works the day loom-exporter can produce it.
That is also why infer takes **kwargs: its arguments are the driver's arguments, and which ones
a model takes is a property of the model. model.driver_source prints the Lua that will run, whose
header comment documents its inputs — that is the authority.
model = loom.Model.from_file("granite_speech_mil.gguf")
model.architecture # 'granite-speech'
model.topologies # ['encoder', 'embed', 'decoder', 'lm_head']
model.hparam("samples_per_chunk") # 192000
print(model.driver_source) # what infer() will run, and what it accepts
Supported models
Seventeen, published at huggingface.co/loom-ai-org and loadable
by id with from_pretrained (needs the [hub] extra). This package has no per-architecture code, so
the list is a property of loom-exporter, not of
anything here.
Language models
These are the models generate works on.
Speech recognition
model.detokenize(model.infer(waveform=audio, audio_samples=len(audio))) — the mel frontend is inside
the graph, so a raw waveform is the input.
Speech synthesis
Supertonic is the one with a text door — model.tokenize works on it and is None for the other
four, which take phoneme ids (see above). model.driver_source is the authority on what each accepts.
The three repos
| loom.cpp | the engine, vendored here as a submodule |
| loom-exporter | produces the GGUFs this runs |
| loom-py | this one |
Installing
pip install loom-py-rt # once published -- `loom-py` on PyPI clashes with `loompy`,
# and `loom-engine` normalizes to the already-taken `loomengine`
pip install loom-py-rt[hub] # + from_pretrained()
From a checkout — note --recursive, since the engine is a submodule:
git clone --recursive https://github.com/loom-ai-org/loom-py
cd loom-py && pip install -e .
No runtime dependencies. Arrays cross the boundary as plain sequences of floats, so numpy is something
you may use rather than something this package makes you install — list, array.array, numpy arrays
and torch tensors all work.
Testing
pytest tests/ci # the Python layer: coercion and error paths. No model. What CI runs.
pytest tests/gate # a real exported GGUF, end to end.
export LOOM_TEST_MODEL=~/loom-fixtures/matcha_mil.gguf
export LOOM_TEST_MODEL_INPUTS='{"tokens":[16,40,22,30,12,3],"n_steps":4,"seed":1234}'
pytest tests/gate -q
The gate suite is written against no particular architecture on purpose: it asserts the shape of what a loom model is, which is the whole of what this package knows. A test that expected one model's inputs would be this package learning about a model, which is the thing the design exists to avoid.
Roadmap
Shared with loom.cpp, because three of the four are the engine's and this package inherits them by having no per-architecture code of its own.
1. GPUs and NPUs. The engine talks to a single ggml_backend_t and uses no ggml_backend_sched,
which is what has to change before a second device can hold part of a graph. Nothing in this package
should need to change with it — device selection is a property of how the engine is built.
2. Wheels for more platforms. Linux x86-64 today; next macOS on Intel, macOS on Apple Silicon and
Linux on ARM. This is the item most visible from here, since it is what pip install loom-py-rt can
resolve to.
3. More models — P5 in the ledger, ordered by coverage per unit of effort: BERT token classifiers (the smallest possible template, and the first non-audio task) → codec decoders → CNN+CTC and SANM encoders → the remaining TTS families → text encoder-decoders → small classifiers → music. Each lands here for free: a model this package has never heard of works the day the exporter can produce it.
4. The follow-ups the docs already name —
BACKLOG.md is the ledger for all three
repos and the authority. The ones that would show up in this API: a permissively-licensed phonemiser,
which is what would give the four phoneme-input TTS models a tokenize; and the KvCache memory
redesign and quantized KV cache, which decide how large a model this can run on a given machine.
Licence
MIT — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file loom_py_rt-1.0.0rc1.tar.gz.
File metadata
- Download URL: loom_py_rt-1.0.0rc1.tar.gz
- Upload date:
- Size: 1.6 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9c043ec0fa123cf65007a0ef853579897cd0d9f7114dc8f954c5d37e63e53b1a
|
|
| MD5 |
492cf65f1efc36228eedb6a34c624d22
|
|
| BLAKE2b-256 |
3ea178d15a46245d70982553a9bd6e32389d79a254bc9e4c1f238c58b13a1d75
|
File details
Details for the file loom_py_rt-1.0.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.
File metadata
- Download URL: loom_py_rt-1.0.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
- Upload date:
- Size: 1.5 MB
- Tags: CPython 3.13, manylinux: glibc 2.27+ x86-64, manylinux: glibc 2.28+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
eb31b5a6af9b1675ee3a673adc11abf54587562f91f3763fcf775019b0363d18
|
|
| MD5 |
ddb79188703fdc815f208cf8d8ed53a6
|
|
| BLAKE2b-256 |
211fa702ea98ff63402612acc13ab935acc13673bc48d4780ed471d34f8f7d1d
|
File details
Details for the file loom_py_rt-1.0.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.
File metadata
- Download URL: loom_py_rt-1.0.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
- Upload date:
- Size: 1.5 MB
- Tags: CPython 3.12, manylinux: glibc 2.27+ x86-64, manylinux: glibc 2.28+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8cf3cb6060face7d64b5de9d490612cc12a229919596862ab5bec4fd8b2142b9
|
|
| MD5 |
78eca546880daf9b70dda717db1763be
|
|
| BLAKE2b-256 |
de57bdf6ae6f05c8683e560962b17a475555199876f091b4697f454258735d9f
|
File details
Details for the file loom_py_rt-1.0.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.
File metadata
- Download URL: loom_py_rt-1.0.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
- Upload date:
- Size: 1.5 MB
- Tags: CPython 3.11, manylinux: glibc 2.27+ x86-64, manylinux: glibc 2.28+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6eb369785c1cf624b939cfc6f2984d1e942d8e2a3aee58512055379140d95add
|
|
| MD5 |
3f69284dba49d2693488784439f1482e
|
|
| BLAKE2b-256 |
97b8a49a8b08edaaa5a2a6a6db996fa895d294a1a1dde2add23fd17db7053928
|
File details
Details for the file loom_py_rt-1.0.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.
File metadata
- Download URL: loom_py_rt-1.0.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
- Upload date:
- Size: 1.5 MB
- Tags: CPython 3.10, manylinux: glibc 2.27+ x86-64, manylinux: glibc 2.28+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
41c77288a759db5e4b5c037cc14f46f3366bdd9e11ee691637d3dada046dcc8b
|
|
| MD5 |
0651a99daa734b917eccf5dd61ab1dee
|
|
| BLAKE2b-256 |
7618e1d4e31c81535d11bad893128d2fae53de9e5cb156428e8e989c6d2ab5e8
|
File details
Details for the file loom_py_rt-1.0.0rc1-cp39-cp39-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.
File metadata
- Download URL: loom_py_rt-1.0.0rc1-cp39-cp39-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
- Upload date:
- Size: 1.5 MB
- Tags: CPython 3.9, manylinux: glibc 2.27+ x86-64, manylinux: glibc 2.28+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
79191a771b13b16dc266375b3f34c3fafb0de8ad9a579c2b4a5e3a3254e2e81e
|
|
| MD5 |
062ec64457dc243733a039e22e7ee751
|
|
| BLAKE2b-256 |
0a404f108058994e9c32943bc7d29ba9f941378d90673636ca39313300e31744
|