Speechloom
Local audio and video transcription using NeMo-Speech.cpp and Parakeet TDT 0.6B v3. Outputs JSON, plain text, SRT, and WebVTT.
Requirements
- Python 3.10+
- FFmpeg and FFprobe
- A supported native runtime profile for managed setup
- Git, a C++17 compiler, Ninja, and CMake 3.26–3.x only for source builds
- CUDA and
nvcconly for CUDA source builds
Setup
Install the Python package from PyPI:
pipx install speechloom
# or: python3 -m pip install speechloom
Install the runtime and ASR model:
speechloom setup
Speechloom selects a usable backend from the capabilities available on the host. You can also select one explicitly:
speechloom setup --backend cuda
To install translation support (about 16 GiB of free space is needed during conversion):
speechloom setup --backend cuda --features translation
Speaker diarization uses the pinned four-speaker Sortformer model:
speechloom setup --features diarization
speechloom transcribe meeting.mp4 --diarize
Multiple optional features can be selected with
--features translation,diarization.
Setup uses the platform's standard per-user configuration, data, and cache
directories. Existing repository-local .runtime assets are imported in
place; they are not moved or deleted.
Portable installations can set SPEECHLOOM_CONFIG_HOME,
SPEECHLOOM_DATA_HOME, and SPEECHLOOM_CACHE_HOME explicitly.
Check setup state or remove setup caches with:
speechloom setup status
speechloom setup clean --all
The repository-local setup remains available for development and troubleshooting.
Usage
Check the installation:
speechloom doctor
Transcribe one or more files:
speechloom transcribe recording.mp4
speechloom transcribe recordings/ --recursive --workers 2
speechloom transcribe recording.mp4 --output-dir ./output
To translate a transcript, install the translation feature and provide the source and target languages:
speechloom transcribe russian.mp4 --source-language ru --translate-to en
The original files remain transcript.* and subtitles.*. Translated files are
named translation.en.* and subtitles.en.*.
Inspect a completed job:
speechloom inspect transcripts/<job-directory>
Each job contains a manifest, canonical transcript.json, and the requested
text and subtitle formats. Completed jobs are reused unless --force is set.
Run speechloom transcribe --help for all options.
Local API
Install the optional server dependencies and allow the directories a desktop client may submit:
python3 -m pip install -e ".[api]"
speechloom serve --allow-root /path/to/media
The API listens on 127.0.0.1:8765; OpenAPI documentation is available at
/docs. Remote binding requires --allow-remote and a
SPEECHLOOM_API_TOKEN bearer token.
Configuration
Settings are read in this order:
- command-line options
SPEECHLOOM_*environment variables- the selected INI file
- defaults
The default file is in the platform's standard user configuration directory.
--config still selects an explicit file. See config.example.ini for
available settings.
Tests
PYTHONDONTWRITEBYTECODE=1 PYTHONPATH=src python3 -m unittest discover -s tests -v
Real-runtime tests are enabled by setting SPEECHLOOM_TEST_NEMO,
SPEECHLOOM_TEST_MODEL, and SPEECHLOOM_TEST_MEDIA.
License
The project is licensed under MIT. Models are distributed separately under their respective licenses.
Metadata
Release files for speechloom 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| speechloom-0.1.1.tar.gz | 60.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| speechloom-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 134.4 kB
Release files / speechloom-0.1.1.tar.gz
| Download URL | speechloom-0.1.1.tar.gz |
|---|---|
| Size | 60.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
833db32d66efb4d003955fa5cf0760b413233f1b095b71935676cffda43fd17a
|
|
BLAKE2b-256 checksum How to use checksums |
46915cb41f0481f6f002a0216d6f179bfd1042bb937657e72872ce5442cab764
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.
Transparency logRelease files / speechloom-0.1.1-py3-none-any.whl
| Download URL | speechloom-0.1.1-py3-none-any.whl |
|---|---|
| Size | 74.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d8743a916b52198c5baa1bcf889bf508e60c4ee1c860e3338da15b607da917f7
|
|
BLAKE2b-256 checksum How to use checksums |
818cabd21ae2d76f3430f72af2e8c54d16be84453e752ca115a8398dc2803211
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.
Transparency log