Skip to main content

Speechloom

Local audio and video transcription using NeMo-Speech.cpp and Parakeet TDT 0.6B v3. Outputs JSON, plain text, SRT, and WebVTT.

Requirements

  • Python 3.10+
  • FFmpeg and FFprobe
  • A supported native runtime profile for managed setup
  • Git, a C++17 compiler, Ninja, and CMake 3.26–3.x only for source builds
  • CUDA and nvcc only for CUDA source builds

Setup

Install the Python package from PyPI:

pipx install speechloom
# or: python3 -m pip install speechloom

Install the runtime and ASR model:

speechloom setup

Speechloom selects a usable backend from the capabilities available on the host. You can also select one explicitly:

speechloom setup --backend cuda

To install translation support (about 16 GiB of free space is needed during conversion):

speechloom setup --backend cuda --features translation

Speaker diarization uses the pinned four-speaker Sortformer model:

speechloom setup --features diarization
speechloom transcribe meeting.mp4 --diarize

Multiple optional features can be selected with --features translation,diarization.

Setup uses the platform's standard per-user configuration, data, and cache directories. Existing repository-local .runtime assets are imported in place; they are not moved or deleted.

Portable installations can set SPEECHLOOM_CONFIG_HOME, SPEECHLOOM_DATA_HOME, and SPEECHLOOM_CACHE_HOME explicitly.

Check setup state or remove setup caches with:

speechloom setup status
speechloom setup clean --all

The repository-local setup remains available for development and troubleshooting.

Usage

Check the installation:

speechloom doctor

Transcribe one or more files:

speechloom transcribe recording.mp4
speechloom transcribe recordings/ --recursive --workers 2
speechloom transcribe recording.mp4 --output-dir ./output

To translate a transcript, install the translation feature and provide the source and target languages:

speechloom transcribe russian.mp4 --source-language ru --translate-to en

The original files remain transcript.* and subtitles.*. Translated files are named translation.en.* and subtitles.en.*.

Inspect a completed job:

speechloom inspect transcripts/<job-directory>

Each job contains a manifest, canonical transcript.json, and the requested text and subtitle formats. Completed jobs are reused unless --force is set.

Run speechloom transcribe --help for all options.

Local API

Install the optional server dependencies and allow the directories a desktop client may submit:

python3 -m pip install -e ".[api]"
speechloom serve --allow-root /path/to/media

The API listens on 127.0.0.1:8765; OpenAPI documentation is available at /docs. Remote binding requires --allow-remote and a SPEECHLOOM_API_TOKEN bearer token.

Configuration

Settings are read in this order:

  1. command-line options
  2. SPEECHLOOM_* environment variables
  3. the selected INI file
  4. defaults

The default file is in the platform's standard user configuration directory. --config still selects an explicit file. See config.example.ini for available settings.

Tests

PYTHONDONTWRITEBYTECODE=1 PYTHONPATH=src python3 -m unittest discover -s tests -v

Real-runtime tests are enabled by setting SPEECHLOOM_TEST_NEMO, SPEECHLOOM_TEST_MODEL, and SPEECHLOOM_TEST_MEDIA.

License

The project is licensed under MIT. Models are distributed separately under their respective licenses.

Metadata

Release files for speechloom 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for speechloom 0.1.1
File Size Uploaded
speechloom-0.1.1.tar.gz 60.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for speechloom 0.1.1
File Interpreter ABI Platform
speechloom-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 134.4 kB

Release files / speechloom-0.1.1.tar.gz

Download URL speechloom-0.1.1.tar.gz
Size 60.3 kB
Tags Source
SHA-256 checksum
How to use checksums
833db32d66efb4d003955fa5cf0760b413233f1b095b71935676cffda43fd17a
BLAKE2b-256 checksum
How to use checksums
46915cb41f0481f6f002a0216d6f179bfd1042bb937657e72872ce5442cab764
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.

Transparency log

Release files / speechloom-0.1.1-py3-none-any.whl

Download URL speechloom-0.1.1-py3-none-any.whl
Size 74.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d8743a916b52198c5baa1bcf889bf508e60c4ee1c860e3338da15b607da917f7
BLAKE2b-256 checksum
How to use checksums
818cabd21ae2d76f3430f72af2e8c54d16be84453e752ca115a8398dc2803211
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page