loudkit
Natural-sounding text-to-speech for products, scripts and experiments.
loudkit runs on your own hardware. It includes twenty voices across ten languages, voice cloning from about ten seconds of audio, and native SDKs for Python, Swift, Go, Rust and TypeScript. Download the model once and run offline with no account, telemetry or usage bill.
Hear all 20 voices | Open in Colab | Model | Documentation
- 20 voices in 10 languages: English, Spanish, French, German, Italian, Polish, Portuguese, Dutch, Swedish and Danish.
- Voice cloning: a reusable profile of about 150 KB from roughly ten seconds of audio.
- Runs locally: PyTorch, ONNX Runtime and CoreML.
- Five native SDKs: Python, Swift, Go, Rust and TypeScript.
Make a WAV
pip install "loudkit[torch,audio,hub]"
import loudkit as lk
engine = lk.load("loudreader/loudr-1")
voice = lk.voice("joe", repo="loudreader/loudr-1")
engine.synthesize("Hello from loudkit.", voice, seed=7).save("hello.wav")
From the shell:
loudkit speak --checkpoint loudreader/loudr-1 --voice joe \
"Hello from loudkit." -o hello.wav
The first run downloads the 747 MB synthesis checkpoint and the voices. Later runs use the local cache. See Getting started for revision pinning, local directories and a complete first-run walkthrough.
Voices
Open the voice gallery to compare all twenty shipped voices. Each generated sample is shown next to the enrollment reference used to build its profile.
The voice roster records the source, licence and consent basis for every profile. The voices come from recordings donated for speech technology or from CC0 and CC-BY speech corpora. No scraped celebrity voices ship with the project.
We have evaluated English by ear. We do not speak the other nine languages well enough to judge their naturalness reliably. If you do, please listen and tell us what sounds good or wrong. Reports from native speakers are especially welcome.
Clone a voice
Use a recording that you own or have permission to use:
pip install "loudkit[torch,audio,enroll,hub]"
loudkit clone my-recording.wav --checkpoint loudreader/loudr-1 \
--name my-voice --language en
loudkit speak --checkpoint loudreader/loudr-1 \
--voice voices/my-voice.safetensors "Now in a cloned voice." -o cloned.wav
The result is a portable profile, not another copy of the model. Python users
can call the same path through lk.enroll. See
Cloning a voice
for recording advice and
Responsible use
for the consent rules.
Download only the runtime you need
The Hugging Face repository contains every supported format. The downloader selects one runtime and leaves the rest behind.
| use case | command |
|---|---|
| Python with PyTorch | loudkit download loudreader/loudr-1 --for torch |
| Python, JS, Go or Rust with ONNX Runtime | loudkit download loudreader/loudr-1 --for onnx --local-dir loudr-1 |
| Swift or Python with CoreML | loudkit download loudreader/loudr-1 --for coreml --local-dir loudr-1 |
Add --with-cloning if that installation also needs enrollment. The
model card lists exact download
sizes and files.
Measured speed
| path | hardware | real-time factor |
|---|---|---|
| PyTorch with CUDA graphs | RTX 3090 | 7.47x |
| PyTorch with CUDA graphs | Jetson Orin Nano | 1.83x |
| split PyTorch engine* | Apple M3 Pro | 3.43x |
| ONNX Runtime, CPU provider | Apple M3 Pro | 1.21x |
| PyTorch CPU reference | Apple M3 Pro | 0.33x |
* "Split" describes device placement, not a different model or checkpoint. The token generator runs on the CPU while the mel and vocoder renderer runs on the Apple GPU through MPS. Adjacent windows can overlap across the two devices.
Higher is faster, and 1.0x means real time. CPU performance depends heavily on the runtime: ONNX Runtime is faster than real time on the measured M3 Pro, while the PyTorch CPU reference path is not.
For batched workloads, the token generator reaches 20.1x aggregate throughput at batch 1 and 153.1x at batch 64 on the RTX 3090. The highest measured result is 170.8x on an A100 at batch 64. These are generator-only throughput numbers, not single-request latency or end-to-end RTF. Full commands, hardware and caveats are in Benchmarks.
SDKs and local APIs
Python is the reference implementation. Swift, Go, Rust and TypeScript are native ports checked against the same conformance fixtures.
| SDK | runtime | install |
|---|---|---|
| Python | PyTorch, ONNX Runtime or CoreML | pip install "loudkit[torch]" plus the extra for your backend |
| Swift | CoreML | Swift Package Manager, from 0.1.0 |
| Go | ONNX Runtime | go get github.com/loudreader/loudkit/go |
| Rust | ONNX Runtime through ort |
cargo add loudkit |
| TypeScript | onnxruntime-node |
npm install loudkit |
For processes that should keep one engine warm, loudkit also includes:
- a local HTTP server with an OpenAI-compatible speech endpoint;
- a small MCP preview server;
- typed gRPC streaming with backpressure;
- a Speech Dispatcher module for Linux screen readers.
Servers and agents documents the exact contract of each integration.
Output and reproducibility
Synthesis results support chunk and estimated word timestamps, pitch-preserving speed from 0.5x to 2.0x, and C2PA Content Credentials. Saved WAVs and server responses carry an unsigned claim-only C2PA manifest by default. It records the model fingerprint, checkpoint and voice digests, seed, backend and audio digest.
For a fixed build, device and backend, the same text, voice and seed produce the same waveform. Different devices or backends can differ at the waveform level because of floating-point arithmetic. The precise guarantees and measurements live in the Identity contract and Measured parity.
Scope
loudkit is an inference toolbox, not a hosted speech platform. It does not provide accounts, billing, multi-tenancy, model training or an emotion control. The local server expects you to provide any public-facing authentication, rate limits and TLS.
The project ships twenty permitted voice profiles and the code to enroll your own. It will not help with undisclosed impersonation, bypassing voice authentication or removing provenance from generated audio.
Documentation
- Getting started: first synthesis, caches and local files.
- Guides: long form, cloning, servers and every SDK.
- Model card: model lineage, limitations and release layout.
- Supported in v0.1: the public compatibility boundary.
- Troubleshooting: common failures and concrete fixes.
- Architecture: components and package layout for contributors.
Licence
The code and loudr-1 release are
Apache-2.0. The
tokenizer and speaker encoder retain their upstream MIT licence from Chatterbox.
NOTICE lists every
upstream component and licence. The voice encoder's chain is in its
provenance record.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file loudkit-0.1.0.tar.gz.
File metadata
- Download URL: loudkit-0.1.0.tar.gz
- Upload date:
- Size: 2.2 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9a78b3dc860be931d8959ed70383b55e710b0d6482020b4201bb234d838de059
|
|
| MD5 |
91f8dd0932cf0ce66a9e1c75bca65ec7
|
|
| BLAKE2b-256 |
e12c8678acba27f273ff8092dfbf3ede88fd3b0cb9c472c5f211491b8bf245a9
|
Provenance
The following attestation bundles were made for loudkit-0.1.0.tar.gz:
Publisher:
release.yml on loudreader/loudkit
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
loudkit-0.1.0.tar.gz -
Subject digest:
9a78b3dc860be931d8959ed70383b55e710b0d6482020b4201bb234d838de059 - Sigstore transparency entry: 2588931236
- Sigstore integration time:
-
Permalink:
loudreader/loudkit@58fd4a58de8980b42c1021492728876d67ea2718 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/loudreader
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@58fd4a58de8980b42c1021492728876d67ea2718 -
Trigger Event:
push
-
Statement type:
File details
Details for the file loudkit-0.1.0-py3-none-any.whl.
File metadata
- Download URL: loudkit-0.1.0-py3-none-any.whl
- Upload date:
- Size: 2.2 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6f166fbd9ae3947f54d42d0f792e57771db7e2f9f61e466acf6269efd16f10d3
|
|
| MD5 |
3e77a8b9126e21e630b7e8619e1603c1
|
|
| BLAKE2b-256 |
7bb04b82a1295d6378e185416dc480391ebfc7f3d92c7252e7cfd99ec0d55dbb
|
Provenance
The following attestation bundles were made for loudkit-0.1.0-py3-none-any.whl:
Publisher:
release.yml on loudreader/loudkit
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
loudkit-0.1.0-py3-none-any.whl -
Subject digest:
6f166fbd9ae3947f54d42d0f792e57771db7e2f9f61e466acf6269efd16f10d3 - Sigstore transparency entry: 2588931805
- Sigstore integration time:
-
Permalink:
loudreader/loudkit@58fd4a58de8980b42c1021492728876d67ea2718 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/loudreader
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@58fd4a58de8980b42c1021492728876d67ea2718 -
Trigger Event:
push
-
Statement type: