Skip to main content

kittensynth

KittenTTS synthesis layer built on kitteng2p + OnnxVoice.

Installation

A base pip install kittensynth is provider-neutral and does not install ONNX Runtime. For CPU inference, install the CPU provider extra:

python -m pip install "kittensynth[cpu]"

For GPU inference or the bundled G2P runtime, use:

python -m pip install "kittensynth[gpu]"
python -m pip install "kittensynth[cpu,bundled-g2p]"

Managed examples require an ONNX Runtime provider and a working G2P runtime. The CPU plus bundled G2P command above is the copy/paste setup for the examples.

Documentation

This MVP follows the same repository/package shape as PiperSynth/PiperG2P:

pyproject.toml
kittensynth/
tests/
docs/

There is no src/ layer and no hard-coded package version.

Architecture

prepared speakable English
        |
        v
kitteng2p
  - eSpeak
  - Kitten phoneme tokenization
  - Kitten v0.8 token IDs
        |
        v
kittensynth
  - voice alias/style row
  - speed-prior policy
        |
        v
onnxvoice
  - catalogs/install/cache/integrity
  - ORT providers/sessions
  - Kitten graph ABI
        |
        v
float32 waveform

kittensynth has no direct eSpeak or phonemizer dependency. That is now entirely owned by kitteng2p.

Required OnnxVoice contract

KittenSynth requires OnnxVoice 0.2.x. OnnxVoice 0.2.0 and later in that supported range register the built-in Kitten adapter and expose:

system = kitten
runtime.infer(token_ids, style=style, speed=effective_speed)

The installation contains at least:

role=model
role=voices

with 24 kHz metadata and Kitten voice aliases/speed priors.

Managed model

from kittensynth import KittenVoice

with KittenVoice.from_pretrained("nano-0.8-int8") as model:
    result = model.synthesize_prepared(
        "Hello from KittenSynth.",
        voice="Jasper",
    )
    result.save_wav("hello.wav")

Local model

from kittensynth import KittenVoice

with KittenVoice.from_local(
    model_path="kitten_tts_nano_v0_8.onnx",
    voices_path="voices.npz",
    config_path="config.json",
) as model:
    result = model.synthesize_prepared("Local synthesis.", voice="Bella")
    result.save_wav("local.wav")

Examples

Run the prepared-speech examples with a managed Kitten model:

python examples/basic.py
python examples/all_voices.py
python examples/run_all.py --list

The first managed run may download model assets. Generated WAVs and the all-voices manifest are written under example-artefacts/ and intentionally gitignored. Use python examples/run_all.py to execute each example in an isolated output directory and validate its WAV files.

Static voice-level calibration

Calibration is an optional deterministic static gain correction. The package default remains off for backward compatibility; opt in explicitly:

from kittensynth import KittenVoice, SynthesisConfig, VoiceLevelConfig

config = SynthesisConfig(
    speed=1.0,
    voice_level=VoiceLevelConfig(mode="calibrated"),
)
with KittenVoice.from_pretrained("nano-0.8-int8") as model:
    result = model.synthesize_prepared("Prepared speech.", voice="Jasper", config=config)

Calibration defaults by interface

Interface Default behavior
Python KittenVoice.synthesize_prepared(...) Calibration off
examples/basic.py Calibrated
examples/all_voices.py Calibrated; --raw disables it
kittensynth CLI Calibration off; use --voice-level calibrated to opt in

The CLI also accepts --gain-db FLOAT for an explicit static gain override. The API and CLI defaults remain off for backward compatibility.

Speed precedence

When config is provided, config.speed is authoritative and the separate speed= argument is ignored:

model.synthesize_prepared(
    "Prepared speech.",
    speed=1.2,  # ignored because config is supplied
    config=SynthesisConfig(speed=0.9),
)

Catalog lookup uses the exact managed model ID and internal voice/style ID, not the friendly alias. The packaged measured catalog covers all eight internal voices for micro-0.8, mini-0.8, nano-0.8-int8, and nano-0.8-fp32; local models do not inherit managed gains. A reviewed explicit gain_db override is available for a local model. Calibration is neither request-time loudness measurement nor dynamic normalization/limiting. See benchmark documentation for the reproducible measurement and promotion workflow and catalog provenance.

Prepared-text boundary

Like PiperG2P/PiperSynth, this MVP expects prepared speakable text. Written-form semantic expansion (numbers, currencies, dates, URLs, abbreviations) stays outside the engine.

That keeps these packages independent:

semantic preparation -> kitteng2p -> kittensynth -> onnxvoice

Voice selection

The current v0.8 aliases are:

Bella   -> expr-voice-2-f
Jasper  -> expr-voice-2-m
Luna    -> expr-voice-3-f
Bruno   -> expr-voice-3-m
Rosie   -> expr-voice-4-f
Hugo    -> expr-voice-4-m
Kiki    -> expr-voice-5-f
Leo     -> expr-voice-5-m

The style row is selected exactly as upstream v0.8:

min(len(text), style_rows - 1)

Dynamic versioning

Both kitteng2p and kittensynth use the same Git-tag-driven setuptools-scm pattern:

git tag vX.Y.Z
python -m build

The source-ZIP fallback is 0.1.dev0; installed __version__ comes from distribution metadata.

Development

For sibling checkout development:

python -m pip install -e ../kitteng2p
python -m pip install -e ".[dev]"
python -m pytest

Managed synthesis is provided by the Kitten adapter in OnnxVoice; see the examples and static calibration sections above for runnable usage.

Metadata

Release files for kittensynth 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for kittensynth 0.1.1
File Size Uploaded
kittensynth-0.1.1.tar.gz 69.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for kittensynth 0.1.1
File Interpreter ABI Platform
kittensynth-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 94.8 kB

Release files / kittensynth-0.1.1.tar.gz

Download URL kittensynth-0.1.1.tar.gz
Size 69.8 kB
Tags Source
SHA-256 checksum
How to use checksums
defd6622e023f2d93a8252a809c5c0c7c74509f6f40b54bb0b85767a072453ca
BLAKE2b-256 checksum
How to use checksums
2ac6fef3ac4b7ae3ba0446f95e1c3c7941df1fadc6410bd67221921f78e1fc5e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.14

Release files / kittensynth-0.1.1-py3-none-any.whl

Download URL kittensynth-0.1.1-py3-none-any.whl
Size 25.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
84a9f50362c62e6b7310e4114369de9274067c9e321059bb89222ce1d55a2525
BLAKE2b-256 checksum
How to use checksums
5d2a3d5fcbf130b69127dd41d284d9f4f2ff18394121e6e12ee4a73a94bad9e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page