Skip to main content

PyKokoro

PyKokoro is a Kokoro speech-synthesis engine. It accepts prepared speech text and explicit pronunciation context, uses KokoroG2P for model-compatible phonemization, and uses OnnxVoice for Kokoro ONNX inference. Document parsing, SSMD interpretation, speech planning, timeline composition, embedded audio, and final-output mastering are intentionally outside PyKokoro.

Quick start

Install one ONNX Runtime provider, then synthesize a prepared string:

pip install "pykokoro[cpu]"
from pykokoro import GenerationConfig, KokoroSynthesizer, SynthesisConfig

config = SynthesisConfig(
    voice="af_sarah",
    generation=GenerationConfig(lang="en-us"),
)
with KokoroSynthesizer(config) as synthesizer:
    rendered = synthesizer.synthesize_text(
        "Hello, world.",
        language="en-us",
        voice="af_sarah",
    )

rendered.save_wav("hello.wav")

The WAV writer stores mono float32 audio. Model assets are provisioned lazily on first use.

Prepared requests and pronunciation context

For orchestration, create a SynthesisSegment with an opaque request ID, prepared text, an explicit pronunciation language, and an actual Kokoro voice. Optional source-aligned pronunciation instructions and linguistic annotations use offsets into that exact text:

from pykokoro import (
    LinguisticToken,
    PronunciationOverride,
    SynthesisSegment,
)

request = SynthesisSegment(
    id="line-001",
    text="Hello Welt.",
    language="en-us",
    voice="af_sarah",
    pronunciation_overrides=(PronunciationOverride(6, 10, language="de"),),
    annotations=(LinguisticToken(0, 5, text="Hello", pos="INTJ"),),
)

Use synthesize() for one request or synthesize_segments() for an ordered iterable of independent requests. Each request yields its own RenderedSegment; PyKokoro does not join batch results or insert cross-request silence. The caller can save, play, or pass each waveform to a separate composition system.

Input text is prepared speech, not a document markup language. PyKokoro does not parse SSMD, YAML front matter, or say-as/voice/pause directives. GenerationConfig.speed is a Kokoro acoustic inference control, not editorial timeline-rate policy.

Engine behavior

The engine owns Kokoro-specific G2P integration, request-local voice/model/style selection, voice blends, token-capacity validation, explicitly configured short-sentence handling, inference, timing reconstruction, waveform validation, tracing, and optional voice-level calibration. discover_models() and discover_lexicons() inspect supported runtime capabilities and lexicon metadata without loading a synthesis session.

Each call returns one request-local RenderedSegment. By default, long_text_split="none" keeps the exact-text path and raises SynthesisInputTooLongError when the prepared text exceeds the model token limit. Set SynthesisConfig.long_text_split="sentence" to enable internal model-safe splitting only when the request is oversized. PhraseSplit is imported lazily for that path; it splits on sentence boundaries, then clauses or safe word boundaries as needed, and joins the audio chunks into one result while preserving the original request text and source-aligned context. long_text_use_spacy=False is the default, so spaCy is not required. Separate caller requests remain separate results, and cross-request composition stays with the caller. See long-text configuration for details.

Installation

Python 3.10 or newer is required. Choose one ONNX Runtime provider extra per environment:

pip install "pykokoro[cpu]"        # CPU
pip install "pykokoro[gpu]"        # NVIDIA CUDA
pip install "pykokoro[openvino]"   # OpenVINO
pip install "pykokoro[directml]"  # DirectML
pip install "pykokoro[coreml]"     # Apple CoreML

For optional direct playback, install pykokoro[cpu,playback]. PyKokoro writes WAV files through soundfile; RenderedSegment.play() uses the optional sounddevice dependency. Install espeak-ng when selecting an eSpeak frontend or fallback.

Documentation and examples

Run the request-centric examples from the repository root with python examples/run_all.py.

Development

pip install -e ".[dev,cpu]"
python -m pytest

See AGENTS.md for the repository's test, lint, and type-check commands.

Release files for pykokoro 0.10.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pykokoro 0.10.0
File Size Uploaded
pykokoro-0.10.0.tar.gz 1.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for pykokoro 0.10.0
File Interpreter ABI Platform
pykokoro-0.10.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.6 MB

Release files / pykokoro-0.10.0.tar.gz

Download URL pykokoro-0.10.0.tar.gz
Size 1.4 MB
Tags Source
SHA-256 checksum
How to use checksums
3116875ddd0cacd30db772895a7836c1a3239be04d0c8a713256662cbec2f30a
BLAKE2b-256 checksum
How to use checksums
71021d8b96e2b3fe3496ee6b1d332fcfaeffec703816429358168e5f8c0c26b0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.14

Release files / pykokoro-0.10.0-py3-none-any.whl

Download URL pykokoro-0.10.0-py3-none-any.whl
Size 124.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e7affd87fe7f7d15c976e4d4ecfbd01ea3c1edd1d7849147fce521b36b949e4b
BLAKE2b-256 checksum
How to use checksums
0dbd01cebcb6d1f6c0e988710f1e08b92779dba1e42f936af0434c9a3f956e07
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.10.0 This release

2 release files

0.9.10

2 release files

0.9.9

2 release files

0.9.8

2 release files

0.9.7

2 release files

0.9.6

2 release files

0.9.5

2 release files

0.9.4

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.8

2 release files

0.8.7

2 release files

0.8.6

2 release files

0.8.5

2 release files

0.8.4

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page