Skip to main content

qwentts-cpp-python

Python bindings and wheel packaging for Pascal's qwentts.cpp C ABI.

This package is intentionally small:

  • it loads libqwen with ctypes
  • it exposes buffered and streaming synthesis
  • it can bundle prebuilt libqwen/libggml binaries in platform wheels
  • it does not bundle GGUF model weights

CUDA development build with an existing qwentts.cpp checkout:

python scripts/build_native.py \
  --source /path/to/qwentts.cpp \
  --backend cuda \
  --clean
QWENTTS_CPP_WHEEL_BUILD_TAG=1cu128 python -m build --wheel

CPU development build:

python scripts/build_native.py \
  --source /path/to/qwentts.cpp \
  --backend cpu \
  --clean
QWENTTS_CPP_WHEEL_BUILD_TAG=1cpu python -m build --wheel

--backend cuda is the default because faster-qwen3-tts is a CUDA-first package. CPU builds are still useful for development and smoke tests, but they are not the primary release target.

Installation

The default PyPI package is built for CUDA 12.8:

pip install qwentts-cpp-python

Additional backend-specific wheels are published to Hugging Face Hub as local-version variants. Use them when the PyPI CUDA 12.8 wheel does not match the runtime or GPU target, for example DGX Spark / GB10 with CUDA 13:

pip install "qwentts-cpp-python==0.3.0+cpu" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cpu

pip install "qwentts-cpp-python==0.3.0+cu124" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124

pip install "qwentts-cpp-python==0.3.0+cu128" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu128

pip install "qwentts-cpp-python==0.3.0+cu130" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu130

These commands use pip's --find-links mode against the Hugging Face directory page for the selected flavor. Dependencies still resolve from PyPI normally. The wheels do not bundle CUDA runtime or cuBLAS libraries; use a base image or system installation that provides the matching CUDA runtime.

The Hugging Face wheel pages may contain multiple Linux compatibility tags for the same backend flavor. For example, the cu128 page can host both manylinux_2_35 wheels for Ubuntu 22.04+ and manylinux_2_39 wheels for Ubuntu 24.04+. Pip selects the newest compatible wheel for the current machine.

Pull requests do not build the wheel matrix. The PyPI and Hugging Face publishing workflows each rebuild fresh wheels from the pinned qwentts.cpp revision; validation artifacts are not reused for publishing.

The CI wheel build defaults to qwentts.cpp 7df559a8ca25f66fee02970514ebe5f01dee9055, which retains ABI v2 and includes the latest static-graph, streaming-decode, and widened voice-route changes.

QWENTTS_CPP_WHEEL_BUILD_TAG is useful for local wheelhouses. For public indexes, publish one backend flavor per package/version/platform compatibility tag; otherwise pip has no way to choose between CPU and CUDA binaries.

Local smoke test with a built library:

QWENTTS_CPP_LIBRARY=/path/to/libqwen.so python - <<'PY'
from qwentts_cpp import QwenLibrary
lib = QwenLibrary()
print(lib.version())
PY

Model files are resolved with huggingface-hub by QwenTTS.from_pretrained(...) or passed directly to QwenTTS(...) as GGUF paths.

Cached voice references

qwentts.cpp ABI v2 can skip reference WAV encoding for Base voice cloning by passing precomputed latents:

  • .spk: raw float32 speaker embedding from qwen-codec --talker
  • .rvq: packed 11-bit reference codec stream from qwen-codec

The wrapper can create those files in-process from decoded mono float32 audio at 24 kHz:

from qwentts_cpp import QwenTTS

tts = QwenTTS.from_pretrained("Qwen/Qwen3-TTS-12Hz-1.7B-Base", quant="Q4_K_M")

# ref_audio_24k is a 1-D numpy float32 array, already resampled to 24 kHz.
voice_ref = tts.extract_voice_ref(ref_audio_24k)
voice_ref.save("reference.spk", "reference.rvq")
from qwentts_cpp import QwenTTS, load_speaker_embedding

tts = QwenTTS.from_pretrained("Qwen/Qwen3-TTS-12Hz-1.7B-Base", quant="Q4_K_M")

spk = load_speaker_embedding("reference.spk")
audio, sr = tts.synthesize(
    text="The sky is blue today.",
    lang="english",
    ref_spk_emb=spk,
    max_new_tokens=128,
)

For ICL clone mode, load the RVQ matrix with the model's codebook count and also pass the reference transcript:

from qwentts_cpp import load_rvq_codes

rvq = load_rvq_codes("reference.rvq", tts.num_codebooks())
audio, sr = tts.synthesize(
    text="The sky is blue today.",
    lang="english",
    ref_spk_emb=spk,
    ref_codes=rvq,
    ref_text="Transcript of the reference audio.",
)

Release files for qwentts-cpp-python 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for qwentts-cpp-python 0.3.1
File Interpreter ABI Platform
qwentts_cpp_python-0.3.1-1cu128-py3-none-manylinux_2_39_x86_64.whl Python 3 none Linux glibc 2.39+ x86-64 Details
qwentts_cpp_python-0.3.1-1cu128-py3-none-manylinux_2_39_aarch64.whl Python 3 none Linux glibc 2.39+ ARM64 Details

Total release size: 208.6 MB

Release files / qwentts_cpp_python-0.3.1-1cu128-py3-none-manylinux_2_39_x86_64.whl

Download URL qwentts_cpp_python-0.3.1-1cu128-py3-none-manylinux_2_39_x86_64.whl
Size 104.4 MB
Tags Linux glibc 2.39+ x86-64 Python 3
SHA-256 checksum
How to use checksums
bd0fbb23c567c3943c9cd620df497b584f9be0d355fa7d44a0ea4045aa574c05
BLAKE2b-256 checksum
How to use checksums
2b00bf8240c30322f9ba942c47c01a8c52da1baa06af3f5b9ae6514e090e50db
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 17, 2026.

Transparency log

Release files / qwentts_cpp_python-0.3.1-1cu128-py3-none-manylinux_2_39_aarch64.whl

Download URL qwentts_cpp_python-0.3.1-1cu128-py3-none-manylinux_2_39_aarch64.whl
Size 104.2 MB
Tags Linux glibc 2.39+ ARM64 Python 3
SHA-256 checksum
How to use checksums
3bee54b0f3b779aa52d4a01c5fd540c18c8c9677a3451020175745368c1a788b
BLAKE2b-256 checksum
How to use checksums
2be909b7db5ee5d0b95061cf035b4068e4c7ff9438e2df28629116709f7f5330
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 17, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page