Skip to main content

🦜 VieNeu-TTS

VieNeu-TTS is an advanced on-device Vietnamese Text-to-Speech (TTS) with instant voice cloning and English–Vietnamese bilingual support. The SDK defaults to VieNeu-TTS v3 Turbo (48 kHz) and the minimal install is torch-free — on CPU it runs entirely on ONNX Runtime.

Hugging Face v3 Turbo License

✨ Key Features

  • v3 Turbo, 48 kHz — high-fidelity, natural Vietnamese speech (default).
  • Torch-free on CPU — minimal install runs on ONNX Runtime; PyTorch is never imported.
  • fp32 backbone by default on CPU — maximum fidelity. Use Vieneu(precision="int8") for ~1.6× speed & ~4× smaller download (needs a VNNI-capable CPU). Need it much faster, or on a phone / ARM board? See v3 Nano (preview) — Vieneu(mode="v3nano"), lower quality but ~3× faster than Turbo fp32.
  • Built-in default voices — call them by name, no reference clip needed.
  • Instant voice cloning — clone any voice from 3–5s of audio.
  • Emotion cues (experimental) — drop [cười], [thở dài], [hắng giọng] into the text.
  • Bilingual (En–Vi) code-switching, fully offline.

📦 Install

CPU (default) — torch-free, runs v3 Turbo via ONNX Runtime. Most users want this:

pip install vieneu

GPU (CUDA) — only if you have an NVIDIA GPU. Install a CUDA build of PyTorch yourself first. Batching then turns on automatically on CUDA — same API, no code change:

pip install torch==2.8.0 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
pip install "transformers==4.57.6"   # Qwen3 backbone + MOSS codec (pinned — most stable for the GPU SDK)
pip install vieneu

ℹ️ When is GPU actually worth it? The GPU win comes from batching, so it only pays off on long text (many chunks generated together in one forward — long-form or bulk synthesis). For short text the torch-free CPU/ONNX path is usually faster (there's no batch to fill). Use CPU for short, interactive calls; reach for GPU for long-form or high-throughput work.


🚀 Quick Start (Python SDK)

from vieneu import Vieneu

# Default = v3 Turbo (48 kHz). GPU → PyTorch (auto-detected).
# On CPU the backbone runs fp32 by default (max quality); pass precision="int8" for speed (VNNI CPU).
vieneu = Vieneu()                    # fp32 backbone (default, max quality)
# vieneu = Vieneu(precision="int8")  # int8 backbone (faster on CPU with VNNI, ~4x smaller)
# vieneu = Vieneu(mode="v3nano")     # v3 Nano (preview): fastest, lower quality — see "v3 Nano" below
# 💡 On a GPU machine you can still switch to ONNX/CPU if you prefer: Vieneu(backend="onnx")

# 1. Built-in voice by name — no reference needed
print("🔊 Generating speech...")
audio = vieneu.infer("Xin chào, đây là VieNeu-TTS.", voice="Minh Quân")
vieneu.save(audio, "output.wav")
print("✅ Saved to output.wav")

# List the built-in voices
voices = vieneu.list_preset_voices()
print(f"\n🎙️  {len(voices)} built-in voices available:")
for label, voice_id in voices:
    print(f"  - {label} ({voice_id})")

# 2. Reading style: DEPRECATED on v3 Turbo — `style` is accepted but IGNORED. The style
#    is already implied by the reference (preset voice / cloned clip), so output is
#    always the natural reading style. Just pick the voice:
audio = vieneu.infer("Bản tin sáng nay.", voice="Minh Quân")

# 3. Emotion / non-verbal cues — EXPERIMENTAL: [cười] [thở dài] [hắng giọng]
audio = vieneu.infer("Nghe hay quá đi [cười].", voice="Minh Quân")

# 4. ⚡ Batch on GPU: infer_batch() runs many texts in ONE batched forward — same API.
#    On a CUDA GPU the chunks from every text share each forward step (big throughput
#    win); on CPU it still works (no error), just sequentially. Batch caps at
#    max_batch_size (default 32; or infer_batch(..., batch_size=64); batch_size=1
#    disables). A single long infer() also auto-batches its own chunks. Uncomment to try:
#
# import time
# texts = [
#     "Chào cả nhà, hôm nay mình sẽ hướng dẫn các bạn cách cài đặt và sử dụng bộ giọng đọc mới.",
#     "Giọng nghe cực kỳ tự nhiên và truyền cảm, lại có thể chuyển đổi biểu cảm một cách linh hoạt.",
#     "Nếu thấy hữu ích, các bạn nhớ để lại một lượt thích và chia sẻ video này cho mọi người nhé!",
# ] * 10   # 30 texts — enough to fill the batch and really show the GPU throughput win
# t0 = time.time()
# audios = vieneu.infer_batch(texts, voice="Minh Quân")
# elapsed = time.time() - t0
# total_audio = sum(len(a) for a in audios) / 48_000
# print(f"⚡ {len(texts)} texts | audio {total_audio:.1f}s | wall {elapsed:.1f}s | RTF {elapsed/total_audio:.3f}")
# for i, a in enumerate(audios):
#     vieneu.save(a, f"batch_{i}.wav")

🔊 Real-time streaming

v3 Turbo streams frame-by-frame (first audio ~300 ms, RTF < 1 on CPU). Streaming runs on the ONNX/CPU engine — the GPU/PyTorch engine is for batch throughput, not streaming, so pin backend="onnx" for realtime. Iterate infer_stream:

vieneu = Vieneu(backend="onnx")   # force ONNX/CPU — the streaming path (int8)
for chunk in vieneu.infer_stream("Xin chào các bạn!", voice="Minh Quân"):
    play(chunk)   # np.float32 @ 48 kHz, play/write as it arrives

A full FastAPI web demo is in apps/web_stream.py (uv run python -m apps.web_stream → http://localhost:8001).

🦜 Zero-shot Voice Cloning

Clone from a short clip; the reference is auto-denoised and trimmed to ≤ 8s.

from vieneu import Vieneu
vieneu = Vieneu()

# Clone straight from a 3–8s clip
audio = vieneu.infer("Chào bạn, đây là giọng của tôi.", ref_audio="path/to/voice.wav", denoise=True)
vieneu.save(audio, "cloned.wav")

# Save a cloned voice and reuse it by name
vieneu.add_voice("Giọng của tôi", "path/to/voice.wav")
audio = vieneu.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi")

# Just clean up a clip (no synthesis)
wav, sr = vieneu.denoise("noisy.wav", out_path="clean.wav")

denoise, add_voice, and cloning work on every backend, including the torch-free CPU/ONNX install — except v3 Nano (preset voices only).

v3 Nano (preview) — edge devices / weak CPUs only

v3 Turbo stays the default. Vieneu(mode="v3nano") loads a 48M-parameter flow model (ONNX, CPU, 24 kHz) for machines where Turbo is too slow: on the same desktop CPU it runs at RTF 0.22 (16 steps) or 0.11 (steps=8, sway=-1) vs 0.37 for Turbo int8 and 0.62 for Turbo fp32. It is lower quality than Turbo, especially on English and code-switched text, ships 6 preset voices only (no cloning), and has no frame-level streaming.

tts = Vieneu(mode="v3nano")
audio = tts.infer("Xin chào, mình là giọng đọc của VieNeu Nano.", voice="Minh Quân")   # 24 kHz

🔬 Model Overview

Model Engine Device Sample Rate Features
VieNeu-TTS v3 Turbo (default) ONNX (CPU) / PyTorch (GPU) CPU/GPU 48 kHz Default voices, cloning, emotion cues
VieNeu-TTS v3 Nano (preview, weak CPUs) ONNX (CPU) weak CPU / edge 24 kHz 6 preset voices, emotion cues — no cloning, weaker English / En-Vi

Made with ❤️ for the Vietnamese TTS community

Metadata

Release files for vieneu 3.8.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vieneu 3.8.1
File Size Uploaded
vieneu-3.8.1.tar.gz 2.6 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for vieneu 3.8.1
File Interpreter ABI Platform
vieneu-3.8.1-py3-none-any.whl Python 3 none any Details

Total release size: 5.3 MB

Release files / vieneu-3.8.1.tar.gz

Download URL vieneu-3.8.1.tar.gz
Size 2.6 MB
Tags Source
SHA-256 checksum
How to use checksums
56be663c6162e9cef248d4b98a134f052719cbc7e0bb35fed76c755ec2fcbd54
BLAKE2b-256 checksum
How to use checksums
01881474f2a85de82461626ba38745b6c130b766828b7785996b8748858c77a3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7

Release files / vieneu-3.8.1-py3-none-any.whl

Download URL vieneu-3.8.1-py3-none-any.whl
Size 2.6 MB
Tags Python 3
SHA-256 checksum
How to use checksums
bf24f88ec95f96459756d897b04a118e923536d36bcfcdccd13ba6f8f7d580ba
BLAKE2b-256 checksum
How to use checksums
350383c6564f834b90b9fe200c4381a8d8fc506c0697233b1a5b2c3d06d9899f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7

Release history Release notifications | RSS feed

3.8.3

2 release files

3.8.2

2 release files

This release

3.8.1 This release

2 release files

3.8.0

2 release files

3.7.1

2 release files

3.7.0

2 release files

3.6.5

2 release files

3.6.4

2 release files

3.6.3

2 release files

3.6.2

2 release files

3.6.1

2 release files

3.6.0

2 release files

3.5.4

2 release files

3.5.3

2 release files

3.5.2

2 release files

3.5.1

2 release files

3.5.0

2 release files

3.4.0

2 release files

3.3.0

2 release files

3.2.12

2 release files

3.2.11

2 release files

3.2.10

2 release files

3.2.9

2 release files

3.2.8

2 release files

3.2.7

2 release files

3.2.6

2 release files

3.2.5

2 release files

3.2.4

2 release files

3.2.3

2 release files

3.2.2

2 release files

3.2.1

2 release files

3.2.0

2 release files

3.1.4

2 release files

3.1.3

2 release files

3.1.2

2 release files

3.1.1

2 release files

3.1.0

2 release files

3.0.11

2 release files

3.0.10

2 release files

3.0.9

2 release files

3.0.7

2 release files

3.0.6

2 release files

3.0.5

2 release files

3.0.4

2 release files

3.0.3

2 release files

3.0.2

2 release files

3.0.1

2 release files

3.0.0

2 release files

2.7.0

2 release files

2.6.1

2 release files

2.6.0

2 release files

2.5.0

2 release files

2.4.3

2 release files

2.4.2

2 release files

2.4.1

2 release files

2.4.0

2 release files

2.3.0

2 release files

2.2.0

2 release files

2.1.3

2 release files

2.1.2

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.3.0

2 release files

1.2.9

2 release files

1.2.8

2 release files

1.2.7

2 release files

1.2.6

2 release files

1.2.5

2 release files

1.2.4

2 release files

1.2.3

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.9

2 release files

1.1.8

2 release files

1.1.7

2 release files

1.1.6

2 release files

1.1.5

2 release files

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.7

2 release files

1.0.6

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page