Skip to main content

StableTTS

Next-generation TTS model using flow-matching and DiT, inspired by Stable Diffusion 3.

Introduction

As the first open-source TTS model that tried to combine flow-matching and DiT, StableTTS is a fast and lightweight TTS model for chinese, english and japanese speech generation. It has 31M parameters.

✨ Huggingface demo: 🤗

News

2024/10: A new autoregressive TTS model is coming soon...

2024/9: 🚀 StableTTS V1.1 Released ⭐ Audio quality is largely improved ⭐

⭐ V1.1 Release Highlights:

  • Fixed critical issues that cause the audio quality being much lower than expected. (Mainly in Mel spectrogram and Attention mask)
  • Introduced U-Net-like long skip connections to the DiT in the Flow-matching Decoder.
  • Use cosine timestep scheduler from Cosyvoice
  • Add support for CFG (Classifier-Free Guidance).
  • Add support for FireflyGAN vocoder.
  • Switched to torchdiffeq for ODE solvers.
  • Improved Chinese text frontend (partially based on gpt-sovits2).
  • Multilingual support (Chinese, English, Japanese) in a single checkpoint.
  • Increased parameters: 10M -> 31M.

Pretrained models

Text-To-Mel model

Download and place the model in the ./checkpoints directory, it is ready for inference, finetuning and webui.

Model Name Task Details Dataset Download Link
StableTTS text to mel 600 hours 🤗

Mel-To-Wav model

Choose a vocoder (vocos or firefly-gan ) and place it in the ./vocoders/pretrained directory.

Model Name Task Details Dataset Download Link
Vocos mel to wav 2k hours 🤗
firefly-gan-base mel to wav HiFi-16kh download from fishaudio

Installation

  1. Install pytorch: Follow the official PyTorch guide to install pytorch and torchaudio. We recommend the latest version (tested with PyTorch 2.4 and Python 3.12).

  2. Install Dependencies: Run the following command to install the required Python packages:

pip install -r requirements.txt

Inference

For detailed inference instructions, please refer to inference.ipynb

We also provide a webui based on gradio, please refer to webui.py

Training

StableTTS is designed to be trained easily. We only need text and audio pairs, without any speaker id or extra feature extraction. Here’s how to get started:

Preparing Your Data

  1. Generate Text and Audio pairs: Generate the text and audio pair filelist as ./filelists/example.txt. Some recipes of open-source datasets could be found in ./recipes.

  2. Run Preprocessing: Adjust the DataConfig in preprocess.py to set your input and output paths, then run the script. This will process the audio and text according to your list, outputting a JSON file with paths to mel features and phonemes.

Note: Process multilingual data separately by changing the language setting in DataConfig

Start training

  1. Adjust Training Configuration: In config.py, modify TrainConfig to set your file list path and adjust training parameters (such as batch_size) as needed.

  2. Start the Training Process: Launch train.py to start training your model.

Note: For finetuning, download the pretrained model and place it in the model_save_path directory specified in TrainConfig. Training script will automatically detect and load the pretrained checkpoint.

(Optional) Vocoder training

The ./vocoder/vocos folder contains the training and finetuning codes for vocos vocoder.

For other types of vocoders, we recommend to train by using fishaudio vocoder: an uniform interface for developing various vocoders. We use the same spectrogram transform so the vocoders trained is compatible with StableTTS.

Model structure

  • We use the Diffusion Convolution Transformer block from Hierspeech++, which is a combination of original DiT and FFT(Feed forward Transformer from fastspeech) for better prosody.

  • In flow-matching decoder, we add a FiLM layer before DiT block to condition timestep embedding into model.

References

The development of our models heavily relies on insights and code from various projects. We express our heartfelt thanks to the creators of the following:

Direct Inspirations

Matcha TTS: Essential flow-matching code.

Grad TTS: Diffusion model structure.

Stable Diffusion 3: Idea of combining flow-matching and DiT.

Vits: Code style and MAS insights, DistributedBucketSampler.

Additional References:

plowtts-pytorch: codes of MAS in training

Bert-VITS2 : numba version of MAS and modern pytorch codes of Vits

fish-speech: dataclass usage and mel-spectrogram transforms using torchaudio, gradio webui

gpt-sovits: melstyle encoder for voice clone

coqui xtts: gradio webui

Chinese Dirtionary Of DiffSinger: Multi-langs_Dictionary and atonyxu's fork

TODO

  • Release pretrained models.
  • Support Japanese language.
  • User friendly preprocess and inference script.
  • Enhance documentation and citations.
  • Release multilingual checkpoint.

Disclaimer

Any organization or individual is prohibited from using any technology in this repo to generate or edit someone's speech without his/her consent, including but not limited to government leaders, political figures, and celebrities. If you do not comply with this item, you could be in violation of copyright laws.

Metadata

Release files for MyanmarTTS 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for MyanmarTTS 1.0.0
File Size Uploaded
myanmartts-1.0.0.tar.gz 41.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for MyanmarTTS 1.0.0
File Interpreter ABI Platform
myanmartts-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 101.9 kB

Release files / myanmartts-1.0.0.tar.gz

Download URL myanmartts-1.0.0.tar.gz
Size 41.9 kB
Tags Source
SHA-256 checksum
How to use checksums
06381b316af9e023757f1c2f87ca78127f99e41ae82ab4e1f2cd2a2bedc5cea3
BLAKE2b-256 checksum
How to use checksums
b592e29f8f457a5ec697019159babf1de87e517b1ffab9140dfa843d5a46476a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / myanmartts-1.0.0-py3-none-any.whl

Download URL myanmartts-1.0.0-py3-none-any.whl
Size 60.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c1ab7d468c0bf2abce2f889676b1f1a73118e701a40f2f61fd33464ebc67f2d9
BLAKE2b-256 checksum
How to use checksums
ceb1385c9d37894b44b9df7600bb2b430264ba214431f06a2dc1e5035c628249
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release history Release notifications | RSS feed

1.0.2

2 release files

1.0.1

2 release files

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page