Production-grade neural TTS with extreme clarity, voice cloning, and multilingual support
Project description
Neural TTS System - Production Grade
A complete, production-ready Text-to-Speech system built in pure Python with PyTorch.
Features
- Extreme Clarity: HiFi-GAN vocoder for near-human audio quality
- Multilingual: Unicode-based phoneme processing for all languages
- Fast: Optimized FastSpeech2 architecture for real-time inference
- Speaker Cloning: Clone any voice with 5-10 seconds of audio
- Flexible: Adjustable pitch, speed, emotion, and tone
- Multi-speaker: Support for multiple voice identities
- Low Latency: CPU and GPU support with optimized inference
- Modular: Clean architecture for easy customization
Architecture
1. Text Processor (text/)
- Phonemizer: Converts text to phonemes using language-specific rules
- Normalizer: Unicode normalization and text cleaning
- Tokenizer: Converts phonemes to model-ready tokens
2. Acoustic Model (models/acoustic/)
- FastSpeech2: Non-autoregressive transformer for mel-spectrogram generation
- Variance Adaptors: Pitch, energy, and duration predictors
- Multi-head Attention: Captures long-range dependencies
3. Vocoder (models/vocoder/)
- HiFi-GAN: High-fidelity generative adversarial network
- Multi-scale discriminators: Ensures audio quality at multiple resolutions
- Multi-period discriminators: Captures periodic patterns in speech
4. Speaker Encoder (models/speaker/)
- Speaker Embedding: Extracts voice characteristics from reference audio
- Voice Cloning: Adapts model to new speakers with minimal data
Project Structure
tts/
├── config/
│ ├── __init__.py
│ ├── acoustic_config.py # FastSpeech2 hyperparameters
│ ├── vocoder_config.py # HiFi-GAN hyperparameters
│ └── training_config.py # Training settings
├── text/
│ ├── __init__.py
│ ├── phonemizer.py # Text to phoneme conversion
│ ├── normalizer.py # Text normalization
│ └── symbols.py # Phoneme symbols and mappings
├── models/
│ ├── __init__.py
│ ├── acoustic/
│ │ ├── __init__.py
│ │ ├── fastspeech2.py # Main acoustic model
│ │ ├── transformer.py # Transformer blocks
│ │ └── variance_adaptor.py # Pitch/duration/energy predictors
│ ├── vocoder/
│ │ ├── __init__.py
│ │ ├── hifigan.py # HiFi-GAN generator
│ │ └── discriminator.py # Multi-scale/period discriminators
│ └── speaker/
│ ├── __init__.py
│ └── encoder.py # Speaker embedding network
├── data/
│ ├── __init__.py
│ ├── dataset.py # PyTorch dataset classes
│ └── preprocessing.py # Audio preprocessing utilities
├── training/
│ ├── __init__.py
│ ├── train_acoustic.py # Train FastSpeech2
│ ├── train_vocoder.py # Train HiFi-GAN
│ └── losses.py # Loss functions
├── inference/
│ ├── __init__.py
│ ├── synthesizer.py # Main TTS inference engine
│ └── voice_cloner.py # Speaker cloning utilities
├── utils/
│ ├── __init__.py
│ ├── audio.py # Audio processing utilities
│ └── tools.py # General utilities
├── checkpoints/ # Model checkpoints (created during training)
├── outputs/ # Generated audio files
├── requirements.txt
└── README.md
Quick Start
Installation
pip install -r requirements.txt
Inference (Pre-trained)
from inference.synthesizer import TTSSynthesizer
# Initialize synthesizer
tts = TTSSynthesizer(
acoustic_checkpoint='checkpoints/acoustic_model.pt',
vocoder_checkpoint='checkpoints/vocoder_model.pt'
)
# Generate speech
audio = tts.synthesize(
text="Hello, this is a test of the neural TTS system.",
language="en",
speaker_id=0,
pitch_scale=1.0,
speed_scale=1.0,
energy_scale=1.0
)
# Save to file
tts.save_audio(audio, 'outputs/output.wav')
Voice Cloning
from inference.voice_cloner import VoiceCloner
cloner = VoiceCloner(
acoustic_checkpoint='checkpoints/acoustic_model.pt',
vocoder_checkpoint='checkpoints/vocoder_model.pt'
)
# Clone voice from reference audio
audio = cloner.clone_voice(
text="This is spoken in the cloned voice.",
reference_audio='path/to/reference.wav',
language="en"
)
cloner.save_audio(audio, 'outputs/cloned.wav')
Training
1. Prepare Dataset
Organize your dataset in the following format:
dataset/
├── metadata.csv # Format: filename|text|speaker_id
└── wavs/
├── audio1.wav
├── audio2.wav
└── ...
2. Train Acoustic Model
python training/train_acoustic.py \
--dataset_path /path/to/dataset \
--output_dir checkpoints/acoustic \
--batch_size 32 \
--epochs 1000 \
--gpu 0
3. Train Vocoder
python training/train_vocoder.py \
--dataset_path /path/to/dataset \
--output_dir checkpoints/vocoder \
--batch_size 16 \
--epochs 500 \
--gpu 0
Configuration
All hyperparameters are in config/:
acoustic_config.py: Model architecture, attention heads, hidden dimensionsvocoder_config.py: Generator/discriminator settings, upsampling ratestraining_config.py: Learning rates, batch sizes, optimization settings
Performance Optimization
Speed
- Use mixed precision training (
--fp16) - Reduce model size in config files
- Use ONNX export for deployment
Quality
- Increase model hidden dimensions
- Train longer with more data
- Use higher sampling rate (48kHz)
Memory
- Reduce batch size
- Use gradient accumulation
- Enable gradient checkpointing
Supported Languages
The phonemizer supports all languages through Unicode normalization:
- English (en)
- Spanish (es)
- French (fr)
- German (de)
- Chinese (zh)
- Japanese (ja)
- Korean (ko)
- Russian (ru)
- Arabic (ar)
- Hindi (hi)
- And many more...
License
MIT License - Free for commercial and personal use.
Citation
If you use this system in your research, please cite:
@software{neural_tts_2025,
title={Neural TTS System - Production Grade},
author={Soham Vyas},
year={2025},
url={https://github.com/vyassoham/neural-tts}
}
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file neural_tts-1.0.0.tar.gz.
File metadata
- Download URL: neural_tts-1.0.0.tar.gz
- Upload date:
- Size: 40.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ed646d0278b3631680a7d7263b97932d3c3ed239d3de60ffb7bd9e23d58367b5
|
|
| MD5 |
df92cf78984b43b83991495247a492ee
|
|
| BLAKE2b-256 |
6eb0cb38d64076748c84f5e68e9b4abad9ca02088edfc2363e93a1a36f29d41d
|
File details
Details for the file neural_tts-1.0.0-py3-none-any.whl.
File metadata
- Download URL: neural_tts-1.0.0-py3-none-any.whl
- Upload date:
- Size: 50.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2af6cb566fcad7ac4dbbcc138aafc7ee16d61cf5afc943fa52e9d85c07f79de4
|
|
| MD5 |
870300a143e8ee6ea357438a87e7440d
|
|
| BLAKE2b-256 |
2ac3756a22997fbc2930036415ee7ec50aceea792ac05b5762e6a3f7876ae839
|