A Python library for Speech-to-Text and Text-to-Speech in Kimbundu, created to promote language inclusion and support the preservation of an under-represented African language.
Project description
kimbundu-speech-ai
kimbundu-speech-ai is a Python library that provides Text-to-Speech (TTS) and Speech-to-Text (STT) capabilities for Kimbundu, a Bantu language spoken in Angola. It allows developers, researchers, and language technology enthusiasts to build voice-enabled applications, assistive tools, and language preservation projects.
Project Information
This project was developed as a Final Year Project for the Informatic Engineering degree at the Catholic University of Angola, supervised by Engº Domingos Fernando.
Features
Text-to-Speech (TTS)
- Convert Kimbundu text into high-quality speech audio using the Kimbundu and WaveGlow models.
- Save audio to a WAV file or play it immediately on your system.
convert_to_file
convert_to_file(
text: str,
model_path: str,
waveglow_model_path: str,
out_wav: str = 'output.wav',
device: str = 'cpu'
)
Description: Converts input text into a spoken audio file.
Parameters:
text(str): The input text to convert.model_path(str): Path to the Kimbundu model checkpoint.waveglow_model_path(str): Path to the WaveGlow model checkpoint.out_wav(str, optional): Path to save the WAV file. Default is'output.wav'.device(str, optional): Device to run inference ('cpu'or'cuda'). Default is'cpu'.
Returns:
str: Path to the generated WAV file.
play_default_player
play_default_player(
text: str,
model_path: str,
waveglow_model_path: str,
device: str = 'cpu'
)
Description: Converts text to speech and plays it immediately using the system's default audio player.
Parameters:
text(str): The text to convert to speech.model_path(str): Path to the Tacotron model checkpoint.waveglow_model_path(str): Path to the WaveGlow model checkpoint.device(str, optional): Device to run inference ('cpu'or'cuda'). Default is'cpu'.
Speech-to-Text (STT)
- Convert Kimbundu speech audio into text using a fine-tuned Whisper model.
convert_to_text
convert_to_text(
audio_path: str,
model_path: str
)
Description: Transcribes speech from an audio file into text.
Parameters:
audio_path(str): Path to the input audio file.model_path(str): Path to the fine-tuned Kimbundu Whisper model.
Returns:
str: The transcribed text from the audio.
Installation
pip install kimbundu-speech-ai
Dependencies:
torchlibrosasoundfilenumpyunidecodeinflecttransformers
Models
This library requires models that are downloaded separately due to their size.
Text-to-Speech (TTS) Models
Kimbundu TTS Model Download: https://drive.google.com/file/d/1iXY7beViczLIG0dEqgiBBM-1b71vzOYh/view?usp=sharing
WaveGlow Vocoder Model Download: https://drive.google.com/file/d/1sAnpQP3q8mOfs8rlZibh42OXiBaHzre-/view?usp=sharing
After downloading, keep note of the local paths to these files and pass them to the TTS functions as model_path and waveglow_model_path.
Speech-to-Text (STT) Model
Whisper Small (Kimbundu Fine-Tuned) Model Download (ZIP file): https://drive.google.com/file/d/15kGW4NLfeNcocNygzBOmuWASuBCPwNOx/view?usp=sharing
⚠️ Important: After downloading the ZIP file, you must extract it. Use the extracted folder path as the model_path argument when calling the STT function.
Example:
STT.convert_to_text(
audio_path="example.wav",
model_path="path/to/extracted/whisper_small_model/"
)
Do not pass the ZIP file itself as the model path.
Example Usage of the Library
from kimbundu_speech_ai import TTS, STT
# Text-to-Speech (save to WAV)
TTS.convert_to_file(
text="ngana nzambi",
model_path="path/to/kimbundu model",
waveglow_model_path="path/to/waveglow model",
out_wav="example.wav"
)
# Speech-to-Text
transcription = STT.convert_to_text(
audio_path="example.wav",
model_path="path/to/kimbundu's whisper fine-tuned model"
)
print(transcription)
Changelog
[0.0.1] – 2026-01-08
Added
- Initial public release.
- Text-to-Speech (TTS) support for Kimbundu using Kimbundu TTS Trained Model + WaveGlow.
- Speech-to-Text (STT) support using a fine-tuned Whisper model.
Notes
- Both Neural Network Models only have CPU support for their inference.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file kimbundu_speech_ai-0.0.4.tar.gz.
File metadata
- Download URL: kimbundu_speech_ai-0.0.4.tar.gz
- Upload date:
- Size: 67.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d0235d81064532cf6ef441c93a65ead032a590359020547cdc9c536431dd05d9
|
|
| MD5 |
c881af13e8c2aa3093253956370914bc
|
|
| BLAKE2b-256 |
1445801a2aaf792c3596aed1c7ef554e7a3ca335a9983726a82ff153a09b9566
|
File details
Details for the file kimbundu_speech_ai-0.0.4-py3-none-any.whl.
File metadata
- Download URL: kimbundu_speech_ai-0.0.4-py3-none-any.whl
- Upload date:
- Size: 67.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7101b84abfe57ee025a5ca1cd33bad054ba176c3cee09ebc32a4937a2810165c
|
|
| MD5 |
f19f70525634f34141113ccada77e133
|
|
| BLAKE2b-256 |
070bdb1bec6da0346cf1503bd7832fa1efb978b6831af6840553a03a84179273
|