Skip to main content

A Python library for Speech-to-Text and Text-to-Speech in Kimbundu, created to promote language inclusion and support the preservation of an under-represented African language.

Project description

kimbundu-speech-ai

kimbundu-speech-ai is a Python library that provides Text-to-Speech (TTS) and Speech-to-Text (STT) capabilities for Kimbundu, a Bantu language spoken in Angola. It allows developers, researchers, and language technology enthusiasts to build voice-enabled applications, assistive tools, and language preservation projects.


Project Information

This project was developed as a Final Year Project for the Informatic Engineering degree at the Catholic University of Angola, supervised by Engº Domingos Fernando.


Features

Text-to-Speech (TTS)

  • Convert Kimbundu text into high-quality speech audio using the Kimbundu and WaveGlow models.
  • Save audio to a WAV file or play it immediately on your system.

convert_to_file

convert_to_file(
    text: str, 
    model_path: str, 
    waveglow_model_path: str, 
    out_wav: str = 'output.wav', 
    device: str = 'cpu'
)

Description: Converts input text into a spoken audio file.

Parameters:

  • text (str): The input text to convert.
  • model_path (str): Path to the Kimbundu model checkpoint.
  • waveglow_model_path (str): Path to the WaveGlow model checkpoint.
  • out_wav (str, optional): Path to save the WAV file. Default is 'output.wav'.
  • device (str, optional): Device to run inference ('cpu' or 'cuda'). Default is 'cpu'.

Returns:

  • str: Path to the generated WAV file.

play_default_player

play_default_player(
    text: str, 
    model_path: str, 
    waveglow_model_path: str, 
    device: str = 'cpu'
)

Description: Converts text to speech and plays it immediately using the system's default audio player.

Parameters:

  • text (str): The text to convert to speech.
  • model_path (str): Path to the Tacotron model checkpoint.
  • waveglow_model_path (str): Path to the WaveGlow model checkpoint.
  • device (str, optional): Device to run inference ('cpu' or 'cuda'). Default is 'cpu'.

Speech-to-Text (STT)

  • Convert Kimbundu speech audio into text using a fine-tuned Whisper model.

convert_to_text

convert_to_text(
    audio_path: str, 
    model_path: str
)

Description: Transcribes speech from an audio file into text.

Parameters:

  • audio_path (str): Path to the input audio file.
  • model_path (str): Path to the fine-tuned Kimbundu Whisper model.

Returns:

  • str: The transcribed text from the audio.

Installation

pip install kimbundu-speech-ai

Dependencies:

  • torch
  • librosa
  • soundfile
  • numpy
  • unidecode
  • inflect
  • transformers

Models

This library requires models that are downloaded separately due to their size.

Text-to-Speech (TTS) Models

Kimbundu TTS Model Download: https://drive.google.com/file/d/1iXY7beViczLIG0dEqgiBBM-1b71vzOYh/view?usp=sharing

WaveGlow Vocoder Model Download: https://drive.google.com/file/d/1sAnpQP3q8mOfs8rlZibh42OXiBaHzre-/view?usp=sharing

After downloading, keep note of the local paths to these files and pass them to the TTS functions as model_path and waveglow_model_path.

Speech-to-Text (STT) Model

Whisper Small (Kimbundu Fine-Tuned) Model Download (ZIP file): https://drive.google.com/file/d/15kGW4NLfeNcocNygzBOmuWASuBCPwNOx/view?usp=sharing

⚠️ Important: After downloading the ZIP file, you must extract it. Use the extracted folder path as the model_path argument when calling the STT function.

Example:
STT.convert_to_text(
    audio_path="example.wav",
    model_path="path/to/extracted/whisper_small_model/"
)

Do not pass the ZIP file itself as the model path.


Example Usage of the Library

from kimbundu_speech_ai import TTS, STT

# Text-to-Speech (save to WAV)
TTS.convert_to_file(
    text="ngana nzambi",
    model_path="path/to/kimbundu model",
    waveglow_model_path="path/to/waveglow model",
    out_wav="example.wav"
)

# Speech-to-Text
transcription = STT.convert_to_text(
    audio_path="example.wav",
    model_path="path/to/kimbundu's whisper fine-tuned model"
)

print(transcription)

Changelog

[0.0.1] – 2026-01-08

Added
  • Initial public release.
  • Text-to-Speech (TTS) support for Kimbundu using Kimbundu TTS Trained Model + WaveGlow.
  • Speech-to-Text (STT) support using a fine-tuned Whisper model.
Notes
  • Both Neural Network Models only have CPU support for their inference.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kimbundu_speech_ai-0.0.5.tar.gz (67.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kimbundu_speech_ai-0.0.5-py3-none-any.whl (67.4 kB view details)

Uploaded Python 3

File details

Details for the file kimbundu_speech_ai-0.0.5.tar.gz.

File metadata

  • Download URL: kimbundu_speech_ai-0.0.5.tar.gz
  • Upload date:
  • Size: 67.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for kimbundu_speech_ai-0.0.5.tar.gz
Algorithm Hash digest
SHA256 d4f0cd7373d1426008761de3544abfa6f815ccfb3631962324947e1e97a99ae4
MD5 cd59f732c25eb8a019291bdc8ea73ec7
BLAKE2b-256 06034bda34645b0c594552f565eba13e154fdbed31a0209187d2def671ee115e

See more details on using hashes here.

File details

Details for the file kimbundu_speech_ai-0.0.5-py3-none-any.whl.

File metadata

File hashes

Hashes for kimbundu_speech_ai-0.0.5-py3-none-any.whl
Algorithm Hash digest
SHA256 a0242268eb8dba52fa751abed7b833d68c8d35564844f23d8c92f12dffa959e4
MD5 60c0c759f1c4eea407eea140d78398fd
BLAKE2b-256 c5aa18add86dfe09ab754720f2cfc6c78a8cdcd5eaf556e24d32e3535618646c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page