Skip to main content

GitHub release Contributor Covenant

🤗 Speechbox offers a set of speech processing tools, such as punctuation restoration.

Installation

With pip (official package)

pip install speechbox

Contributing

We ❤️ contributions from the open-source community! If you want to contribute to this library, please check out our Contribution guide. You can look out for issues you'd like to tackle to contribute to the library.

Also, say 👋 in our public Discord channel Join us on Discord under ML for Audio and Speech. We discuss the new trends about machine learning methods for speech, help each other with contributions, personal projects or just hang out ☕.

Tasks

Task Description Author
Punctuation Restoration Punctuation restoration allows one to predict capitalized words as well as punctuation by using Whisper. Patrick von Platen
ASR With Speaker Diarization Transcribe long audio files, such as meeting recordings, with speaker information (who spoke when) and the transcribed text. Sanchit Gandhi

Punctuation Restoration

Punctuation restoration relies on the premise that Whisper can understand universal speech. The model is forced to predict the passed words, but is allowed to capitalized letters, remove or add blank spaces as well as add punctuation. Punctuation is simply defined as the offial Python string.Punctuation characters.

Note: For now this package has only been tested with:

and only on some 80 audio samples of patrickvonplaten/librispeech_asr_dummy.

See some transcribed results here.

Web Demo

If you want to try out the punctuation restoration, you can try out the following 🚀 Spaces:

Hugging Face Spaces

Example

In order to use the punctuation restoration task, you need to install Transformers:

pip install --upgrade transformers

For this example, we will additionally make use of datasets to load a sample audio file:

pip install --upgrade datasets

Now we stream a single audio sample, load the punctuation restoring class with "openai/whisper-tiny.en" and add punctuation to the transcription.

from speechbox import PunctuationRestorer
from datasets import load_dataset

streamed_dataset = load_dataset("librispeech_asr", "clean", split="validation", streaming=True)

# get first sample
sample = next(iter(streamed_dataset))

# print out normalized transcript
print(sample["text"])
# => "HE WAS IN A FEVERED STATE OF MIND OWING TO THE BLIGHT HIS WIFE'S ACTION THREATENED TO CAST UPON HIS ENTIRE FUTURE"

# load the restoring class
restorer = PunctuationRestorer.from_pretrained("openai/whisper-tiny.en")
restorer.to("cuda")

restored_text, log_probs = restorer(sample["audio"]["array"], sample["text"], sampling_rate=sample["audio"]["sampling_rate"], num_beams=1)

print("Restored text:\n", restored_text)

See examples/restore for more information.

ASR With Speaker Diarization

Given an unlabelled audio segment, a speaker diarization model is used to predict "who spoke when". These speaker predictions are paired with the output of a speech recognition system (e.g. Whisper) to give speaker-labelled transcriptions.

The combined ASR + Diarization pipeline can be applied directly to long audio samples, such as meeting recordings, to give fully annotated meeting transcriptions.

Web Demo

If you want to try out the ASR + Diarization pipeline, you can try out the following Space:

Hugging Face Spaces

Example

In order to use the ASR + Diarization pipeline, you need to install 🤗 Transformers and pyannote.audio:

pip install --upgrade transformers pyannote.audio

For this example, we will additionally make use of 🤗 Datasets to load a sample audio file:

pip install --upgrade datasets

Now we stream a single audio sample, pass it to the ASR + Diarization pipeline, and return the speaker-segmented transcription:

import torch
from speechbox import ASRDiarizationPipeline
from datasets import load_dataset

device = "cuda:0" if torch.cuda.is_available() else "cpu"
pipeline = ASRDiarizationPipeline.from_pretrained("openai/whisper-tiny", device=device)

# load dataset of concatenated LibriSpeech samples
concatenated_librispeech = load_dataset("sanchit-gandhi/concatenated_librispeech", split="train", streaming=True)
# get first sample
sample = next(iter(concatenated_librispeech))

out = pipeline(sample["audio"])
print(out)

Release files for speechbox 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for speechbox 0.2.1
File Size Uploaded
speechbox-0.2.1.tar.gz 22.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for speechbox 0.2.1
File Interpreter ABI Platform
speechbox-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 42.6 kB

Release files / speechbox-0.2.1.tar.gz

Download URL speechbox-0.2.1.tar.gz
Size 22.4 kB
Tags Source
SHA-256 checksum
How to use checksums
250de696210e2390af61b7204d84cb9c29a9789919ebbfdf5bebf65c4bf35ce4
BLAKE2b-256 checksum
How to use checksums
bcdfa8e3a1ecd01896f98be8d23dbc2d488e3b06c03ab75d5b9e87014199c1f8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.13

Release files / speechbox-0.2.1-py3-none-any.whl

Download URL speechbox-0.2.1-py3-none-any.whl
Size 20.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bfd4c63afa57a4dc26179f0143636d1ebc224a8333618bed9c8c971b06500fb5
BLAKE2b-256 checksum
How to use checksums
32e84cb10f042ea08fd234545e0d386243d2b77f94d3b39ee6432b842242d8c3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.13

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page