Transcribe Allign TextGrid
A small wrapper package around whisper-timestamped. Create force-aligned transcription TextGrids from raw audio.
Installation
Requirements
Python3.8topython3.11.- Use the executable
python3.xon Unix, available in most package managers, orpy -3.xon Windows. - This command line executable of will be referred to as
[python-executable]for the rest of the instructions - Install pip on old python versions with
[python-executable] -m ensurepip --default-pip
- Use the executable
ffmpegUsually preinstalled on Linux. For Windows see instructions for installation on the whisper repositorygitUsually preinstalled on Linux. For Windows, visit the git site.- Needed for installation of whisper-timestamped, as it is not available on PyPI
- Note that it needs to be available from the command line; git-bash might not work.
Installing Torch
Torch, on which Whisper is built, is quite a low-level library, meaning which version you'll need depends on your OS and type of GPU. On Mac and Windows, pip will by default install a non-accelerated CPU version of the library. If you are on Linux, it will presume you have a CUDA-capable (which is to say Nvidia branded) GPU. If you are on Windows and have an Nvidia GPU you can use, or are on Linux and either do not have a GPU or have an AMD GPU, you should check out the more detailed torch installation instructions here.
This should be done before installing transcribe_allign_textgrid and whisper_timestamped.
Installing
Once the requirements are satisfied, you can install whisper-timestamped and this package:
Whisper-timestamped is not on Pypi, so a separate git+ install is needed. (If you only want to use the package as a library instead of a cli, whisper-timestamped is not a dependency, and this manual install of it is not needed.)
[python-executable] -m pip install git+https://github.com/linto-ai/whisper-timestamped
[python-executable] -m pip install transcribe_allign_textgrid
Running from the command line
Once the application is installed, you can run it with:
[python-executable] -m transcribe_allign_textgrid [path]
here path is the path to the audio files.
- If a directory path is passed, all audio files in the directory will be transcribed, and force-aligned transcription TextGrids of the same name will be generated in this directory.
- If a file path is passed, a force-aligned transcription TextGrid will be generated into the same directory with the same name as the original file.
- If a glob is passed, the glob will be resolved and all matches will be processed as if the files were passed individually
- By default, if a non-audio file is passed, an error is raised. To skip those instead, pass the
--skipflag.
Selecting a different model
By default, this will run on the smallest, that is, least accurate and fastest, model, tiny. To run with another model, pass it as an argument:
[python-executable] -m transcribe_allign_textgrid [path] --model [model]
The available models are:
| name | Parameters | Required VRAM | Relative speed |
|---|---|---|---|
| tiny | 39 M | ~1 GB | ~32x |
| base | 74 M | ~1 GB | ~16x |
| small | 244 M | ~2 GB | ~6x |
| medium | 769 M | ~5 GB | ~2x |
| large | 1550 M | ~10 GB | 1x |
Specifying what language to use
By default, the application will try to detect what language is used automatically. However, you can also specify this manually:
[python-executable] -m transcribe_allign_textgrid [path] --language [language]
# Or also specifying what model to use:
[python-executable] -m transcribe_allign_textgrid [path] --model [model] --language [language]
To see what languages are available, please see the tokenizer.py file in the Whisper source (Yes, the OpenAI team themselves recommends finding it this way, too.)
Using as a library
The tool can also be used as a library. It exports one function: whisper_to_textgrid() Which takes in a transcription object (nested dictionary) from whisper-timestamped and returns a Textgrid object from praatio. The typical Json output from whisper-timestamped works, too.
This library part of the package does not depend on whisper-timestamped, to make it fully installable and usable as a requirement via pipy.
Output
The output TextGrids have four TextGridTiers:
segments_textThe text in a given segment (Speaker's turn)segments_confidenceThe confidence the model has that this is the correct labeling and segmentation for the segmentwords_textThe text of a given wordwords_confidenceThe confidence the model has that this is the current labeling and segmentation for this word.
If one of these tiers would have been empty per the output of whisper-timestamped, to satisfy Praat's error handling, a tier with an empty interval (0.0, 0.1) is generated.
In praat, it will look a little like this:
Development
The package is quite trivial, but, if you want to work on it, here are some instructions
Style
All code is formatted with the Black code-formatter. As for casing, python standards are used except in cases where dependencies don't.
I am dyslectic, and quite likely to make spelling errors in variables. If you find any, don't hesitate to send me a pull request!
Running Tests
After cloning the repository, moving into it, and installing pytest and pytest-cov with pip, run tests with:
# Install the current version of the package locally to be able to test it.
[python-executable] -m pip install -e .
[python-executable] -m pytest --cov=transcribe_allign_textgrid tests/
Release files for transcribe-allign-textgrid 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| transcribe_allign_textgrid-0.1.5.tar.gz | 21.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| transcribe_allign_textgrid-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size:41.8 kB
Release files / transcribe_allign_textgrid-0.1.5.tar.gz
| Download URL | transcribe_allign_textgrid-0.1.5.tar.gz |
|---|---|
| Size | 21.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5b68dd6c2506ddeb93f42bff51124548ce10152935f5a1fb2106dffb26fd9771
|
|
BLAKE2b-256 checksum How to use checksums |
bda4b614688568e55186c95b9f69022bee9040cb70642ad795275ebbb99f1ffe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.11.6
|
Release files / transcribe_allign_textgrid-0.1.5-py3-none-any.whl
| Download URL | transcribe_allign_textgrid-0.1.5-py3-none-any.whl |
|---|---|
| Size | 20.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ae87dfb448429de3a5f43703ac4bd5c635b7bb33430c2cfd5ce9f2cc355dd8f4
|
|
BLAKE2b-256 checksum How to use checksums |
59665a5d54d4046182135267f8615da65d82fbe2d054a59418a939e8b3eac2e9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.11.6
|