TensorFlowASR :zap:
Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2
TensorFlowASR implements some automatic speech recognition architectures such as DeepSpeech2, Jasper, RNN Transducer, ContextNet, Conformer, etc. These models can be converted to TFLite to reduce memory and computation for deployment :smile:
What's New?
Table of Contents
- What's New?
- Table of Contents
- 😋 Supported Models
- Installation
- Training & Testing Tutorial
- Features Extraction
- Decoders
- Inference
- Augmentations
- TFLite Convertion
- Pretrained Models
- Corpus Sources
- How to contribute
- References & Credits
- Contact
😋 Supported Models
Baselines
- Transducer Models (End2end models using RNNT Loss for training, currently supported Conformer, ContextNet, Streaming Transducer)
- CTCModel (End2end models using CTC Loss for training, currently supported DeepSpeech2, Jasper)
Publications
- Conformer Transducer (Reference: https://arxiv.org/abs/2005.08100) See examples/models/transducer/conformer
- Streaming Conformer (Reference: http://arxiv.org/abs/2010.11395) See examples/models/transducer/conformer
- ContextNet (Reference: http://arxiv.org/abs/2005.03191) See examples/models/transducer/contextnet
- RNN Transducer (Reference: https://arxiv.org/abs/1811.06621) See examples/models/transducer/rnnt
- Deep Speech 2 (Reference: https://arxiv.org/abs/1512.02595) See examples/models/ctc/deepspeech2
- Jasper (Reference: https://arxiv.org/abs/1904.03288) See examples/models/ctc/jasper
Installation
For training and testing, you should use git clone for installing necessary packages from other authors (ctc_decoders, rnnt_loss, etc.)
TensorFlowASR uses uv as its package manager and requires python 3.12 or 3.13.
A plain uv sync installs tensorflow + tensorflow-text, which is all that CPU
and Apple Silicon need. Accelerators are opt-in extras:
git clone https://github.com/TensorSpeech/TensorFlowASR.git
cd TensorFlowASR
uv sync # CPU / Apple Silicon
uv sync --extra cuda # NVIDIA GPU
uv sync --extra dev # add the development tooling
Extras are declared in pyproject.toml:
| Extra | Contents |
|---|---|
| (none) | tensorflow + tensorflow-text, works on CPU and Apple Silicon |
cuda |
tensorflow[and-cuda] for NVIDIA GPUs |
dev |
pytest, ruff, pre-commit, plotting and export tooling |
Run commands inside the environment with uv run, e.g. uv run pytest or
uv run tensorflow_asr --help.
Cloud TPU is not an extra:
uv sync && ./scripts/install_tpu.sh
tensorflow-tpu ships its own tensorflow distribution, and tensorflow is a base
dependency, so an extra adding it would install both and leave whichever landed last in
place. The script uninstalls the stock build first, then installs the TPU one. Re-run it
after any later uv sync, which puts the stock tensorflow back.
TPU note:
tensorflow-tpuships its owntensorflowdistribution and overwrites the base one. This matches the previoussetup.sh tpubehaviour, which uninstalledtensorflowand force-installedtensorflow-tpuover it.
Running in a container
docker-compose up -d
Training & Testing Tutorial
- For training, please read tutorial_training
- For testing, please read tutorial_testing
FYI: Keras builtin training uses infinite dataset, which avoids the potential last partial batch.
See examples for some predefined ASR models and results
Features Extraction
Decoders
Greedy and beam search decoding for CTC and Transducer models, including the ALSD++ transducer beam search
See decoders
Inference
ASRInference transcribes with either a checkpoint or an exported tflite model, in one pass or
streaming chunk by chunk
See inferences, and examples/inferences for runnable scripts including a live microphone
Augmentations
See augmentations
TFLite Convertion
After converting to tflite, the tflite model is like a function that transforms directly from an audio signal to text and tokens
Pretrained Models
See the results on each example folder, e.g. ./examples/models//transducer/conformer/results/sentencepiece/README.md
Corpus Sources
English
| Name | Source | Hours |
|---|---|---|
| LibriSpeech | LibriSpeech | 970h |
| Common Voice | https://commonvoice.mozilla.org | 1932h |
Vietnamese
| Name | Source | Hours |
|---|---|---|
| Vivos | https://ailab.hcmus.edu.vn/vivos | 15h |
| InfoRe Technology 1 | InfoRe1 (passwd: BroughtToYouByInfoRe) | 25h |
| InfoRe Technology 2 (used in VLSP2019) | InfoRe2 (passwd: BroughtToYouByInfoRe) | 415h |
| VietBud500 | https://huggingface.co/datasets/linhtran92/viet_bud500 | 500h |
How to contribute
- Fork the project
- Install for development
- Create a branch
- Make a pull request to this repo
References & Credits
- NVIDIA OpenSeq2Seq Toolkit
- https://github.com/noahchalifour/warp-transducer
- Sequence Transduction with Recurrent Neural Network
- End-to-End Speech Processing Toolkit in PyTorch
- https://github.com/iankur/ContextNet
Contact
Huy Le Nguyen
Email: nlhuy.cs.16@gmail.com
Metadata
Release files for TensorFlowASR 3.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tensorflowasr-3.1.0.tar.gz | 279.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tensorflowasr-3.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 538.4 kB
Release files / tensorflowasr-3.1.0.tar.gz
| Download URL | tensorflowasr-3.1.0.tar.gz |
|---|---|
| Size | 279.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5b64c9b2e7160ae12e659c27fe1c90a3c1ad1993fece1f8a9f5359ee958fbda0
|
|
BLAKE2b-256 checksum How to use checksums |
ca0ea585857f5c67929f65cc7151029e832a9912d7f0d095e7a61da783170693
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / tensorflowasr-3.1.0-py3-none-any.whl
| Download URL | tensorflowasr-3.1.0-py3-none-any.whl |
|---|---|
| Size | 258.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8c2bead5e1d822a6904fd17ca7895d9dd92e7b4f5e31847cb41415bfedfb2d32
|
|
BLAKE2b-256 checksum How to use checksums |
f502d5585cfdbe1fa824f95a9931abaf08a5e864dcbbfa71ea84672702fe185f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|