Skip to main content

speech-recognition-inference

This repo provides a command-line tool for performing automatic speech-to-text tasks (i.e., "transcription") using open source models from Hugging Face Hub. For interactive tasks, it allows users to spin up an inference API. For bulk processing, it furnishes a pipeline for running inference on the contents of a specified directory.

Quickstart

Pip

First, install ffmpeg if it is not already installed on your machine.

Next, clone the repository and install it:

git clone https://github.com/princeton-ddss/speech-recognition-inference.git
cd speech-recognition-inference
python -m venv .venv
pip install --upgrade pip
pip install . # or pip install -e . for development

To start an API from the command-line, simply run:

speech_recognition launch \
  --port 8000:8000 \
  --model-id openai/whisper-tiny \
  --model-dir $HOME/.cache/huggingface/hub

Once the application startup is complete, you can submit requests using any HTTP request library or tool, e.g.,

curl localhost:8000/transcribe \
  -X POST \
  -d '{"audio_file": "/tmp/female.wav", "response_format": "json"}' \
  -H 'Content-Type: application/json'

To run batch processing, run:

speech_recognition pipeline /data/audio \
  --model-id openai/whisper-tiny \
  --model-dir $HOME/.cache/huggingface/hub

Docker

We also provide a Docker image, ghcr.io/princeton-ddss/speech-recognition-inference. To run the API via Docker use the following command:

docker run \
  -p 8000:8000 \
  -v $HOME/.cache/huggingface/hub:/data/models \
  -v /tmp:/data/audio \
  ghcr.io/princeton-ddss/speech-recognition-inference:latest \
  launch \
  --port 8000 \
  --model_id openai/whisper-large-v3 \
  --model_dir /data/models

This command makes the API available on localhost:8000 and makes host model and audio files available to the container via bind mounting. Sending requests is the same as above, but note that the container only has access to bind mounted files. Above, this means that requests should replace /tmp with /data/audio:

curl localhost:8000/transcribe \
  -X POST \
  -d '{"audio_file": "/data/audio/female.wav", "response_format": "json"}' \
  -H 'Content-Type: application/json'

To run batch processing via Docker, replace the launch command with pipeline and update the options:

docker run \
  -v $HOME/.cache/huggingface/hub:/data/models \
  -v /tmp:/data/audio \
  ghcr.io/princeton-ddss/speech-recognition-inference:latest \
  pipeline \
  /data/audio \
  --model_id openai/whisper-large-v3 \
  --model_dir /data/models

Again, note that /tmp is bound to /data/audio on the host, so this command runs inference on all audio files in /tmp on the host.

Detailed Usage

Full usage details are available via the --help option. For example,

❯ speech_recognition --help

 Usage: speech_recognition [OPTIONS] COMMAND [ARGS]...

╭─ Options ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ --install-completion          Install completion for the current shell.                                                                          │
│ --show-completion             Show completion for the current shell, to copy it or customize the installation.                                   │
│ --help                        Show this message and exit.                                                                                        │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭─ Commands ───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ launch     Launch a speech-to-text API.                                                                                                          │
│ pipeline   Perform batch speech-to-text inference.                                                                                               │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

Environment Variables

In addition to specifying settings at the command line, some settings can be provided through environment variables. These settings are:

  • SRI_HOST - Host to use for the API.
  • SRI_PORT - Port to use for the API.
  • SRI_TOKEN - Authentication token to use for authenticated API access.
  • HF_ACCESS_TOKEN - Authentication token to use for authentication with the Hugging Face Hub API.

Users can set environment variables by exporting or via a .env file. Command line arguments always take precedence over environment variables.

Release files for speech-recognition-inference 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for speech-recognition-inference 0.2.0
File Size Uploaded
speech_recognition_inference-0.2.0.tar.gz 131.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for speech-recognition-inference 0.2.0
File Interpreter ABI Platform
speech_recognition_inference-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 146.3 kB

Release files / speech_recognition_inference-0.2.0.tar.gz

Download URL speech_recognition_inference-0.2.0.tar.gz
Size 131.8 kB
Tags Source
SHA-256 checksum
How to use checksums
278286483b822832558112b0ac4c6523b9befe491e3a09ab3dbf126102a368e7
BLAKE2b-256 checksum
How to use checksums
d8cc272da91a8fc85d5a108994e3b7f7e8b7520364430615b53d03e6b3d4e780
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.6.6

Release files / speech_recognition_inference-0.2.0-py3-none-any.whl

Download URL speech_recognition_inference-0.2.0-py3-none-any.whl
Size 14.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
396f2a9a0970b1085d0341bb61e2765dc6eed9d61029824921415a918c889670
BLAKE2b-256 checksum
How to use checksums
7012db493711cd8bd5f7497ffadc4a5686b5a450a0ccdadf525620497537e427
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.6.6

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page