Skip to main content

A radio streaming and transcription application.

Project description

# README for audio_miner

![Test Status](https://github.com/smilchsack/audio_miner/actions/workflows/ci.yml/badge.svg)

## audio_miner

audio_miner is a Python application that records audio streams and transcribes them using the Whisper model. It allows users to capture radio broadcasts and convert them into text files for easy access and analysis.

### Features

  • Record audio streams in specified segments.

  • Transcribe recorded audio using the Whisper model.

  • Organize recordings and transcriptions in a structured directory.

  • Optional speaker diarization using PyAnnote (requires a Hugging Face token).

### Installation

To install audio_miner, you can use pip. Clone the repository and run the following command:

`bash pip install . `

Additionally, you need to have ffmpeg installed on your system. You can install it using the following command:

`bash sudo apt-get install ffmpeg `

### Usage

After installation, you can run audio_miner from the command line. Use the following command format:

`bash audio_miner --stream-url <STREAM_URL> --sender <SENDER_NAME> [--segment-time <SEGMENT_TIME>] [--base-dir <BASE_DIR>] [--poll-interval <POLL_INTERVAL>] [--whisper-model <WHISPER_MODEL>] [--quality <QUALITY>] [--record-only] [--transcribe-only] [--verbose] [--ffmpeg-path <FFMPEG_PATH>] `

#### Parameters

  • –stream-url: The URL of the audio stream to record (required).

  • –sender: The name of the radio station (required).

  • –segment-time: Length of each audio segment in seconds (default: 3600).

  • –base-dir: Base directory for storing audio and transcription files (default: current directory).

  • –poll-interval: Interval in seconds between recordings (default: 5).

  • –whisper-model: The Whisper model to use for transcription (default: TURBO; options include TINY, BASE, SMALL, MEDIUM, LARGE, TURBO). Warning: TURBO requires ~6GB VRAM and LARGE requires ~10GB VRAM. More info: [https://github.com/openai/whisper](https://github.com/openai/whisper)

  • –quality: Audio bitrate for re-encoding (e.g., 64k). If not specified, the original stream quality will be copied.

  • –start-time: Start time for transcription in YYYYMMDD_HHMMSS format. Only relevant when using –transcribe-only.

  • –end-time: End time for transcription in YYYYMMDD_HHMMSS format. Only relevant when using –transcribe-only.

  • –token: Hugging Face token for PyAnnote speaker diarization model (optional). If provided, diarization will be performed.

  • –record-only: Record audio without transcribing.

  • –transcribe-only: Transcribe existing audio files without recording.

  • –verbose: Enable detailed output.

  • –ffmpeg-path: Path to the ffmpeg executable. This is only necessary if ffmpeg cannot be started directly from the terminal.

### Example

To record from a stream and transcribe it, you can use:

`bash audio_miner --stream-url 'https://liveradio.swr.de/sw282p3/swr1rp/' --sender 'swr1' --segment-time 300 --base-dir './output' --poll-interval 5 --whisper-model TURBO `

### Contributing

Contributions are welcome! Please feel free to submit a pull request or open an issue for any suggestions or improvements.

### License

This project is licensed under the MIT License. See the LICENSE file for more details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

audio_miner-0.0.6.tar.gz (16.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

audio_miner-0.0.6-py3-none-any.whl (19.7 kB view details)

Uploaded Python 3

File details

Details for the file audio_miner-0.0.6.tar.gz.

File metadata

  • Download URL: audio_miner-0.0.6.tar.gz
  • Upload date:
  • Size: 16.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.17

File hashes

Hashes for audio_miner-0.0.6.tar.gz
Algorithm Hash digest
SHA256 e49cd03a7478e5146d63ecd77aa956e0192eeaffd1b84b59cfdaac5f2d9720ee
MD5 fb0b0f3909684391fc74e8c0bd261178
BLAKE2b-256 e6872f2356e5152a38dcae67b4f1e0f0f975a41924349a68998d8ec077c52cca

See more details on using hashes here.

File details

Details for the file audio_miner-0.0.6-py3-none-any.whl.

File metadata

  • Download URL: audio_miner-0.0.6-py3-none-any.whl
  • Upload date:
  • Size: 19.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.17

File hashes

Hashes for audio_miner-0.0.6-py3-none-any.whl
Algorithm Hash digest
SHA256 6c71491558ef9e84221d161e9ed763578ded5ecbb4aa3c2a4da37b307f4c6d15
MD5 201b6143283ecebf7a76c8a63763772a
BLAKE2b-256 ab9cb0b9d3bf81de070c00a48e0cc0aa6b43d37b14e385e7cb46745d15c9932a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page