Skip to main content
Audiogrep
=========

Audiogrep transcribes audio files and then creates "audio supercuts" based on search phrases. It uses [CMU Pocketsphinx](http://cmusphinx.sourceforge.net/) for speech-to-text and [pydub](http://pydub.com/) to stitch things together.

Here's some [sample output](http://lav.io/2015/02/audiogrep-automatic-audio-supercuts/).

## Requirements
Install using pip
```
pip install audiogrep
```
Install [ffmpeg](http://ffmpeg.org/) with Ogg/Vorbis support. If you're on a mac with [homebrew](http://brew.sh/) you can install ffmpeg with:
```
brew install ffmpeg --with-libvpx --with-libvorbis
```
Finally, install [CMU Pocketsphinx](http://cmusphinx.sourceforge.net/). For mac
users I followed [these instructions](https://github.com/watsonbox/homebrew-cmu-sphinx) to get it working:
```
brew tap watsonbox/cmu-sphinx
brew install --HEAD watsonbox/cmu-sphinx/cmu-sphinxbase
brew install --HEAD watsonbox/cmu-sphinx/cmu-sphinxtrain # optional
brew install --HEAD watsonbox/cmu-sphinx/cmu-pocketsphinx
```

## How to use it
First, transcribe the audio (you'll only need to do this once per audio track, but it can take some time)
```
# transcribes all mp3s in the selected folder
audiogrep --input path/to/*.mp3 --transcribe
```
Then, basic use:
```
# returns all phrases with the word 'word' in them
audiogrep --input path/to/*.mp3 --search 'word'
```
The previous example will extract phrase chunks containing the search term, but you can also just get individual words:
```
audiogrep --input path/to/*.mp3 --search 'word' --output-mode word
```
If you add the '--regex' flag you can use regular expressions. For example:
```
# creates a supercut of every instance of the words "spectre", "haunting" and "europe"
audiogrep --input path/to/*.mp3 --search 'spectre|haunting|europe' --output-mode word --regex
```
You can also construct 'frankenstein' sentences (mileage may vary):
```
# stupid joke
audiogrep --input path/to/*.mp3 --search 'my voice is my passport' --output-mode franken
```
Or you can just extract individual words into files.
```
# extracts each individual word into its own file in a directory called 'extracted_words'
audiogrep --input path/to/*.mp3 --extract

Exporting to: extracted_words/i.mp3
Exporting to: extracted_words/am.mp3
Exporting to: extracted_words/the.mp3
Exporting to: extracted_words/key.mp3
Exporting to: extracted_words/master.mp3
```

### Options

audiogrep can take a number of options:

#### --input / -i
mp3 file or pattern for input

#### --output / -o
Name of the file to generate. By default this is "supercut.mp3"

#### --search / -s
Search term

#### --output-mode / -m
Splice together phrases, single words, fragments with wildcards, or "frankenstein" sentences.
Options are:
* sentence: (this is the default)
* word
* fragment
* franken

#### --padding / -p
Time in milliseconds to add between audio segments. Default is 0.

#### --crossfade / -c
Time in milliseconds to crossfade audio segments. Default is 0.

#### --extract / -x

#### --demo / -d
Show the results of the search without outputing a file

Metadata

Release files for audiogrep 0.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for audiogrep 0.1.5
File Size Uploaded
audiogrep-0.1.5.tar.gz 7.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for audiogrep 0.1.5
File Interpreter ABI Platform
audiogrep-0.1.5-py3-none-any.whl Python 3 none any Details

Total release size: 16.6 kB

Release files / audiogrep-0.1.5.tar.gz

Download URL audiogrep-0.1.5.tar.gz
Size 7.6 kB
Tags Source
SHA-256 checksum
How to use checksums
2f1d4142a5f09ede52432c772631de11285b7f651ee8747311380092824a208a
BLAKE2b-256 checksum
How to use checksums
e8aff6ec8b523e78fc00d2bc17a2a8f0129518588e77c17773cfac558731e281
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release files / audiogrep-0.1.5-py3-none-any.whl

Download URL audiogrep-0.1.5-py3-none-any.whl
Size 9.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c4583cee07adf421ec38dbf4b60efb733e58694b2918e34e325e6a1abffbf077
BLAKE2b-256 checksum
How to use checksums
f2a9d4a17342d8e6ecb4ddd72e9da5b7b437abbbfcfc0c25b18c0ebd77fb1aa9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release history Release notifications | RSS feed

This release

0.1.5 This release

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page