video-transcriber
Extracts visually distinct frames from videos and transcribes audio using Whisper. Creates a portable zip file containing a markdown transcript with slide images in the img sub-directory.
Features
- Smart slide detection - Uses perceptual hashing to capture only distinct frames, not every frame
- Audio transcription - Uses Whisper AI locally to transcribe speech to text
- Timeline merging - Associates transcribed audio with the corresponding slides
- Portable output - Generates a zip file with markdown and images that works anywhere
Requirements
- Python 3.10+
- ffmpeg (for audio extraction)
# Ubuntu/Debian
apt install ffmpeg
# macOS
brew install ffmpeg
# Windows (using winget)
winget install ffmpeg
Installation
uv venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
uv pip install video-transcriber
Hugging Face
The first time you run the transcriber, it downloads Whisper models from Hugging Face. This may take several minutes.
You may see this warning:
Warning: You are sending unauthenticated requests to the HF Hub.
This is harmless. Everything will still work.
Usage
Python API
from video_transcriber.transcribe import transcribe_video
zip_path = transcribe_video("my-presentation.mp4", "output/")
The output zip file contains:
output/my-presentation_transcript.zip
├── transcript.md
└── img/
├── frame_000.png
├── frame_001.png
└── ...
Options
zip_path = transcribe_video(
"my-presentation.mp4",
"output/",
model_size="large-v3", # Whisper model: tiny, base, small, medium, large-v3
sample_interval=15, # Check for new slides every N frames (default: 30)
include_timestamps=True, # Include timestamps in markdown output (default: False)
audio_only=False # Transcribe audio only, skip frame extraction (default: False)
)
Audio-Only Mode
For videos where you only need the audio transcript (podcasts, interviews, etc.):
zip_path = transcribe_video(
"podcast.mp4",
"output/",
audio_only=True,
include_timestamps=True
)
This skips frame extraction entirely, producing a zip with just the text transcript.
Development
git clone https://github.com/romilly/video-transcriber.git
cd video-transcriber
python -m venv venv
source venv/bin/activate
pip install -e .[test]
pytest
Quick Demo
Run the demo script to process the included test video:
python demo_create_zip.py
This processes tests/data/demo.mp4 and creates tests/data/generated/demo.zip.
Metadata
Release files for video-transcriber 0.3.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| video_transcriber-0.3.1.tar.gz | 14.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| video_transcriber-0.3.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 32.2 kB
Release files / video_transcriber-0.3.1.tar.gz
| Download URL | video_transcriber-0.3.1.tar.gz |
|---|---|
| Size | 14.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4ebab8927117b5ee628c2c49780d0220c9e9e14683b55db2b3a19cb27df9f2b6
|
|
BLAKE2b-256 checksum How to use checksums |
484af119774d096fd46c14bb58f8391a1a6624ab575273085295c07889f65cf1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|
Release files / video_transcriber-0.3.1-py3-none-any.whl
| Download URL | video_transcriber-0.3.1-py3-none-any.whl |
|---|---|
| Size | 17.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cd75ffd31b489fa4cee1a4ddcc4b633d5f4f93316b050ed37dd40d500ff4f0ea
|
|
BLAKE2b-256 checksum How to use checksums |
928fa7de575da72a0cf99e106111b0848e5e055adea1a653a4d55a11c9907473
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|