A Python library for downloading videos and extracting frames at specified intervals
Project description
video-frame-extractor-cv
A lightweight Python library for downloading videos and extracting frames at precise intervals. It handles direct URLs, supports sub-second extraction, and includes metadata analysis and an interactive player for debugging.
Features
- Direct Download: Process videos directly from URLs without manual downloading.
- Precise Extraction: Support for decimal intervals (e.g., every 0.5 seconds).
- Smart resizing: Resize frames on the fly to save storage.
- Metadata: Automatically extracts FPS, duration, and resolution data.
- Interactive Mode: Optional built-in player to preview or control the process.
- Robust: Includes retry logic, logging, and summary reports.
Installation
pip install video-frame-extractor-cv
For development or building from source:
git clone https://github.com/chibuezedev/video-frame-extractor.git
cd video-frame-extractor
pip install -e .
Quick Start
Command Line Interface (CLI)
The library ships with a video-extractor entry point for quick operations.
# basic: download and extract frames every 5s (default)
video-extractor "https://example.com/video.mp4"
# advanced: extract every 0.5s, resize to width 1280px, skip playback
video-extractor "https://example.com/video.mp4" -i 0.5 -w 1280 --no-play
# specific range: extract from 00:30 to 01:00
video-extractor "https://example.com/video.mp4" -s 30 -e 60
Python Usage
The recommended way to use the library is via the context manager, which handles cleanup automatically.
from video_frame_extractor import VideoFrameExtractor
# use context manager to handle resources automatically
with VideoFrameExtractor("https://example.com/video.mp4") as extractor:
# this downloads the video and extracts metadata
extractor.download_video()
# get info before processing
meta = extractor.get_video_metadata()
print(f"processing {meta['duration_seconds']}s video...")
# run extraction
count = extractor.extract_frames()
print(f"done. extracted {count} frames.")
Advanced Configuration
You can customize the extractor behavior extensively via the constructor.
extractor = VideoFrameExtractor(
video_url="https://example.com/video.mp4",
output_folder="dataset/train",
interval=2.5, # extract frame every 2.5 seconds
quality=90, # jpeg quality (1-100)
max_width=1280, # downscale if width > 1280
start_time=10, # start at 10s mark
end_time=60, # stop at 60s mark
log_level="DEBUG"
)
# run() wraps download, extraction, and reporting in one call
extractor.run(play_video=False, create_report=True)
Output Structure
The library organizes outputs into a clean directory structure:
output_folder/
├── frame_0000_time_0.0s.jpg # extracted frames
├── frame_0001_time_2.5s.jpg
├── video_metadata.json # resolution, fps, source info
├── extraction_report.txt # human-readable summary
└── extraction_log.txt # debug logs
CLI Options Reference
| Flag | Short | Description | Default |
|---|---|---|---|
--output |
-o |
Output directory | frames |
--interval |
-i |
Time between frames (seconds) | 5 |
--quality |
-q |
JPEG quality (1-100) | 95 |
--width |
-w |
Max frame width (px) | None |
--start |
-s |
Start timestamp (seconds) | 0 |
--end |
-e |
End timestamp (seconds) | None |
--no-play |
Disable interactive player | False |
Interactive Player Controls
If you run without --no-play, an OpenCV window will open.
q: Quitp: Pause/Resumer: Restartf/b: Seek forward/back 10s
Recipes
Batch Processing
Process a list of URLs and organize them into separate folders.
urls = [
"https://example.com/clip1.mp4",
"https://example.com/clip2.mp4"
]
for i, url in enumerate(urls):
folder = f"data/clip_{i}"
# initialize and run in one go
extractor = VideoFrameExtractor(url, output_folder=folder)
if extractor.run(play_video=False):
print(f"finished {url}")
else:
print(f"failed {url}")
Scene Extraction
Extract frames from specific time ranges within a single video.
# (start_time, end_time) tuples
scenes = [(30, 60), (120, 180)]
for start, end in scenes:
extractor = VideoFrameExtractor(
"https://example.com/movie.mp4",
output_folder=f"frames/{start}_{end}",
start_time=start,
end_time=end,
interval=1.0
)
extractor.run(play_video=False)
Individual Operations
from video_frame_extractor import VideoFrameExtractor
extractor = VideoFrameExtractor("https://example.com/video.mp4", interval=1.0)
# download only
if extractor.download_video():
print("Video downloaded successfully")
# metadata
metadata = extractor.get_video_metadata()
print(f"Video duration: {metadata.get('duration_seconds', 0):.1f} seconds")
# extract frames without playing
frames_extracted = extractor.extract_frames()
print(f"Extracted {frames_extracted} frames")
# play video separately
extractor.play_video(show_controls=True)
# create report
report_path = extractor.create_summary_report()
print(f"Report saved to: {report_path}")
Using the Video Player Separately
from video_frame_extractor import VideoPlayer
player = VideoPlayer()
player.play("path/to/video.mp4", start_time=10, end_time=60)
Utility Functions
from video_frame_extractor import validate_url, sanitize_filename
# validate video URL
is_valid = validate_url("https://example.com/video.mp4")
print(f"URL is valid: {is_valid}")
# clean filename
clean_name = sanitize_filename("my video [1080p].mp4")
print(f"Clean filename: {clean_name}")
Output Files
The library creates several output files:
output_folder/
├── frame_0000_time_0.0s.jpg # Extracted frames
├── frame_0001_time_5.0s.jpg
├── frame_0002_time_10.0s.jpg
├── ...
├── video_metadata.json # Video information
├── extraction_report.txt # Summary report
└── extraction_log.txt # Detailed logs
Metadata JSON Structure
{
"source_url": "https://example.com/video.mp4",
"fps": 30.0,
"total_frames": 1800,
"width": 1920,
"height": 1080,
"duration_seconds": 60.0,
"extraction_interval": 5.0,
"extraction_time": "2024-01-15T10:30:00",
"start_time": 0,
"end_time": null,
"quality": 95,
"max_width": null,
"frames_extracted": 12
}
Error Handling
from video_frame_extractor import VideoFrameExtractor
try:
extractor = VideoFrameExtractor("https://invalid-url.com/video.mp4")
success = extractor.run()
if not success:
print("Extraction failed - check logs for details")
except Exception as e:
print(f"Unexpected error: {e}")
API Reference
VideoFrameExtractor Class
Constructor Parameters
video_url(str): URL of the video to download and processoutput_folder(str, optional): Directory to save extracted frames (default: "frames")interval(float, optional): Time interval in seconds between frame extractions (default: 5.0)quality(int, optional): JPEG quality for saved frames, 1-100 (default: 95)max_width(int, optional): Maximum width for extracted frames (default: None)start_time(float, optional): Start time in seconds for extraction (default: 0)end_time(float, optional): End time in seconds for extraction (default: None)log_level(str, optional): Logging level (default: "INFO")
Methods
download_video(timeout=30, chunk_size=8192): Download video from URLget_video_metadata(): Extract and return video metadataextract_frames(): Extract frames at specified intervalsplay_video(show_controls=True): Play the downloaded videocreate_summary_report(): Generate extraction reportrun(play_video=True, create_report=True): Execute complete process
VideoPlayer Class
Methods
play(video_path, start_time=0, end_time=None, show_controls=True): Play video file
Utility Functions
validate_url(url, timeout=10): Check if URL is accessible videosanitize_filename(filename): Clean filename for filesystem compatibilitysetup_logging(output_folder, log_level="INFO"): Configure logging
Troubleshooting
OpenCV Errors:
If you see errors related to cv2 or libGL, you might need the headless version of OpenCV for server environments:
pip install opencv-python-headless
Download Failures:
Ensure the URL is a direct link to a file (ends in .mp4, .avi, etc). For YouTube links, use a tool like yt-dlp to get the direct stream URL first.
Contributing
- Fork the repo
- Create your feature branch (
git checkout -b feature/cool-feature) - Commit changes (
git commit -m 'add cool feature') - Push to branch (
git push origin feature/cool-feature) - Open a Pull Request
For support, please:
- Check the troubleshooting section
- Search existing issues
- Create a new issue
License
Distributed under the MIT License. See LICENSE for more information.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file video_frame_extractor_cv-0.1.4.tar.gz.
File metadata
- Download URL: video_frame_extractor_cv-0.1.4.tar.gz
- Upload date:
- Size: 19.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
82af3e40519041382defa4fefd3b55cd3f218ac47e7e753f6928cb52624bfb42
|
|
| MD5 |
ada2cd85dd8568accd52061193f6e4a1
|
|
| BLAKE2b-256 |
46119cbaaf7322690550fc09ca901a213e7f3e26cb12872bd5581a75c682b150
|
Provenance
The following attestation bundles were made for video_frame_extractor_cv-0.1.4.tar.gz:
Publisher:
python-publish.yml on chibuezedev/video-frame-extractor
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
video_frame_extractor_cv-0.1.4.tar.gz -
Subject digest:
82af3e40519041382defa4fefd3b55cd3f218ac47e7e753f6928cb52624bfb42 - Sigstore transparency entry: 732430515
- Sigstore integration time:
-
Permalink:
chibuezedev/video-frame-extractor@2aa805476d773283acf715bfe0263c0b1e78ef93 -
Branch / Tag:
refs/tags/v.0.1.5 - Owner: https://github.com/chibuezedev
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@2aa805476d773283acf715bfe0263c0b1e78ef93 -
Trigger Event:
release
-
Statement type:
File details
Details for the file video_frame_extractor_cv-0.1.4-py3-none-any.whl.
File metadata
- Download URL: video_frame_extractor_cv-0.1.4-py3-none-any.whl
- Upload date:
- Size: 14.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
83feacee88bfe18e01ee18d944dbadfe7c6c37e1d2e62e1e9d19cf4cd9b0c01c
|
|
| MD5 |
0c282758fbc18246854861cc224c8191
|
|
| BLAKE2b-256 |
3ec8fe30ffc330dcd6301bf6fdd328691e9de10274bfb087e2baa4d3d79e0bc6
|
Provenance
The following attestation bundles were made for video_frame_extractor_cv-0.1.4-py3-none-any.whl:
Publisher:
python-publish.yml on chibuezedev/video-frame-extractor
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
video_frame_extractor_cv-0.1.4-py3-none-any.whl -
Subject digest:
83feacee88bfe18e01ee18d944dbadfe7c6c37e1d2e62e1e9d19cf4cd9b0c01c - Sigstore transparency entry: 732430522
- Sigstore integration time:
-
Permalink:
chibuezedev/video-frame-extractor@2aa805476d773283acf715bfe0263c0b1e78ef93 -
Branch / Tag:
refs/tags/v.0.1.5 - Owner: https://github.com/chibuezedev
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@2aa805476d773283acf715bfe0263c0b1e78ef93 -
Trigger Event:
release
-
Statement type: