Video Offset Finder
Find the temporal offset between two videos using perceptual hashing or direct pixel comparison (SAD).
Contents:
- Why Do We Need This?
- Requirements and Installation
- Usage
- Output Format
- How Does It Work?
- API
- License
Why Do We Need This?
I've too often encountered slightly offset video files, which are a pain to sync for calculating full-reference video quality metrics (like VMAF). Based on an earlier, PSNR-based Python script, this is now a fully-featured – and much faster! – tool to find the temporal offset between two videos.
This tool is generally useful for:
- Synchronizing videos from different sources
- A/V sync analysis
- Video quality comparison (aligning reference and test videos)
- Finding where a clip appears in a longer video
The default algorithm uses perceptual hashing and therefore is robust to:
- Different resolutions
- Different quality/compression levels
- Color grading differences
- Minor geometric distortions
Requirements and Installation
Using uv:
uvx video-offset-finder
Using pipx:
pipx install video-offset-finder
Or, using pip:
pip install video-offset-finder
Usage
Let's say you have two video files, reference.mp4 and distorted.mp4, and you want to find the temporal offset between them. You can use the command-line tool as follows:
# Find offset between reference and distorted/delayed video
uvx video-offset-finder reference.mp4 distorted.mp4
# With hints about expected offset (faster)
uvx video-offset-finder ref.mp4 dist.mp4 --start-offset 10 --max-search-offset 15
# Verbose output
uvx video-offset-finder ref.mp4 dist.mp4 -v
The tool will output JSON with the detected offset and confidence score. For the output format, see Output Format.
Full usage:
usage: video-offset-finder [-h] [-t {phash,dhash,ahash,whash,sad}]
[--hash-size HASH_SIZE] [--coarse-fps COARSE_FPS]
[--fine-fps FINE_FPS] [-o START_OFFSET]
[-s MAX_SEARCH_OFFSET] [-m MAX_DURATION]
[--refine-window REFINE_WINDOW] [-v] [-q] [--version]
ref dist
positional arguments:
ref Reference video
dist Distorted/delayed video
options:
-h, --help show this help message and exit
-t, --compare-type {phash,dhash,ahash,whash,sad}
Comparison algorithm: phash (default, best quality), dhash
(fast), ahash (fastest), whash (most robust), sad (direct
pixel comparison)
--hash-size HASH_SIZE
Hash size in bits (default: 16, larger = more precise)
--coarse-fps COARSE_FPS
FPS for coarse search (default: 1.0)
--fine-fps FINE_FPS FPS for fine search (default: 10.0)
-o, --start-offset START_OFFSET
Minimum offset to search in seconds (default: unlimited)
-s, --max-search-offset MAX_SEARCH_OFFSET
Maximum offset to search in seconds (default: unlimited)
-m, --max-duration MAX_DURATION
Maximum duration to analyze in seconds (default: unlimited)
--refine-window REFINE_WINDOW
Window size around coarse result for refinement (default: 2.0s)
-v, --verbose Enable debug logging
-q, --quiet Suppress progress bars and logging (only output JSON)
--version show program's version number and exit
Output Format
The tool outputs JSON to stdout:
{
"date": "2025-01-09T20:15:30.123456",
"reference": "reference.mp4",
"distorted": "distorted.mp4",
"offset_frames": 150,
"offset_seconds": 5.005,
"offset_timestamp": "00:00:05.005",
"confidence": 2.34,
"second_best_confidence": 12.81,
"overlap_frames": 91,
"fps_used": 29.97,
"method": "frame_accurate_phash",
"settings": {
"compare_type": "phash",
"hash_size": 16,
"coarse_fps": 1.0,
"fine_fps": 10.0,
"start_offset": null,
"max_search_offset": null,
"max_duration": null,
"refine_window": 2.0,
"compute_time": 12.45
}
}
The fields are as follows:
| Field | Description |
|---|---|
offset_frames |
Offset in frames (at fps_used rate) |
offset_seconds |
Offset in seconds |
offset_timestamp |
Offset in HH:MM:SS.sss format |
confidence |
Average distance (lower = better match, 0 = identical). Hamming distance for hash algorithms, SAD for pixel comparison. |
second_best_confidence |
Distance of the second-best candidate, useful for judging ambiguity. |
overlap_frames |
Number of frames compared for the selected candidate. |
fps_used |
Frame rate used for final measurement |
method |
Algorithm used for final result |
compute_time |
Processing time in seconds |
How Does It Work?
This section explains the frame comparison methods, the overall search algorithm, and visualizes how the search parameters affect the process.
Hashing/Comparison Algorithms
There are different algorithms available for comparing frames, each with their own trade-offs:
| Algorithm | Speed | Robustness | Best For |
|---|---|---|---|
phash |
Medium | High | General use (default) |
dhash |
Fast | Medium | Fast processing |
ahash |
Fastest | Lower | Very fast estimates |
whash |
Slowest | Highest | Difficult comparisons |
sad |
Fast | Medium | Identical/similar quality videos |
The first four are "perceptual hash" algorithms from the ImageHash library:
- phash (Perceptual Hash): Applies a Discrete Cosine Transform (DCT) to capture low-frequency components, similar to JPEG compression. Most robust to scaling and minor edits.
- dhash (Difference Hash): Compares the brightness of adjacent pixels horizontally. Fast and effective for detecting shifts/translations.
- ahash (Average Hash): Compares each pixel to the average brightness of the image. Simplest and fastest, but less robust to changes.
- whash (Wavelet Hash): Uses Haar wavelet decomposition for multi-resolution analysis. Most robust to compression artifacts and color changes.
All hash algorithms reduce an image to a compact binary fingerprint. For more details, see the ImageHash library documentation.
The last algorithm is direct pixel comparison:
- sad (Sum of Absolute Differences): Directly compares pixel values between frames after resizing both inputs to 64x64 grayscale. It is fast and effective when videos have similar quality/encoding, but less robust to compression artifacts or color grading differences than perceptual hashes.
Overall Flow
The tool uses a hierarchical coarse-to-fine search. For ordinary clips, it decodes and hashes each input once at the highest required cadence, then selects timestamped subsets from that cache for each pass:
- Coarse pass (1 fps): Compute signatures for both videos at low frame rate, find approximate offset via cross-correlation
- Fine pass (10 fps): Compute signatures only within a ±2s window around the coarse result, refine the offset
- Frame-accurate pass (native fps): Compute signatures within a ±0.5s window around the fine result for exact frame matching
For very long videos, the tool reads only the section needed for each search step to limit memory use. It uses each frame's timestamp to choose samples. If no new frame exists for a sample time, it uses the previous frame again. The decoder resizes frames before hashing to save work. Wavelet hashing keeps the original frame size because resizing it first would change the hash.
For each allowed offset, the tool compares the overlapping frames and averages their differences. At least half of the shorter sequence must overlap. This stops a single matching frame at the edge from winning. Hash modes count different bits, while SAD adds up pixel differences. The result includes the best score, the second-best score, and the number of frames compared.
Search Parameters Visualized
The following diagrams show how the offset detection and search parameters work.
Default Case: Cut Video Within Source
The most common scenario: a shorter "distorted" video is a clip extracted from the longer "reference" video:
Reference (source):
|======================================================|
0s 60s
Distorted (cut):
|=================|
15s 35s
↑
└── offset = 15s (positive: distorted
starts later in timeline)
Result: offset_seconds = 15.0
Negative Offset: Distorted Starts Earlier
When the distorted video contains content that appears before the reference:
Reference:
|==============================|
10s 50s
Distorted:
|============================================|
0s 40s
↑
└── offset = -10s (negative: distorted starts earlier in timeline)
Result: offset_seconds = -10.0
Using --start-offset to Skip Reference Start
If you know the match is not before N seconds, use -o/--start-offset to set the minimum candidate offset:
Reference (60s total):
|======================================================|
0s 60s
With --start-offset 20, frames extracted from reference:
|xxxxxxxxxxxxxxxxxxxx|=================================|
0s (not extracted) 20s 60s
Distorted (20s clip that matches at 30s):
|=================|
30s 50s
Offset found = 30s
Matches before 20s cannot be returned. On long inputs that use phase-specific extraction, this also avoids decoding the beginning of the reference.
Using --max-search-offset to Bound the Search
Use -s/--max-search-offset to set the maximum candidate offset:
Reference (60s), Distorted (20s), --max-search-offset 25:
Reference frames extracted (25s + 20s = 45s):
|==========================================|xxxxxxxxxxx|
0s 45s 60s
(not extracted)
Distorted query:
|===================|
0s 20s
Only candidates at or before 25s are considered. The reference only needs to cover the search range plus the distorted query duration.
Using --max-duration to Limit Analysis Length
Use -m/--max-duration to limit the query duration used from both videos:
Reference (60s), --max-duration 30:
Reference and distorted query interval:
|==============================|xxxxxxxxxxxxxxxxxxxxxxxxxxx|
0s 30s 60s
(not extracted)
This is useful when a shorter excerpt contains enough distinctive content to locate the match.
API
Use as a library in your Python code:
from pathlib import Path
from video_offset_finder import find_offset, CompareType
# Basic usage
result = find_offset(
ref_path=Path("reference.mp4"),
dist_path=Path("distorted.mp4"),
)
print(f"Offset: {result.offset_seconds:.3f}s ({result.offset_frames} frames)")
# With options (using perceptual hash)
result = find_offset(
ref_path=Path("reference.mp4"),
dist_path=Path("distorted.mp4"),
compare_type=CompareType.DHASH, # Faster hash algorithm
coarse_fps=2.0, # More samples in coarse pass
fine_fps=15.0, # Higher precision in fine pass
start_offset=5.0, # Known minimum offset
max_search_offset=20.0, # Limit search range
max_duration=60.0, # Only analyze first 60s
frame_accurate=True, # Final pass at native FPS
quiet=True, # Suppress progress bars
)
# Using SAD (direct pixel comparison)
result = find_offset(
ref_path=Path("reference.mp4"),
dist_path=Path("distorted.mp4"),
compare_type=CompareType.SAD, # Sum of Absolute Differences
)
Available Functions
from video_offset_finder import (
# Main function
find_offset,
# Models
CompareType, # Enum: PHASH, DHASH, AHASH, WHASH, SAD
VideoInfo, # Dataclass with video metadata
OffsetResult, # Dataclass with detection result
CorrelationResult, # Detailed correlation result
# Video utilities
get_video_info, # Extract video metadata
extract_frames, # Generator yielding (timestamp, PIL.Image) tuples
# Comparison utilities
compute_hash, # Compute perceptual hash for a single image
compute_sad_signature, # Compute SAD signature for a single image
compute_video_signatures, # Compute signatures for all frames in a video
cross_correlate_signatures, # Find best alignment between signature sequences
cross_correlate_signatures_detailed, # Include second-best score and overlap
)
OffsetResult Fields
@dataclass
class OffsetResult:
offset_frames: int # Offset in frames
offset_seconds: float # Offset in seconds
confidence: float # Distance metric (lower = better)
fps_used: float # FPS used for measurement
method: str # Algorithm identifier
second_best_confidence: float | None # Runner-up distance
overlap_frames: int # Frames compared for the selected candidate
License
MIT License
Copyright (c) 2025 Werner Robitza
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
Metadata
Release files for video-offset-finder 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| video_offset_finder-0.4.0.tar.gz | 16.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| video_offset_finder-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 37.0 kB
Release files / video_offset_finder-0.4.0.tar.gz
| Download URL | video_offset_finder-0.4.0.tar.gz |
|---|---|
| Size | 16.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
731767bc82e510d80c6e0276b58218a66d626462c91adbcb7309d12910e5c886
|
|
BLAKE2b-256 checksum How to use checksums |
43da51c42b96c9c64375402c1aaeb804fe143423514cffa51588f83ee1846650
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / video_offset_finder-0.4.0-py3-none-any.whl
| Download URL | video_offset_finder-0.4.0-py3-none-any.whl |
|---|---|
| Size | 20.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
43c8dd7cb7e73991c6bd25ae628a5298d300b5ebfa7cee536577299e08e0d7ab
|
|
BLAKE2b-256 checksum How to use checksums |
c22b4495f9094321796e74ae9f32959e6dc0cf740b501bb7b27f18aad06853b9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|