depth_estimation
A unified Python library for monocular depth estimation
Inference · Video · Training · Evaluation · ONNX Export · Pruning · Quantization
depth_estimation is a unified model framework for monocular depth estimation. It provides one consistent API across 12 model families and 28 variants, so you can swap models, compare them, fine-tune them, and prepare them for deployment without rewriting your pipeline.
It covers the workflow end-to-end: run inference with one line, process video, visualize results, evaluate standard benchmarks, fine-tune on custom depth data, optimize models, and export them to ONNX.
Latest release: v0.1.3 — Documentation Refresh. It adds clearer installation paths, a documentation index, updated deployment guidance, and corrected ONNX verification dependencies. Read the release notes or view the GitHub release.
Installation
# Core library
pip install --upgrade depth-estimation
# ONNX export
pip install "depth-estimation[export]"
# Export verification and ONNX quantization
pip install "depth-estimation[export]" onnxruntime
Python 3.10–3.12 is supported. For GPU execution of exported models, replace onnxruntime with onnxruntime-gpu. See docs/dependencies.md for the complete dependency guide.
Documentation
| Goal | Guide |
|---|---|
| Choose a model | Models |
| Run video or visualize results | Video · Visualization |
| Load data, evaluate, or fine-tune | Data · Evaluation · Training |
| Prepare a model for deployment | ONNX export · Pruning · Quantization |
| Use the command line | CLI reference |
| Extend the library | Adding a model |
Quickstart
The pipeline API is the fastest way to get a depth map from any image:
from depth_estimation import pipeline
pipe = pipeline("depth-estimation", model="depth-anything-v2-vitb")
result = pipe("image.jpg")
depth_map = result.depth # np.ndarray, float32, (H, W)
colored = result.colored_depth # np.ndarray, uint8, (H, W, 3)
For full control over each step — preprocessing, forward pass, postprocessing — use Auto Classes:
from depth_estimation import AutoDepthModel, AutoProcessor
import torch
model = AutoDepthModel.from_pretrained("zoedepth")
processor = AutoProcessor.from_pretrained("zoedepth")
inputs = processor("image.jpg")
with torch.no_grad():
depth = model(inputs["pixel_values"])
result = processor.postprocess(depth, inputs["original_sizes"])
Or from the command line:
depth-estimate predict image.jpg --model depth-anything-v2-vitb
Why use depth_estimation?
1. One API, every model. Switch from Depth Anything to DepthPro to MoGe by changing a single string. Preprocessing, postprocessing, and output format are identical across all models.
2. The full depth workflow in one place. Most libraries stop at inference. This one also covers video, training, benchmark evaluation, dataset loading, pruning, quantization, and ONNX export, so you don't have to stitch together separate tools.
3. Modular, single-file model design.
Each model lives in one self-contained file. No hidden abstractions. If you need to understand or modify a model, there's exactly one place to look. New models self-register — AutoDepthModel and pipeline() resolve them automatically.
4. Designed for research.
Trainable models with backbone freeze schedules, proper batch-level metric accumulation (no mean-of-means), and a compare() function that shows a formatted table across models.
5. Built for reliable deployment. Export verification checks ONNX output against PyTorch, quantization verifies accuracy by default, and unsupported model paths fail early with actionable errors.
Supported Models
12 model families · 28 variants — see docs/models.md for the full list.
All models support inference and CLI. The Trainable column indicates fine-tuning support via DepthTrainer.
| Family | Variants | Depth type | Trainable |
|---|---|---|---|
| Depth Anything v1 | vits / vitb / vitl | Relative | ✅ |
| Depth Anything v2 | vits / vitb / vitl | Relative | ✅ |
| Depth Anything v3 | small / base / large / giant / mono / metric | Relative + Metric | ✅ |
| Depth Anything v3 Nested | nested-giant-large | Relative | ✅ |
| ZoeDepth | nyu / kitti | Metric | ❌ |
| MiDaS | dpt-large / dpt-hybrid / beit-large | Relative | ✅ |
| Apple DepthPro | — | Metric | ✅ |
| Pixel-Perfect Depth | — | Relative | ❌ |
| Marigold-DC | — | Relative (depth completion) | ❌ |
| MoGe | v1 vitl / v2 vitl / v2 vitb / v2 vits (+ normal variants) | Metric | ❌ |
| OmniVGGT | vitl | Metric | ✅ |
| VGGT | standard / commercial | Metric | ✅ |
What can you do?
Inference — single image, batch, or video
# Single image
result = pipe("image.jpg")
# Batch
results = pipe(["img1.jpg", "img2.jpg"], batch_size=2)
# CLI — batch predict
depth-estimate predict "images/*.jpg" --model depth-anything-v2-vitb --output-dir results/
Video & Streaming — frame-by-frame depth from video, webcam, or image sequences
from depth_estimation import pipeline
pipe = pipeline("depth-estimation", model="depth-anything-v2-vitb")
# Stream a video file — yields DepthOutput per frame
for result in pipe.stream("video.mp4", temporal_smoothing=0.5):
depth = result.depth # (H, W) float32
colored = result.colored_depth # (H, W, 3) uint8
print(result.metadata["frame_index"])
# Webcam stream
for result in pipe.stream(0): # device index
...
# Frame glob (sorted alphabetically)
for result in pipe.stream("frames/*.png"):
...
# Write output video to disk
pipe.process_video(
"input.mp4",
"output_depth.mp4",
colormap="inferno",
side_by_side=True, # RGB | depth composite
temporal_smoothing=0.5,
)
# CLI — video prediction
depth-estimate predict video.mp4 --model depth-anything-v2-vitb --output depth_video.mp4
See docs/video.md.
Visualization — depth maps, comparisons, overlays, 3D animations, error maps
from depth_estimation.viz import (
show_depth, compare_depths, overlay_depth,
create_anaglyph, animate_3d, plot_error_map,
)
# Display a depth result
show_depth(result, colormap="Spectral_r", title="Depth Anything V2")
# Side-by-side comparison of multiple models
compare_depths([result_v2, result_pro], labels=["DA V2", "DepthPro"], save="compare.png")
# Blend depth over RGB image
overlay = overlay_depth(image, result.depth, alpha=0.5, colormap="inferno")
# Red-cyan anaglyph stereo image
anaglyph = create_anaglyph(image, result.depth, baseline=0.065)
# Rotating 3D surface animation
animate_3d(image, result.depth, "rotation.gif", frames=60)
# Per-pixel error heatmap (requires ground truth)
plot_error_map(pred_depth, gt_depth, metric="abs_rel", save="errors.png")
See docs/viz.md.
Evaluation — standard benchmarks, custom predictions
from depth_estimation.evaluation import evaluate, compare, Evaluator
# Single model on NYU Depth V2
results = evaluate("depth-anything-v2-vitb", "nyu_depth_v2", split="test")
# Compare multiple models — prints table with best values marked (*)
compare(["depth-anything-v2-vits", "depth-anything-v2-vitb"], dataset="nyu_depth_v2")
# Accumulate metrics over your own dataloader
ev = Evaluator()
for pred, gt, mask in dataloader:
ev.update(pred, gt, mask)
final = ev.compute() # abs_rel, sq_rel, rmse, rmse_log, delta1/2/3
See docs/evaluation.md.
Fine-Tuning — any trainable model, any depth dataset
from depth_estimation import DepthTrainer, DepthTrainingArguments, load_dataset
from depth_estimation.models.depth_anything_v2 import DepthAnythingV2Model
from depth_estimation.data.transforms import get_train_transforms, get_val_transforms
model = DepthAnythingV2Model.from_pretrained("depth-anything-v2-vits", for_training=True)
train_ds = load_dataset("nyu_depth_v2", split="train", transform=get_train_transforms(518))
val_ds = load_dataset("nyu_depth_v2", split="test", transform=get_val_transforms(518))
args = DepthTrainingArguments(output_dir="./checkpoints", num_epochs=25, batch_size=8,
freeze_backbone_epochs=5, mixed_precision=True)
DepthTrainer(model=model, args=args, train_dataset=train_ds, eval_dataset=val_ds).train()
Any torch.utils.data.Dataset returning pixel_values / depth_map / valid_mask works directly — no subclassing needed. See docs/training.md.
Dataset Loading — standard benchmarks, custom folders
from depth_estimation import load_dataset
ds = load_dataset("nyu_depth_v2", split="test") # auto-downloads ~2.8 GB
ds = load_dataset("diode", split="val", scene_type="indoors") # auto-downloads ~2.6 GB
ds = load_dataset("kitti_eigen", split="test", root="/data/kitti") # local path
ds = load_dataset("folder", image_dir="rgb/", depth_dir="depth/") # any folder
See docs/data.md.
ONNX Export — deploy outside PyTorch
from depth_estimation import AutoDepthModel, export_onnx
model = AutoDepthModel.from_pretrained("depth-anything-v2-vitb")
export_onnx(model, "depth_anything_v2_vitb.onnx", input_size=518, verify=True)
# CLI
depth-estimate export --model depth-anything-v2-vitb --output model.onnx --verify
Verified working for depth-anything-v1/v2/v3, depth-pro, and moge, with model-specific limitations documented in the compatibility table. zoedepth and marigold-dc are not exportable because both wrap opaque external pipelines; export_onnx() fails early with a clear error. Export requires pip install "depth-estimation[export]"; using verify=True additionally requires onnxruntime or onnxruntime-gpu. See docs/export.md for the full supported-models table and GPU setup.
Pruning — reduce effective parameter count
from depth_estimation import AutoDepthModel, prune_model, compute_sparsity
model = AutoDepthModel.from_pretrained("depth-anything-v2-vitb")
prune_model(model, amount=0.3) # zero out the smallest-magnitude 30% of weights
print(compute_sparsity(model)["overall"]) # ~0.3
model.export_onnx("pruned.onnx", verify=True) # pruned models export like any other
Pure PyTorch (torch.nn.utils.prune) — no special hardware or SDK. See docs/pruning.md for prune-aware fine-tuning and what pruning does/doesn't buy you (it zeros weights, it doesn't shrink tensors).
For a complete prune → fine-tune → export → quantize workflow, see examples/optimize.py.
Quantization — float16/bfloat16/int8, plus ONNX int8/uint8
from depth_estimation import AutoDepthModel, quantize_onnx
model = AutoDepthModel.from_pretrained("depth-anything-v2-vitb")
model.quantize(dtype="float16") # in-place GPU precision cast
model.export_onnx("model.onnx")
quantize_onnx("model.onnx", "model_uint8.onnx") # default weight_type="uint8", verify=True
verify defaults to True — quantization accuracy is model-dependent, not reliably safe: testing across the 28 registered variants found uint8 produces badly wrong output for several real pretrained checkpoints, not just an edge case. See docs/quantization.md — including which formats are actually verified working (int16/uint16 are not, despite being accepted by the underlying APIs, and int8 needs onnxruntime>=1.26.0 for models with Conv2d layers).
Adding a New Model
- Create
src/depth_estimation/models/your_model/ - Add
configuration_your_model.py(inheritBaseDepthConfig) - Add
modeling_your_model.py(inheritBaseDepthModel, single file) - Add
__init__.pywithMODEL_REGISTRY.register(...)
AutoDepthModel, AutoProcessor, and pipeline() resolve the new model automatically. See docs/adding_a_model.md for a step-by-step guide.
Acknowledgments
This library builds upon the work of 12 research teams — see docs/models.md#citations for the full list.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file depth_estimation-0.1.3.tar.gz.
File metadata
- Download URL: depth_estimation-0.1.3.tar.gz
- Upload date:
- Size: 210.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8425fe4e7a152c7630781203ba11ad6f24045784cecdbdcecd2eec9486cec77f
|
|
| MD5 |
da605c62f01f92bda7a9e241ce7b3160
|
|
| BLAKE2b-256 |
3fbd244548f34dcafbf875fe3393661ecce6da45782ae222252fb8bb15f64fa1
|
Provenance
The following attestation bundles were made for depth_estimation-0.1.3.tar.gz:
Publisher:
python-publish.yml on shriarul5273/depth_estimation
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
depth_estimation-0.1.3.tar.gz -
Subject digest:
8425fe4e7a152c7630781203ba11ad6f24045784cecdbdcecd2eec9486cec77f - Sigstore transparency entry: 2549345438
- Sigstore integration time:
-
Permalink:
shriarul5273/depth_estimation@8d6e2b68a32792f5ac720d7289f125b9a130400d -
Branch / Tag:
refs/tags/v0.1.3 - Owner: https://github.com/shriarul5273
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@8d6e2b68a32792f5ac720d7289f125b9a130400d -
Trigger Event:
release
-
Statement type:
File details
Details for the file depth_estimation-0.1.3-py3-none-any.whl.
File metadata
- Download URL: depth_estimation-0.1.3-py3-none-any.whl
- Upload date:
- Size: 207.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
00ace63a6275c971db3026b413ba8e1620d3c9528b2eacf5d7214fa2cdd1bf3e
|
|
| MD5 |
141471db1a1a9b3f13d8ec48f2d09d96
|
|
| BLAKE2b-256 |
a82594cf88edee68d8702c9dbc6e540ef218762128bc094003b480b290e793dd
|
Provenance
The following attestation bundles were made for depth_estimation-0.1.3-py3-none-any.whl:
Publisher:
python-publish.yml on shriarul5273/depth_estimation
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
depth_estimation-0.1.3-py3-none-any.whl -
Subject digest:
00ace63a6275c971db3026b413ba8e1620d3c9528b2eacf5d7214fa2cdd1bf3e - Sigstore transparency entry: 2549345499
- Sigstore integration time:
-
Permalink:
shriarul5273/depth_estimation@8d6e2b68a32792f5ac720d7289f125b9a130400d -
Branch / Tag:
refs/tags/v0.1.3 - Owner: https://github.com/shriarul5273
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@8d6e2b68a32792f5ac720d7289f125b9a130400d -
Trigger Event:
release
-
Statement type: