Depth Anything V2: Robust Monocular Depth Estimation

These details have not been verified by PyPI

Project links

Project description

Depth Anything V2

Lihe Yang¹ · Bingyi Kang^2† · Zilong Huang²
Zhen Zhao · Xiaogang Xu · Jiashi Feng² · Hengshuang Zhao^1*

¹HKU ²TikTok
†project lead *corresponding author

This work presents Depth Anything V2. It significantly outperforms V1 in fine-grained details and robustness. Compared with SD-based models, it enjoys faster inference speed, fewer parameters, and higher depth accuracy.

teaser

News

2025-01-22: Video Depth Anything has been released. It generates consistent depth maps for super-long videos (e.g., over 5 minutes).
2024-12-22: Prompt Depth Anything has been released. It supports 4K resolution metric depth estimation when low-res LiDAR is used to prompt the DA models.
2024-07-06: Depth Anything V2 is supported in Transformers. See the instructions for convenient usage.
2024-06-25: Depth Anything is integrated into Apple Core ML Models. See the instructions (V1, V2) for usage.
2024-06-22: We release smaller metric depth models based on Depth-Anything-V2-Small and Base.
2024-06-20: Our repository and project page are flagged by GitHub and removed from the public for 6 days. Sorry for the inconvenience.
2024-06-14: Paper, project page, code, models, demo, and benchmark are all released.

Pre-trained Models

We provide four models of varying scales for robust relative depth estimation:

Model	Params	Checkpoint
Depth-Anything-V2-Small	24.8M	Download
Depth-Anything-V2-Base	97.5M	Download
Depth-Anything-V2-Large	335.3M	Download
Depth-Anything-V2-Giant	1.3B	Coming soon

Installation

From PyPI

You can install Depth Anything V2 directly from PyPI:

pip install jkp-depth-anything-v2
# Or with uv
uv pip install jkp-depth-anything-v2

From Source

git clone https://github.com/DepthAnything/Depth-Anything-V2
cd Depth-Anything-V2
pip install -e .
# Or with uv
uv pip install -e .

Usage

Prepraration

git clone https://github.com/DepthAnything/Depth-Anything-V2
cd Depth-Anything-V2
pip install -r requirements.txt

Download the checkpoints listed here and put them under the checkpoints directory.

Use our models

import cv2
import torch

from depth_anything_v2.dpt import DepthAnythingV2

DEVICE = 'cuda' if torch.cuda.is_available() else 'mps' if torch.backends.mps.is_available() else 'cpu'

model_configs = {
    'vits': {'encoder': 'vits', 'features': 64, 'out_channels': [48, 96, 192, 384]},
    'vitb': {'encoder': 'vitb', 'features': 128, 'out_channels': [96, 192, 384, 768]},
    'vitl': {'encoder': 'vitl', 'features': 256, 'out_channels': [256, 512, 1024, 1024]},
    'vitg': {'encoder': 'vitg', 'features': 384, 'out_channels': [1536, 1536, 1536, 1536]}
}

encoder = 'vitl' # or 'vits', 'vitb', 'vitg'

model = DepthAnythingV2(**model_configs[encoder])
model.load_state_dict(torch.load(f'checkpoints/depth_anything_v2_{encoder}.pth', map_location='cpu'))
model = model.to(DEVICE).eval()

raw_img = cv2.imread('your/image/path')
depth = model.infer_image(raw_img) # HxW raw depth map in numpy

If you do not want to clone this repository, you can also load our models through Transformers. Below is a simple code snippet. Please refer to the official page for more details.

Note 1: Make sure you can connect to Hugging Face and have installed the latest Transformers.
Note 2: Due to the upsampling difference between OpenCV (we used) and Pillow (HF used), predictions may differ slightly. So you are more recommended to use our models through the way introduced above.

from transformers import pipeline
from PIL import Image

pipe = pipeline(task="depth-estimation", model="depth-anything/Depth-Anything-V2-Small-hf")
image = Image.open('your/image/path')
depth = pipe(image)["depth"]

Running script on images

python run.py \
  --encoder <vits | vitb | vitl | vitg> \
  --img-path <path> --outdir <outdir> \
  [--input-size <size>] [--pred-only] [--grayscale]

Options:

--img-path: You can either 1) point it to an image directory storing all interested images, 2) point it to a single image, or 3) point it to a text file storing all image paths.
--input-size (optional): By default, we use input size 518 for model inference. You can increase the size for even more fine-grained results.
--pred-only (optional): Only save the predicted depth map, without raw image.
--grayscale (optional): Save the grayscale depth map, without applying color palette.

For example:

python run.py --encoder vitl --img-path assets/examples --outdir depth_vis

Running script on videos

python run_video.py \
  --encoder <vits | vitb | vitl | vitg> \
  --video-path assets/examples_video --outdir video_depth_vis \
  [--input-size <size>] [--pred-only] [--grayscale]

Our larger model has better temporal consistency on videos.

Gradio demo

To use our gradio demo locally:

python app.py

You can also try our online demo.

Note: Compared to V1, we have made a minor modification to the DINOv2-DPT architecture (originating from this issue). In V1, we unintentionally used features from the last four layers of DINOv2 for decoding. In V2, we use intermediate features instead. Although this modification did not improve details or accuracy, we decided to follow this common practice.

Fine-tuned to Metric Depth Estimation

Please refer to metric depth estimation.

DA-2K Evaluation Benchmark

Please refer to DA-2K benchmark.

Community Support

We sincerely appreciate all the community support for our Depth Anything series. Thank you a lot!

Apple Core ML:
Transformers:
- https://huggingface.co/docs/transformers/main/en/model_doc/depth_anything_v2
- https://huggingface.co/docs/transformers/main/en/model_doc/depth_anything
TensorRT:
- https://github.com/spacewalk01/depth-anything-tensorrt
- https://github.com/zhujiajian98/Depth-Anythingv2-TensorRT-python
ONNX: https://github.com/fabio-sim/Depth-Anything-ONNX
ComfyUI: https://github.com/kijai/ComfyUI-DepthAnythingV2
Transformers.js (real-time depth in web): https://huggingface.co/spaces/Xenova/webgpu-realtime-depth-estimation
Android:
- https://github.com/shubham0204/Depth-Anything-Android
- https://github.com/FeiGeChuanShu/ncnn-android-depth_anything

Acknowledgement

We are sincerely grateful to the awesome Hugging Face team (@Pedro Cuenca, @Niels Rogge, @Merve Noyan, @Amy Roberts, et al.) for their huge efforts in supporting our models in Transformers and Apple Core ML.

We also thank the DINOv2 team for contributing such impressive models to our community.

LICENSE

Depth-Anything-V2-Small model is under the Apache-2.0 license. Depth-Anything-V2-Base/Large/Giant models are under the CC-BY-NC-4.0 license.

Citation

If you find this project useful, please consider citing:

@article{depth_anything_v2,
  title={Depth Anything V2},
  author={Yang, Lihe and Kang, Bingyi and Huang, Zilong and Zhao, Zhen and Xu, Xiaogang and Feng, Jiashi and Zhao, Hengshuang},
  journal={arXiv:2406.09414},
  year={2024}
}

@inproceedings{depth_anything_v1,
  title={Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data}, 
  author={Yang, Lihe and Kang, Bingyi and Huang, Zilong and Xu, Xiaogang and Feng, Jiashi and Zhao, Hengshuang},
  booktitle={CVPR},
  year={2024}
}

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

0.1.1

Mar 23, 2025

This version

0.1.0

Mar 23, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

jkp_depth_anything_v2-0.1.0.tar.gz (26.1 kB view details)

Uploaded Mar 23, 2025 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

jkp_depth_anything_v2-0.1.0-py3-none-any.whl (25.0 kB view details)

Uploaded Mar 23, 2025 Python 3

File details

Details for the file jkp_depth_anything_v2-0.1.0.tar.gz.

File metadata

Download URL: jkp_depth_anything_v2-0.1.0.tar.gz
Upload date: Mar 23, 2025
Size: 26.1 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: uv/0.6.3

File hashes

Hashes for jkp_depth_anything_v2-0.1.0.tar.gz
Algorithm	Hash digest
SHA256	`0e0fdff57cde04bd2c3db40e6c61d608c96eb1b50d68ac1d6e29b8a5dc3f78be`
MD5	`ac936bef7d3870124b2080fb8ca5ac6d`
BLAKE2b-256	`a064c30f61e26fc66229e7000b3ead41d5ad2797cb9b1bc8c64a42461de8bce7`

See more details on using hashes here.

File details

Details for the file jkp_depth_anything_v2-0.1.0-py3-none-any.whl.

File metadata

Download URL: jkp_depth_anything_v2-0.1.0-py3-none-any.whl
Upload date: Mar 23, 2025
Size: 25.0 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: uv/0.6.3

File hashes

Hashes for jkp_depth_anything_v2-0.1.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`fff8af6d0a0457e1194e41531c992c280ed5e534825daabf43f6231264aa961f`
MD5	`3c4e3e675359b2e43f6462f75115af4f`
BLAKE2b-256	`0e432c664f38690846d50898f5c2859baead2a249e45cdf090377193f643b434`

See more details on using hashes here.

jkp-depth-anything-v2 0.1.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Depth Anything V2

News

Pre-trained Models

Installation

From PyPI

From Source

Usage

Prepraration

Use our models

Running script on images

Running script on videos

Gradio demo

Fine-tuned to Metric Depth Estimation

DA-2K Evaluation Benchmark

Community Support

Acknowledgement

LICENSE

Citation

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes