img-depth-anything / README.md
arraypress's picture
Depth Anything V2 Small for Core AI, nine input sizes
594e017 verified
|
Raw History Blame Contribute Delete
2.1 kB
metadata
license: apache-2.0
base_model: depth-anything/Depth-Anything-V2-Small
tags:
  - depth-estimation
  - monocular-depth
  - depth-anything
  - core-ai
  - apple-silicon
  - macos
library_name: swift-depth-estimator

img-depth-anything — Depth Anything V2 Small for Core AI (macOS 27)

Depth Anything V2 Small (Yang, Kang, Ming, Xu, Feng and Zhao, NeurIPS 2024; Apache 2.0) exported from the authors' own code and checkpoint (depth_anything_v2_vits.pth) to an Apple Core AI .aimodel for on-device monocular depth. Read by swift-depth-estimator and the img depth verb.

Files

Path What
img-depth-anything-v2-small-float32.aimodel/ Nine entry points, one per input size the authors' Resize (shorter side 518, multiples of 14, lower_bound) produces for common aspect ratios: h518w518, h518w686, h686w518, h518w784, h784w518, h518w924, h924w518, h518w826, h826w518. Each takes an ImageNet-normalised RGB image [1, 3, H, W] and returns relative inverse depth [1, H, W] (larger is nearer). Float32, 102 MB — the weights are shared.

The host reproduces infer_image: OpenCV INTER_CUBIC on the image in 0…1, ImageNet mean and standard deviation, the network, then bilinear upsampling with align_corners=True back to the picture's own size. A picture whose aspect ratio maps to a size not in the list uses the nearest.

Fidelity

Against the authors' PyTorch model on the CPU, on their demo images cropped to every exported ratio: prepared tensors above 100 dB PSNR (the cubic resampler is exact), network outputs 111–127 dB, whole images end to end 113–130 dB.

Use

hf download arraypress/img-depth-anything --local-dir models
img depth install models/img-depth-anything-v2-small-float32.aimodel
img depth photo.jpg        # photo-depth.png, nearer is brighter

Requires macOS 27 (Core AI) on Apple silicon. Converted with coreai-torch 0.4.2 / torch 2.13.