Depth Anything V2
Paper • 2406.09414 • Published • 105
Core ML conversions of Depth Anything V2 Small at smaller fixed input sizes than Apple's 518×392 conversion (apple/coreml-depth-anything-v2-small), for real-time depth on the Neural Engine. Used by Arkestra for live depth.
| Package | Input | Neural Engine latency (M-series) |
|---|---|---|
DepthAnythingV2SmallF16_364x210.mlpackage |
364×210 (16:9) | ~6 ms |
DepthAnythingV2SmallF16_448x252.mlpackage |
448×252 (16:9) | ~12 ms |
Apple's DepthAnythingV2SmallF16.mlpackage (for comparison) |
518×392 | ~23 ms |
image: RGB image, exactly the size above. ImageNet normalisation is built in.depth: Float16 grayscale image of the same size, relative inverse depth (larger is nearer).
Unlike Apple's conversion it is not min/max-normalised per frame.FP16, ML Program, macOS 14+ / iOS 17+. DINOv2's bicubic positional-embedding resize is computed once at
export time for the fixed input size (no accuracy loss). Export script:
scripts/depth-export/export_depth_model.py in the Arkestra repository.
Apache-2.0, same as the original Depth Anything V2 Small weights.
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao. Depth Anything V2, 2024. arXiv:2406.09414
Base model
depth-anything/Depth-Anything-V2-Small-hf