Depth Anything V2 Small — Core ML, small fixed inputs

Core ML conversions of Depth Anything V2 Small at smaller fixed input sizes than Apple's 518×392 conversion (apple/coreml-depth-anything-v2-small), for real-time depth on the Neural Engine. Used by Arkestra for live depth.

Package Input Neural Engine latency (M-series)
DepthAnythingV2SmallF16_364x210.mlpackage 364×210 (16:9) ~6 ms
DepthAnythingV2SmallF16_448x252.mlpackage 448×252 (16:9) ~12 ms
Apple's DepthAnythingV2SmallF16.mlpackage (for comparison) 518×392 ~23 ms

Interface

  • Input image: RGB image, exactly the size above. ImageNet normalisation is built in.
  • Output depth: Float16 grayscale image of the same size, relative inverse depth (larger is nearer). Unlike Apple's conversion it is not min/max-normalised per frame.

FP16, ML Program, macOS 14+ / iOS 17+. DINOv2's bicubic positional-embedding resize is computed once at export time for the fixed input size (no accuracy loss). Export script: scripts/depth-export/export_depth_model.py in the Arkestra repository.

License and credit

Apache-2.0, same as the original Depth Anything V2 Small weights.

Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao. Depth Anything V2, 2024. arXiv:2406.09414

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rhythmic-visions/depth-anything-v2-small-coreml

Quantized
(10)
this model

Paper for rhythmic-visions/depth-anything-v2-small-coreml