Depth Anything V2 Small for Core AI, nine input sizes
Browse files
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
img-depth-anything-v2-small-float32.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,46 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: depth-anything/Depth-Anything-V2-Small
|
| 4 |
+
tags:
|
| 5 |
+
- depth-estimation
|
| 6 |
+
- monocular-depth
|
| 7 |
+
- depth-anything
|
| 8 |
+
- core-ai
|
| 9 |
+
- apple-silicon
|
| 10 |
+
- macos
|
| 11 |
+
library_name: swift-depth-estimator
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# img-depth-anything — Depth Anything V2 Small for Core AI (macOS 27)
|
| 15 |
+
|
| 16 |
+
[Depth Anything V2](https://github.com/DepthAnything/Depth-Anything-V2) Small (Yang, Kang, Ming,
|
| 17 |
+
Xu, Feng and Zhao, NeurIPS 2024; Apache 2.0) exported from the authors' own code and checkpoint
|
| 18 |
+
(`depth_anything_v2_vits.pth`) to an Apple Core AI `.aimodel` for on-device monocular depth. Read
|
| 19 |
+
by [swift-depth-estimator](https://github.com/arraypress/swift-depth-estimator) and the
|
| 20 |
+
[`img depth`](https://github.com/arraypress/swift-img-cli) verb.
|
| 21 |
+
|
| 22 |
+
## Files
|
| 23 |
+
|
| 24 |
+
| Path | What |
|
| 25 |
+
|---|---|
|
| 26 |
+
| `img-depth-anything-v2-small-float32.aimodel/` | Nine entry points, one per input size the authors' `Resize` (shorter side 518, multiples of 14, `lower_bound`) produces for common aspect ratios: `h518w518`, `h518w686`, `h686w518`, `h518w784`, `h784w518`, `h518w924`, `h924w518`, `h518w826`, `h826w518`. Each takes an ImageNet-normalised RGB image `[1, 3, H, W]` and returns relative inverse depth `[1, H, W]` (larger is nearer). Float32, 102 MB — the weights are shared. |
|
| 27 |
+
|
| 28 |
+
The host reproduces `infer_image`: OpenCV `INTER_CUBIC` on the image in 0…1, ImageNet mean and
|
| 29 |
+
standard deviation, the network, then bilinear upsampling with `align_corners=True` back to the
|
| 30 |
+
picture's own size. A picture whose aspect ratio maps to a size not in the list uses the nearest.
|
| 31 |
+
|
| 32 |
+
## Fidelity
|
| 33 |
+
|
| 34 |
+
Against the authors' PyTorch model on the CPU, on their demo images cropped to every exported
|
| 35 |
+
ratio: prepared tensors above 100 dB PSNR (the cubic resampler is exact), network outputs
|
| 36 |
+
111–127 dB, whole images end to end 113–130 dB.
|
| 37 |
+
|
| 38 |
+
## Use
|
| 39 |
+
|
| 40 |
+
```sh
|
| 41 |
+
hf download arraypress/img-depth-anything --local-dir models
|
| 42 |
+
img depth install models/img-depth-anything-v2-small-float32.aimodel
|
| 43 |
+
img depth photo.jpg # photo-depth.png, nearer is brighter
|
| 44 |
+
```
|
| 45 |
+
|
| 46 |
+
Requires macOS 27 (Core AI) on Apple silicon. Converted with coreai-torch 0.4.2 / torch 2.13.
|
img-depth-anything-v2-small-float32.aimodel/main.hash
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
����F#E|����o�\�.�})|9Ai�r
|
img-depth-anything-v2-small-float32.aimodel/main.mlirb
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8aebc9f94623457c9b07ded1da6ffd5c1e962e12d096cb7d297c0e394169eb72
|
| 3 |
+
size 101694147
|
img-depth-anything-v2-small-float32.aimodel/metadata.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"license" : "Apache-2.0",
|
| 3 |
+
"producer" : "coreai-core 1.0.0b2",
|
| 4 |
+
"creationDate" : "20260924T152434Z",
|
| 5 |
+
"assetVersion" : "2.0",
|
| 6 |
+
"author" : "Yang, Kang, Ming, Xu, Feng, Zhao — Depth Anything V2 (NeurIPS 2024), Small; Core AI export by DepthEstimator",
|
| 7 |
+
"description" : "Depth Anything V2 Small: entry points h518w518, h518w686, h686w518, h518w784, h784w518, h518w924, h924w518, h518w826, h826w518 take an ImageNet-normalised RGB image [1,3,H,W] and return relative inverse depth [1,H,W]. Preprocess and upsample in the host as the authors' infer_image does."
|
| 8 |
+
}
|