depth-anything-v2-small-coreml
Model Summary
This is an unofficial Core ML conversion of depth-anything/Depth-Anything-V2-Small-hf. All credit for the original model goes to its authors (TikTok / The HuggingFace team).
Conversion notes
Depth estimation is a single forward pass with no autoregressive decoder, so
no KV-cache complexity is needed -- torch.jit.trace + coremltools converts
it directly. The one adjustment: DINOv2's position-embedding interpolation
(interpolate_pos_encoding) always forces a bicubic interpolate under
tracing, which coremltools doesn't support. At this model's native training
resolution of 518x518, that interpolation is mathematically a no-op, which
was verified numerically before skipping it.
Usage (Python / coremltools)
import numpy as np
import coremltools as ct
from PIL import Image
mlmodel = ct.models.MLModel("depth-anything-v2-small_fp16.mlpackage")
image = Image.open("photo.jpg").convert("RGB").resize((518, 518))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32) # (1, 3, 518, 518), NCHW
out = mlmodel.predict({"pixel_values": pixel_values})
depth = out["predicted_depth"][0] # (518, 518)
Input is a fixed 518x518, NCHW, normalized with mean
[0.485, 0.456, 0.406] / std [0.229, 0.224, 0.225].
Accuracy
Compared against the PyTorch fp32 reference on a COCO validation image:
| Metric | Result |
|---|---|
| Cosine similarity | 0.999998 |
| Mean absolute error (relative) | 0.14% |
Specs
| Item | Value |
|---|---|
| Base model | depth-anything/Depth-Anything-V2-Small-hf (DINOv2-Small, 24.8M params) |
| Precision | float16 |
| Input | 518x518, NCHW, fixed size (dynamic resolution not supported) |
| Framework | Core ML (mlprogram, minimum_deployment_target=macOS14) |
Notes
- This is a community conversion, not an official release from the Depth Anything authors.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
depth-anything/Depth-Anything-V2-Small-hf のCore ML版です。元モデルの著作権はその作者(TikTok / HuggingFaceチーム)に帰属します。
変換について
深度推定モデルは自己回帰デコーダーを持たない単純な1回のフォワードパスのため、KVキャッシュなどの
複雑さは不要です。torch.jit.trace + coremltoolsでそのまま変換できます。唯一の調整点は、
DINOv2の位置埋め込み補間(interpolate_pos_encoding)がトレース時に常にbicubic補間を強制する
実装になっており、これがcoremltoolsでサポートされていないため、本モデルの学習解像度である
518x518固定入力では補間が数学的に恒等写像になることを確認した上で、この補間をスキップしています
(実際に同一であることをテストで確認済み)。
使い方(Python / coremltools)
import numpy as np
import coremltools as ct
from PIL import Image
mlmodel = ct.models.MLModel("depth-anything-v2-small_fp16.mlpackage")
image = Image.open("photo.jpg").convert("RGB").resize((518, 518))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32) # (1, 3, 518, 518), NCHW
out = mlmodel.predict({"pixel_values": pixel_values})
depth = out["predicted_depth"][0] # (518, 518)
入力は固定サイズ518x518、NCHW([0.485,0.456,0.406]/[0.229,0.224,0.225]で正規化)。
精度検証
PyTorch fp32リファレンスと、COCO検証画像1枚で比較:
| 項目 | 結果 |
|---|---|
| コサイン類似度 | 0.999998 |
| 平均絶対誤差(相対) | 0.14% |
Specs
| Item | Value |
|---|---|
| ベースモデル | depth-anything/Depth-Anything-V2-Small-hf(DINOv2-Small、24.8M params) |
| 精度 | float16 |
| 入力 | 518x518、NCHW、固定サイズ(動的解像度は未対応) |
| フレームワーク | Core ML(mlprogram、minimum_deployment_target=macOS14) |
備考
- 本変換は非公式のコミュニティ版です。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
- Downloads last month
- 9
Model tree for masahiroid/depth-anything-v2-small-coreml
Base model
depth-anything/Depth-Anything-V2-Small-hf