depth-anything-v2-small-coreml

English | 日本語

Model Summary

This is an unofficial Core ML conversion of depth-anything/Depth-Anything-V2-Small-hf. All credit for the original model goes to its authors (TikTok / The HuggingFace team).

Conversion notes

Depth estimation is a single forward pass with no autoregressive decoder, so no KV-cache complexity is needed -- torch.jit.trace + coremltools converts it directly. The one adjustment: DINOv2's position-embedding interpolation (interpolate_pos_encoding) always forces a bicubic interpolate under tracing, which coremltools doesn't support. At this model's native training resolution of 518x518, that interpolation is mathematically a no-op, which was verified numerically before skipping it.

Usage (Python / coremltools)

import numpy as np
import coremltools as ct
from PIL import Image

mlmodel = ct.models.MLModel("depth-anything-v2-small_fp16.mlpackage")

image = Image.open("photo.jpg").convert("RGB").resize((518, 518))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32)  # (1, 3, 518, 518), NCHW

out = mlmodel.predict({"pixel_values": pixel_values})
depth = out["predicted_depth"][0]  # (518, 518)

Input is a fixed 518x518, NCHW, normalized with mean [0.485, 0.456, 0.406] / std [0.229, 0.224, 0.225].

Accuracy

Compared against the PyTorch fp32 reference on a COCO validation image:

Metric Result
Cosine similarity 0.999998
Mean absolute error (relative) 0.14%

Specs

Item Value
Base model depth-anything/Depth-Anything-V2-Small-hf (DINOv2-Small, 24.8M params)
Precision float16
Input 518x518, NCHW, fixed size (dynamic resolution not supported)
Framework Core ML (mlprogram, minimum_deployment_target=macOS14)

Notes

  • This is a community conversion, not an official release from the Depth Anything authors.
  • Security audit uses model-audit-lite (see SECURITY.md for details).

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

depth-anything/Depth-Anything-V2-Small-hf のCore ML版です。元モデルの著作権はその作者(TikTok / HuggingFaceチーム)に帰属します。

変換について

深度推定モデルは自己回帰デコーダーを持たない単純な1回のフォワードパスのため、KVキャッシュなどの 複雑さは不要です。torch.jit.trace + coremltoolsでそのまま変換できます。唯一の調整点は、 DINOv2の位置埋め込み補間(interpolate_pos_encoding)がトレース時に常にbicubic補間を強制する 実装になっており、これがcoremltoolsでサポートされていないため、本モデルの学習解像度である 518x518固定入力では補間が数学的に恒等写像になることを確認した上で、この補間をスキップしています (実際に同一であることをテストで確認済み)。

使い方(Python / coremltools)

import numpy as np
import coremltools as ct
from PIL import Image

mlmodel = ct.models.MLModel("depth-anything-v2-small_fp16.mlpackage")

image = Image.open("photo.jpg").convert("RGB").resize((518, 518))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = pixel_values.transpose(2, 0, 1)[None].astype(np.float32)  # (1, 3, 518, 518), NCHW

out = mlmodel.predict({"pixel_values": pixel_values})
depth = out["predicted_depth"][0]  # (518, 518)

入力は固定サイズ518x518、NCHW([0.485,0.456,0.406]/[0.229,0.224,0.225]で正規化)。

精度検証

PyTorch fp32リファレンスと、COCO検証画像1枚で比較:

項目 結果
コサイン類似度 0.999998
平均絶対誤差(相対) 0.14%

Specs

Item Value
ベースモデル depth-anything/Depth-Anything-V2-Small-hf(DINOv2-Small、24.8M params)
精度 float16
入力 518x518、NCHW、固定サイズ(動的解像度は未対応)
フレームワーク Core ML(mlprogram、minimum_deployment_target=macOS14)

備考

  • 本変換は非公式のコミュニティ版です。
  • セキュリティー監査にはmodel-audit-liteを 使用しています(詳細はSECURITY.md)。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/depth-anything-v2-small-coreml

Quantized
(9)
this model