depth-anything-v2-small-mlx

English | 日本語

Model Summary

This is an unofficial MLX conversion of depth-anything/Depth-Anything-V2-Small-hf (a monocular depth-estimation model with a DINOv2-Small backbone and a DPT-style reassemble/fusion neck). All credit for the original model goes to its authors (TikTok / The HuggingFace team).

This cannot be loaded with mlx-vlm

A DINOv2 backbone + DPT-style reassemble/fusion neck isn't in mlx-vlm's list of supported architectures, so this was reimplemented from scratch for MLX and requires the bundled depth_anything_mlx.py.

Usage

import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np

from depth_anything_mlx import DepthAnythingMLX, IMAGE_SIZE

model = DepthAnythingMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())

image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16))  # (1, 518, 518, 3), NHWC

depth = model(pixel_values)  # (1, 518, 518)

Input is a fixed 518x518, NHWC (note the channel order differs from Core ML/torch's NCHW). Weights are stored in float16.

Accuracy

Compared against the PyTorch fp32 reference for depth-anything/Depth-Anything-V2-Small-hf on a COCO validation image:

Precision Cosine similarity MAE (relative)
MLX fp32 1.0 1.1e-6
MLX fp16 (this release) 1.0000001 0.08%

Specs

Item Value
Base model depth-anything/Depth-Anything-V2-Small-hf (DINOv2-Small, 24.8M params)
Precision float16
Input 518x518, NHWC, fixed size (dynamic resolution not supported)
Framework MLX (from-scratch depth_anything_mlx.py)

Notes

  • This is a community conversion, not an official release from the Depth Anything authors.
  • Security audit uses model-audit-lite (see SECURITY.md for details).

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

depth-anything/Depth-Anything-V2-Small-hf (DINOv2-Smallバックボーン + DPT系reassemble/fusionネックによる単眼深度推定モデル)の MLX版です。元モデルの著作権はその作者(TikTok / HuggingFaceチーム)に帰属します。

mlx-vlmでは読み込めません

DINOv2バックボーン + DPT系のreassemble/fusionネックという構成はmlx-vlmの対応アーキテクチャ 一覧に含まれていないため、MLXでの実装をゼロから書き起こして変換しています。同梱の depth_anything_mlx.pyが必要です。

使い方

import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np

from depth_anything_mlx import DepthAnythingMLX, IMAGE_SIZE

model = DepthAnythingMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())

image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16))  # (1, 518, 518, 3), NHWC

depth = model(pixel_values)  # (1, 518, 518)

入力は固定サイズ518x518、NHWC(Core MLやtorch (NCHW) とはチャンネル順が異なる点に注意)。 重みはfloat16で保存しています。

精度検証

depth-anything/Depth-Anything-V2-Small-hfのPyTorch fp32リファレンスと、COCO検証画像1枚で比較:

精度 コサイン類似度 MAE(相対)
MLX fp32 1.0 1.1e-6
MLX fp16(本リリース) 1.0000001 0.08%

Specs

Item Value
ベースモデル depth-anything/Depth-Anything-V2-Small-hf(DINOv2-Small、24.8M params)
精度 float16
入力 518x518、NHWC、固定サイズ(動的解像度は未対応)
フレームワーク MLX(ゼロから実装したdepth_anything_mlx.py)

備考

  • 本変換は非公式のコミュニティ版です。
  • セキュリティー監査にはmodel-audit-liteを 使用しています(詳細はSECURITY.md)。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
24.8M params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/depth-anything-v2-small-mlx

Finetuned
(6)
this model