Instructions to use masahiroid/depth-anything-v2-small-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use masahiroid/depth-anything-v2-small-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download masahiroid/depth-anything-v2-small-mlx --local-dir depth-anything-v2-small-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
depth-anything-v2-small-mlx
Model Summary
This is an unofficial MLX conversion of depth-anything/Depth-Anything-V2-Small-hf (a monocular depth-estimation model with a DINOv2-Small backbone and a DPT-style reassemble/fusion neck). All credit for the original model goes to its authors (TikTok / The HuggingFace team).
This cannot be loaded with mlx-vlm
A DINOv2 backbone + DPT-style reassemble/fusion neck isn't in mlx-vlm's
list of supported architectures, so this was reimplemented from scratch
for MLX and requires the bundled depth_anything_mlx.py.
Usage
import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np
from depth_anything_mlx import DepthAnythingMLX, IMAGE_SIZE
model = DepthAnythingMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16)) # (1, 518, 518, 3), NHWC
depth = model(pixel_values) # (1, 518, 518)
Input is a fixed 518x518, NHWC (note the channel order differs from Core ML/torch's NCHW). Weights are stored in float16.
Accuracy
Compared against the PyTorch fp32 reference for
depth-anything/Depth-Anything-V2-Small-hf on a COCO validation image:
| Precision | Cosine similarity | MAE (relative) |
|---|---|---|
| MLX fp32 | 1.0 | 1.1e-6 |
| MLX fp16 (this release) | 1.0000001 | 0.08% |
Specs
| Item | Value |
|---|---|
| Base model | depth-anything/Depth-Anything-V2-Small-hf (DINOv2-Small, 24.8M params) |
| Precision | float16 |
| Input | 518x518, NHWC, fixed size (dynamic resolution not supported) |
| Framework | MLX (from-scratch depth_anything_mlx.py) |
Notes
- This is a community conversion, not an official release from the Depth Anything authors.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
depth-anything/Depth-Anything-V2-Small-hf (DINOv2-Smallバックボーン + DPT系reassemble/fusionネックによる単眼深度推定モデル)の MLX版です。元モデルの著作権はその作者(TikTok / HuggingFaceチーム)に帰属します。
mlx-vlmでは読み込めません
DINOv2バックボーン + DPT系のreassemble/fusionネックという構成はmlx-vlmの対応アーキテクチャ
一覧に含まれていないため、MLXでの実装をゼロから書き起こして変換しています。同梱の
depth_anything_mlx.pyが必要です。
使い方
import mlx.core as mx
from mlx.utils import tree_unflatten
from PIL import Image
import numpy as np
from depth_anything_mlx import DepthAnythingMLX, IMAGE_SIZE
model = DepthAnythingMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
image = Image.open("photo.jpg").convert("RGB").resize((IMAGE_SIZE, IMAGE_SIZE))
pixel_values = np.asarray(image, dtype=np.float32) / 255.0
pixel_values = (pixel_values - np.array([0.485, 0.456, 0.406])) / np.array([0.229, 0.224, 0.225])
pixel_values = mx.array(pixel_values[None].astype(np.float16)) # (1, 518, 518, 3), NHWC
depth = model(pixel_values) # (1, 518, 518)
入力は固定サイズ518x518、NHWC(Core MLやtorch (NCHW) とはチャンネル順が異なる点に注意)。 重みはfloat16で保存しています。
精度検証
depth-anything/Depth-Anything-V2-Small-hfのPyTorch fp32リファレンスと、COCO検証画像1枚で比較:
| 精度 | コサイン類似度 | MAE(相対) |
|---|---|---|
| MLX fp32 | 1.0 | 1.1e-6 |
| MLX fp16(本リリース) | 1.0000001 | 0.08% |
Specs
| Item | Value |
|---|---|
| ベースモデル | depth-anything/Depth-Anything-V2-Small-hf(DINOv2-Small、24.8M params) |
| 精度 | float16 |
| 入力 | 518x518、NHWC、固定サイズ(動的解像度は未対応) |
| フレームワーク | MLX(ゼロから実装したdepth_anything_mlx.py) |
備考
- 本変換は非公式のコミュニティ版です。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
Quantized
Model tree for masahiroid/depth-anything-v2-small-mlx
Base model
depth-anything/Depth-Anything-V2-Small-hf