Instructions to use Arm/depth-anything-v3-small-mix-precision with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use Arm/depth-anything-v3-small-mix-precision with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Depth Anything 3 Small - Static Core FP16 Compute
This repository packages the official
depth-anything/DA3-SMALL checkpoint for
monocular relative-depth estimation. The package combines a static LiteRT depth model operating at
504x378 pixels (width x height) with dynamic ExecuTorch input and output resizing.
โจ Key Highlights
- Fast single-thread execution โ median packaged inference latency is 157.8 ms for a 504x378 image on an Androidโข Vivo X300 smartphone using one Armยฎ CPU thread.
- Performance uplift โ compared with the FP32 package, median latency is 2.36ร faster and average memory is 6.6% lower.
- Almost identical accuracy โ no material accuracy regression was observed against the FP32 implementation.
- Memory efficient โ peak process RSS is 772.1 MiB on an Androidโข Vivo X300 smartphone.
- Compact deployment โ the packaged model files total 100.4 MB.
๐ฆ Model Details
Model Description
Depth Anything 3 is a visual-geometry model developed by ByteDance Seed. This package exposes its single-image relative-depth capability.
- Developed by: ByteDance Seed
- Model type: Monocular relative-depth estimation
- License: Apache License 2.0
- Base model:
depth-anything/DA3-SMALL - Package form: LiteRT and ExecuTorch model artifacts
Model Sources
- Base model: https://huggingface.co/depth-anything/DA3-SMALL
- Upstream repository: https://github.com/ByteDance-Seed/Depth-Anything-3
- Technical report: https://arxiv.org/abs/2511.10647
๐ Get Started with the Model
๐ Compute Flow โ Early Access
The inference engine for this model package is available through the Compute Flow Early Access Program.
Want to try it?
๐ฉ Contact us at ai-early-access@arm.com to request access.
๐ Quality Evaluation
Quality was measured with the complete packaged model on the standard 654-image NYU Depth V2 labeled test split. The evaluation follows the affine-invariant relative-depth protocol used by the Depth Anything 3 technical report:
- Raw ground-truth depths from 0.001 m to 10 m were evaluated within the Eigen crop.
- Each prediction was aligned to ground truth with a least-squares scale and shift in disparity space, then converted to depth and clipped to the evaluation range.
- Metrics were calculated per image and averaged across all 654 images.
- ฮด1 is the fraction of valid pixels where predicted and ground-truth depth differ by less than a factor of 1.25; higher is better.
- AbsRel is the mean absolute depth error divided by ground-truth depth; lower is better.
The FP32 and FP16-compute packages retain the same weights and model operations. The optimized package adds only the execution-precision request, and both variants produce the same reported quality metrics at the displayed precision.
| Metric | FP32 baseline | FP16 compute | Delta |
|---|---|---|---|
| ฮด1 accuracy | 91.63% | 91.63% | 0.00 pp |
| AbsRel | 8.85% | 8.85% | 0.00 pp |
| Evaluation samples | 654 | 654 | โ |
๐ฏ Performance Evaluation
Performance was measured under the following conditions:
- One Armยฎ CPU thread.
- 5 warmups followed by 30 consecutive measured runs with no pause between runs.
- The Androidโข Vivo X300 smartphone screen was kept on and thermal status remained normal.
- The model remained loaded and its state was reset between runs.
- The packaged 504x378 sample image was used, matching the static depth-core resolution.
The following methodology and definitions were used:
- Runtime: LiteRT and ExecuTorch on CPU, with XNNPACK and KleidiAI enabled.
- End-to-end latency covers the complete packaged model execution, including input resizing, depth estimation, and output resizing. It excludes model setup and image file decoding.
- Throughput is derived from median end-to-end latency.
- Average memory is the median of the mean process RSS from three independent executions, sampled throughout setup and inference.
- Peak memory is the median process high-water mark from the same three executions.
- Latency and memory were collected in separate executions to prevent memory sampling from affecting latency.
Compared with the FP32 package, the FP16-compute package provides the following uplift.
| Metric | Androidโข Vivo X300: FP32 | Androidโข Vivo X300: FP16 compute | Uplift |
|---|---|---|---|
| Throughput | 2.68 images/s | 6.34 images/s | 2.36ร higher |
| End-to-end latency, p50 | 372.8 ms | 157.8 ms | 2.36ร faster |
| End-to-end latency, p90 | 384.7 ms | 168.2 ms | 2.29ร faster |
| End-to-end latency, p99 | 392.2 ms | 172.6 ms | 2.27ร faster |
| Peak memory | 811.5 MiB | 772.1 MiB | 4.9% lower |
| Average memory | 799.0 MiB | 746.4 MiB | 6.6% lower |
๐ ๏ธ Technical Specifications
Objective
Estimate a dense relative inverse-depth map from a single RGB image.
Runtime Architecture
| Component role | Framework / format |
|---|---|
| Input image resizing | ExecuTorch / PTE |
| Relative-depth model | LiteRT / TFLite |
| Output map resizing | ExecuTorch / PTE |
Precision and Quantization
The LiteRT model retains FP32 weights and tensor interfaces. An FP16 marker requests FP16 arithmetic for supported XNNPACK-delegated operations; unsupported operations continue to use their available precision.
Input Specification
The model accepts one planar FP32 RGB image normalized to the range [0, 1]. Each image dimension must not exceed 4096 pixels, and the image must contain no more than 1,440,000 pixels.
Output Specification
The model returns a two-dimensional FP32 relative inverse-depth map at the input image resolution. Larger values indicate points nearer to the camera. The values do not represent metric distance.
Repository Contents
da3_small_static_core_fp16_compute.tfliteโ static 504x378 (width x height) LiteRT depth model.da3_prepare_image.pteandda3_restore_depth.pteโ dynamic ExecuTorch resize transforms.da3_small_manifest.jsonโ model package manifest.samples/โ packaged input and relative-depth visualization.metadata.yamlandbenchmarks/โ model and benchmark metadata.LICENSEโ upstream Apache License 2.0 text.SHA256SUMSโ model-bundle checksums for reproducibility.
โ ๏ธ Known Limitations
- The output is relative inverse depth and requires scale-and-shift alignment for comparison with metric ground truth.
- This package exposes monocular relative depth only; it does not expose the upstream model's multi-view geometry or camera-pose outputs.
- The depth core operates at 504x378 pixels (width x height), so output detail is bounded by that internal resolution.
๐๏ธ Model and Asset Origin
Models
da3_small_static_core_fp16_compute.tflitewas derived frommodel.safetensorsin the official DA3-SMALL repository using source code from the Depth Anything 3 repository, then augmented with an unused FP16 marker constant to request XNNPACK FP16 computation.da3_prepare_image.pteandda3_restore_depth.ptewere generated from PyTorch image-resize transforms.- Depth Anything 3 is distributed under the Apache License 2.0 reproduced in
LICENSE.
Sample
samples/000.pngis an example image from the Depth Anything 3 repository.samples/000_relative_depth.pngwas generated fromsamples/000.pngwith the packaged model and rendered as a relative-depth visualization.
๐ Checksums
SHA256SUMS was generated by recursively hashing every regular file in the model bundle, including
files in subdirectories, except the generated root SHA256SUMS and paths with a dotfile component.
From the model bundle root, verify the checked-out files with:
shasum -a 256 -c SHA256SUMS
- Downloads last month
- 96
Model tree for Arm/depth-anything-v3-small-mix-precision
Base model
depth-anything/DA3-SMALL