Depth Anything 3 Small - Static Core FP16 Compute

This repository packages the official depth-anything/DA3-SMALL checkpoint for monocular relative-depth estimation. The package combines a static LiteRT depth model operating at 504x378 pixels (width x height) with dynamic ExecuTorch input and output resizing.

โœจ Key Highlights

  • Fast single-thread execution โ€” median packaged inference latency is 157.8 ms for a 504x378 image on an Androidโ„ข Vivo X300 smartphone using one Armยฎ CPU thread.
  • Performance uplift โ€” compared with the FP32 package, median latency is 2.36ร— faster and average memory is 6.6% lower.
  • Almost identical accuracy โ€” no material accuracy regression was observed against the FP32 implementation.
  • Memory efficient โ€” peak process RSS is 772.1 MiB on an Androidโ„ข Vivo X300 smartphone.
  • Compact deployment โ€” the packaged model files total 100.4 MB.

๐Ÿ“ฆ Model Details

Model Description

Depth Anything 3 is a visual-geometry model developed by ByteDance Seed. This package exposes its single-image relative-depth capability.

  • Developed by: ByteDance Seed
  • Model type: Monocular relative-depth estimation
  • License: Apache License 2.0
  • Base model: depth-anything/DA3-SMALL
  • Package form: LiteRT and ExecuTorch model artifacts

Model Sources

๐Ÿš€ Get Started with the Model

๐Ÿ”“ Compute Flow โ€” Early Access

The inference engine for this model package is available through the Compute Flow Early Access Program.

Want to try it?

๐Ÿ“ฉ Contact us at ai-early-access@arm.com to request access.

๐Ÿ“Š Quality Evaluation

Quality was measured with the complete packaged model on the standard 654-image NYU Depth V2 labeled test split. The evaluation follows the affine-invariant relative-depth protocol used by the Depth Anything 3 technical report:

  • Raw ground-truth depths from 0.001 m to 10 m were evaluated within the Eigen crop.
  • Each prediction was aligned to ground truth with a least-squares scale and shift in disparity space, then converted to depth and clipped to the evaluation range.
  • Metrics were calculated per image and averaged across all 654 images.
  • ฮด1 is the fraction of valid pixels where predicted and ground-truth depth differ by less than a factor of 1.25; higher is better.
  • AbsRel is the mean absolute depth error divided by ground-truth depth; lower is better.

The FP32 and FP16-compute packages retain the same weights and model operations. The optimized package adds only the execution-precision request, and both variants produce the same reported quality metrics at the displayed precision.

Metric FP32 baseline FP16 compute Delta
ฮด1 accuracy 91.63% 91.63% 0.00 pp
AbsRel 8.85% 8.85% 0.00 pp
Evaluation samples 654 654 โ€”

๐ŸŽฏ Performance Evaluation

Performance was measured under the following conditions:

  • One Armยฎ CPU thread.
  • 5 warmups followed by 30 consecutive measured runs with no pause between runs.
  • The Androidโ„ข Vivo X300 smartphone screen was kept on and thermal status remained normal.
  • The model remained loaded and its state was reset between runs.
  • The packaged 504x378 sample image was used, matching the static depth-core resolution.

The following methodology and definitions were used:

  • Runtime: LiteRT and ExecuTorch on CPU, with XNNPACK and KleidiAI enabled.
  • End-to-end latency covers the complete packaged model execution, including input resizing, depth estimation, and output resizing. It excludes model setup and image file decoding.
  • Throughput is derived from median end-to-end latency.
  • Average memory is the median of the mean process RSS from three independent executions, sampled throughout setup and inference.
  • Peak memory is the median process high-water mark from the same three executions.
  • Latency and memory were collected in separate executions to prevent memory sampling from affecting latency.

Compared with the FP32 package, the FP16-compute package provides the following uplift.

Metric Androidโ„ข Vivo X300: FP32 Androidโ„ข Vivo X300: FP16 compute Uplift
Throughput 2.68 images/s 6.34 images/s 2.36ร— higher
End-to-end latency, p50 372.8 ms 157.8 ms 2.36ร— faster
End-to-end latency, p90 384.7 ms 168.2 ms 2.29ร— faster
End-to-end latency, p99 392.2 ms 172.6 ms 2.27ร— faster
Peak memory 811.5 MiB 772.1 MiB 4.9% lower
Average memory 799.0 MiB 746.4 MiB 6.6% lower

๐Ÿ› ๏ธ Technical Specifications

Objective

Estimate a dense relative inverse-depth map from a single RGB image.

Runtime Architecture

Component role Framework / format
Input image resizing ExecuTorch / PTE
Relative-depth model LiteRT / TFLite
Output map resizing ExecuTorch / PTE

Precision and Quantization

The LiteRT model retains FP32 weights and tensor interfaces. An FP16 marker requests FP16 arithmetic for supported XNNPACK-delegated operations; unsupported operations continue to use their available precision.

Input Specification

The model accepts one planar FP32 RGB image normalized to the range [0, 1]. Each image dimension must not exceed 4096 pixels, and the image must contain no more than 1,440,000 pixels.

Output Specification

The model returns a two-dimensional FP32 relative inverse-depth map at the input image resolution. Larger values indicate points nearer to the camera. The values do not represent metric distance.

Repository Contents

  • da3_small_static_core_fp16_compute.tflite โ€” static 504x378 (width x height) LiteRT depth model.
  • da3_prepare_image.pte and da3_restore_depth.pte โ€” dynamic ExecuTorch resize transforms.
  • da3_small_manifest.json โ€” model package manifest.
  • samples/ โ€” packaged input and relative-depth visualization.
  • metadata.yaml and benchmarks/ โ€” model and benchmark metadata.
  • LICENSE โ€” upstream Apache License 2.0 text.
  • SHA256SUMS โ€” model-bundle checksums for reproducibility.

โš ๏ธ Known Limitations

  • The output is relative inverse depth and requires scale-and-shift alignment for comparison with metric ground truth.
  • This package exposes monocular relative depth only; it does not expose the upstream model's multi-view geometry or camera-pose outputs.
  • The depth core operates at 504x378 pixels (width x height), so output detail is bounded by that internal resolution.

๐Ÿ—‚๏ธ Model and Asset Origin

Models

  • da3_small_static_core_fp16_compute.tflite was derived from model.safetensors in the official DA3-SMALL repository using source code from the Depth Anything 3 repository, then augmented with an unused FP16 marker constant to request XNNPACK FP16 computation.
  • da3_prepare_image.pte and da3_restore_depth.pte were generated from PyTorch image-resize transforms.
  • Depth Anything 3 is distributed under the Apache License 2.0 reproduced in LICENSE.

Sample

  • samples/000.png is an example image from the Depth Anything 3 repository.
  • samples/000_relative_depth.png was generated from samples/000.png with the packaged model and rendered as a relative-depth visualization.

๐Ÿ” Checksums

SHA256SUMS was generated by recursively hashing every regular file in the model bundle, including files in subdirectories, except the generated root SHA256SUMS and paths with a dotfile component.

From the model bundle root, verify the checked-out files with:

shasum -a 256 -c SHA256SUMS
Downloads last month
96
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Arm/depth-anything-v3-small-mix-precision

Quantized
(6)
this model

Collection including Arm/depth-anything-v3-small-mix-precision

Paper for Arm/depth-anything-v3-small-mix-precision