File size: 7,167 Bytes
f56db95 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 | ---
library_name: executorch
license: mit
base_model: fabiotosi92/ZipDepth
display_name: ZipDepth Base FP16 — ExecuTorch
tags:
- depth-estimation
- monocular-depth-estimation
- zipdepth
- executorch
- xnnpack
- fp16
- arm
- edge-ai
pipeline_tag: depth-estimation
---
<!--
SPDX-FileCopyrightText: Copyright 2026 Arm Limited and/or its affiliates <open-source-office@arm.com>
SPDX-License-Identifier: Apache-2.0
-->
# ZipDepth Base FP16 — ExecuTorch
This repository packages the unfold-free ZipDepth Base mobile/NPU checkpoint as a static 384×384
ExecuTorch model. The depth network uses FP16 activations with FP32 serialized convolution weights
and external tensors.
## ✨ Key Highlights
- **2.06× faster on Android™ Vivo X300** — one-thread p50 depth-model latency is **20.431 ms**,
compared with **41.987 ms** for FP32 under the same inference-only protocol.
- **Negligible quality change** — across all 654 NYUv2 test images, δ1 changes by only
**−0.046 percentage points** and AbsRel improves by **0.011 percentage points** relative to the
FP32 depth model under identical fixed-shape preprocessing and evaluation.
## 📦 Model Details
ZipDepth predicts affine-invariant inverse depth from one RGB image. The package accepts arbitrary
image dimensions up to 4096×4096, crops the largest top-left square, resizes it to the static
384×384 depth-model input, and restores the result to the original tensor shape with zero padding.
- **Developed by:** Fabio Tosi, Luca Bartolomei, Matteo Poggi, and Stefano Mattoccia
- **Model type:** Zero-shot monocular relative-depth estimation
- **License:** MIT
- **Base model:** [ZipDepth](https://github.com/fabiotosi92/ZipDepth)
- **Source revision:** [`a302e543`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470)
- **Source checkpoint:** `zipdepth_base_npu.pth`
- **ExecuTorch revision:** `d13c78971338b35a219e0a27f7045b29c7f4e8c1`
- **Parameters:** Approximately 6.1 million after inference-time fusion
- **Artifact size:** 23.47 MiB, 0.27% smaller than the 23.53 MiB FP32 artifact
## 🚀 Get Started with the Model
### 🔓 Compute Flow — Early Access
The inference engine for this model package is available through the Compute Flow Early Access
Program. Contact ai-early-access@arm.com to request access.
## 📊 Quality evaluation
Quality was evaluated on all 654 images in the NYUv2 test split. Both models received the same
top-left square crop resized to 384×384. Their predictions were resized to the ground-truth shape,
aligned independently in inverse-depth space using a least-squares scale and shift, and evaluated
over the NYUv2 Eigen crop and 0.001–10 m depth range. This isolates the static depth models; it is
not an evaluation of the dynamic prepare/restore transforms.
| Metric | FP32 | FP16 ExecuTorch | Change |
| --- | ---: | ---: | ---: |
| AbsRel ↓ | 15.796% | **15.785%** | 0.011 percentage points lower |
| δ1 ↑ | 77.055% | 77.009% | 0.046 percentage points lower |
The FP16 candidate uses a different official ZipDepth checkpoint because the mobile/NPU checkpoint
replaces convex unfold upsampling with a deployment-friendly learned blend. The comparison
therefore measures the complete mobile graph and FP16 conversion, not FP16 rounding in isolation.
On the distributed 384×384 sample, its output has 0.99645 Pearson correlation with the FP32 output.
## 🎯 Performance evaluation
The static depth models were profiled using the distributed `samples/im0.jpg` image, which is
already 384×384. The dynamic prepare and restore models were therefore bypassed.
- **CPU:** One Arm CPU thread.
- **Runs:** One unmeasured in-process inference to initialize each runtime, followed by 5 warmup
and 30 measured fresh-process runs with 1 second between processes.
- **Timing scope:** The setup call is not included.
- **Runtimes:** LiteRT 2.1.6 FP32 baseline and ExecuTorch `d13c7897` FP16 candidate, both with the
XNNPACK CPU backend.
- **Target:** Android™ Vivo X300 smartphone (`V2502A`) on Android™ 16, USB powered at 100% battery.
| Metric | Android™ Vivo X300: FP32 | Android™ Vivo X300: FP16 | Uplift |
| --- | ---: | ---: | ---: |
| p50 inference latency | 41.987 ms | **20.431 ms** | **2.06× faster** |
| p90 inference latency | 42.612 ms | **21.342 ms** | **2.00× faster** |
| p99 inference latency | 43.236 ms | **27.181 ms** | **1.59× faster** |
## 🛠️ Technical Specifications
### Runtime Architecture
| Component role | Framework / format |
| --- | --- |
| Input square crop and resize | ExecuTorch / PTE |
| ZipDepth Base inference | ExecuTorch / PTE |
| Output resize and zero padding | ExecuTorch / PTE |
### Precision
- Internal depth-model activations use FP16.
- Public input and output tensors remain FP32 to preserve the existing brick contract.
- Dynamic image-shape transforms use FP32.
- The fixed 24×24-to-12×12 adaptive average pool is represented as an equivalent 2×2,
stride-2 average pool supported by the deployed ExecuTorch kernel set.
### Input Specification
| Input | Data type | Shape | Value range |
| --- | --- | --- | --- |
| RGB image | FP32 | `3 × H × W`, with `1 ≤ H,W ≤ 4096` | `[0, 1]` |
### Output Specification
| Output | Data type | Shape | Description |
| --- | --- | --- | --- |
| Relative depth | FP32 | `H × W` | Affine-invariant inverse depth; padded regions are zero |
### Repository Contents
- `zipdepth_base_fp16.pte` — static FP16 ZipDepth Base mobile/NPU depth model.
- `zipdepth_prepare_image.pte` — dynamic input crop-and-resize transform.
- `zipdepth_restore_depth.pte` — dynamic output resize-and-padding transform.
- `zipdepth_manifest.json` — model tensor contracts and runtime configuration.
- `samples/im0.jpg` — 384×384 upstream sample used for profiling.
- `metadata.yaml` and `benchmarks/` — model and benchmark metadata.
- `LICENSE` — upstream MIT license.
- `SHA256SUMS` — reproducibility checksums.
## ⚠️ Known Limitations
- The square-crop package pipeline does not preserve the aspect-ratio preprocessing used for
ZipDepth's published benchmark results.
- The output is affine-invariant relative depth, not metric depth.
- The model has no temporal-consistency mechanism and may flicker when applied independently to
video frames.
## 🗂️ Model and Asset Origin
- [`zipdepth_base_npu.pth`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470/checkpoints)
is the source checkpoint for `zipdepth_base_fp16.pte`.
- [`samples/im0.jpg`](https://github.com/fabiotosi92/ZipDepth/blob/a302e5437bc58f15c4efd41d3e8222bf24f7d470/assets/examples/im0.jpg)
is an upstream example crop.
## 🔐 Checksums
Verify the checked-out files with:
```sh
shasum -a 256 -c SHA256SUMS
```
## Citation
```bibtex
@inproceedings{tosi2026zipdepth,
title = {{ZipDepth}: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device},
author = {Tosi, Fabio and Bartolomei, Luca and Poggi, Matteo and Mattoccia, Stefano},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}
```
|