File size: 7,167 Bytes
f56db95
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
---
library_name: executorch
license: mit
base_model: fabiotosi92/ZipDepth
display_name: ZipDepth Base FP16 — ExecuTorch
tags:
  - depth-estimation
  - monocular-depth-estimation
  - zipdepth
  - executorch
  - xnnpack
  - fp16
  - arm
  - edge-ai
pipeline_tag: depth-estimation
---

<!--
    SPDX-FileCopyrightText: Copyright 2026 Arm Limited and/or its affiliates <open-source-office@arm.com>

    SPDX-License-Identifier: Apache-2.0
-->

# ZipDepth Base FP16 — ExecuTorch

This repository packages the unfold-free ZipDepth Base mobile/NPU checkpoint as a static 384×384
ExecuTorch model. The depth network uses FP16 activations with FP32 serialized convolution weights
and external tensors.

## ✨ Key Highlights

- **2.06× faster on Android™ Vivo X300** — one-thread p50 depth-model latency is **20.431 ms**,
  compared with **41.987 ms** for FP32 under the same inference-only protocol.
- **Negligible quality change** — across all 654 NYUv2 test images, δ1 changes by only
  **−0.046 percentage points** and AbsRel improves by **0.011 percentage points** relative to the
  FP32 depth model under identical fixed-shape preprocessing and evaluation.

## 📦 Model Details

ZipDepth predicts affine-invariant inverse depth from one RGB image. The package accepts arbitrary
image dimensions up to 4096×4096, crops the largest top-left square, resizes it to the static
384×384 depth-model input, and restores the result to the original tensor shape with zero padding.

- **Developed by:** Fabio Tosi, Luca Bartolomei, Matteo Poggi, and Stefano Mattoccia
- **Model type:** Zero-shot monocular relative-depth estimation
- **License:** MIT
- **Base model:** [ZipDepth](https://github.com/fabiotosi92/ZipDepth)
- **Source revision:** [`a302e543`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470)
- **Source checkpoint:** `zipdepth_base_npu.pth`
- **ExecuTorch revision:** `d13c78971338b35a219e0a27f7045b29c7f4e8c1`
- **Parameters:** Approximately 6.1 million after inference-time fusion
- **Artifact size:** 23.47 MiB, 0.27% smaller than the 23.53 MiB FP32 artifact

## 🚀 Get Started with the Model

### 🔓 Compute Flow — Early Access

The inference engine for this model package is available through the Compute Flow Early Access
Program. Contact ai-early-access@arm.com to request access.

## 📊 Quality evaluation

Quality was evaluated on all 654 images in the NYUv2 test split. Both models received the same
top-left square crop resized to 384×384. Their predictions were resized to the ground-truth shape,
aligned independently in inverse-depth space using a least-squares scale and shift, and evaluated
over the NYUv2 Eigen crop and 0.001–10 m depth range. This isolates the static depth models; it is
not an evaluation of the dynamic prepare/restore transforms.

| Metric | FP32 | FP16 ExecuTorch | Change |
| --- | ---: | ---: | ---: |
| AbsRel ↓ | 15.796% | **15.785%** | 0.011 percentage points lower |
| δ1 ↑ | 77.055% | 77.009% | 0.046 percentage points lower |

The FP16 candidate uses a different official ZipDepth checkpoint because the mobile/NPU checkpoint
replaces convex unfold upsampling with a deployment-friendly learned blend. The comparison
therefore measures the complete mobile graph and FP16 conversion, not FP16 rounding in isolation.
On the distributed 384×384 sample, its output has 0.99645 Pearson correlation with the FP32 output.

## 🎯 Performance evaluation

The static depth models were profiled using the distributed `samples/im0.jpg` image, which is
already 384×384. The dynamic prepare and restore models were therefore bypassed.

- **CPU:** One Arm CPU thread.
- **Runs:** One unmeasured in-process inference to initialize each runtime, followed by 5 warmup
  and 30 measured fresh-process runs with 1 second between processes.
- **Timing scope:** The setup call is not included.
- **Runtimes:** LiteRT 2.1.6 FP32 baseline and ExecuTorch `d13c7897` FP16 candidate, both with the
  XNNPACK CPU backend.
- **Target:** Android™ Vivo X300 smartphone (`V2502A`) on Android™ 16, USB powered at 100% battery.

| Metric | Android™ Vivo X300: FP32 | Android™ Vivo X300: FP16 | Uplift |
| --- | ---: | ---: | ---: |
| p50 inference latency | 41.987 ms | **20.431 ms** | **2.06× faster** |
| p90 inference latency | 42.612 ms | **21.342 ms** | **2.00× faster** |
| p99 inference latency | 43.236 ms | **27.181 ms** | **1.59× faster** |

## 🛠️ Technical Specifications

### Runtime Architecture

| Component role | Framework / format |
| --- | --- |
| Input square crop and resize | ExecuTorch / PTE |
| ZipDepth Base inference | ExecuTorch / PTE |
| Output resize and zero padding | ExecuTorch / PTE |

### Precision

- Internal depth-model activations use FP16.
- Public input and output tensors remain FP32 to preserve the existing brick contract.
- Dynamic image-shape transforms use FP32.
- The fixed 24×24-to-12×12 adaptive average pool is represented as an equivalent 2×2,
  stride-2 average pool supported by the deployed ExecuTorch kernel set.

### Input Specification

| Input | Data type | Shape | Value range |
| --- | --- | --- | --- |
| RGB image | FP32 | `3 × H × W`, with `1 ≤ H,W ≤ 4096` | `[0, 1]` |

### Output Specification

| Output | Data type | Shape | Description |
| --- | --- | --- | --- |
| Relative depth | FP32 | `H × W` | Affine-invariant inverse depth; padded regions are zero |

### Repository Contents

- `zipdepth_base_fp16.pte` — static FP16 ZipDepth Base mobile/NPU depth model.
- `zipdepth_prepare_image.pte` — dynamic input crop-and-resize transform.
- `zipdepth_restore_depth.pte` — dynamic output resize-and-padding transform.
- `zipdepth_manifest.json` — model tensor contracts and runtime configuration.
- `samples/im0.jpg` — 384×384 upstream sample used for profiling.
- `metadata.yaml` and `benchmarks/` — model and benchmark metadata.
- `LICENSE` — upstream MIT license.
- `SHA256SUMS` — reproducibility checksums.

## ⚠️ Known Limitations

- The square-crop package pipeline does not preserve the aspect-ratio preprocessing used for
  ZipDepth's published benchmark results.
- The output is affine-invariant relative depth, not metric depth.
- The model has no temporal-consistency mechanism and may flicker when applied independently to
  video frames.

## 🗂️ Model and Asset Origin

- [`zipdepth_base_npu.pth`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470/checkpoints)
  is the source checkpoint for `zipdepth_base_fp16.pte`.
- [`samples/im0.jpg`](https://github.com/fabiotosi92/ZipDepth/blob/a302e5437bc58f15c4efd41d3e8222bf24f7d470/assets/examples/im0.jpg)
  is an upstream example crop.

## 🔐 Checksums

Verify the checked-out files with:

```sh
shasum -a 256 -c SHA256SUMS
```

## Citation

```bibtex
@inproceedings{tosi2026zipdepth,
  title     = {{ZipDepth}: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device},
  author    = {Tosi, Fabio and Bartolomei, Luca and Poggi, Matteo and Mattoccia, Stefano},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026}
}
```