gimmy87's picture
Sync model repo (text/metadata)
f56db95 verified
|
Raw History Blame Contribute Delete
7.17 kB
---
library_name: executorch
license: mit
base_model: fabiotosi92/ZipDepth
display_name: ZipDepth Base FP16 — ExecuTorch
tags:
- depth-estimation
- monocular-depth-estimation
- zipdepth
- executorch
- xnnpack
- fp16
- arm
- edge-ai
pipeline_tag: depth-estimation
---
<!--
SPDX-FileCopyrightText: Copyright 2026 Arm Limited and/or its affiliates <open-source-office@arm.com>
SPDX-License-Identifier: Apache-2.0
-->
# ZipDepth Base FP16 — ExecuTorch
This repository packages the unfold-free ZipDepth Base mobile/NPU checkpoint as a static 384×384
ExecuTorch model. The depth network uses FP16 activations with FP32 serialized convolution weights
and external tensors.
## ✨ Key Highlights
- **2.06× faster on Android™ Vivo X300** — one-thread p50 depth-model latency is **20.431 ms**,
compared with **41.987 ms** for FP32 under the same inference-only protocol.
- **Negligible quality change** — across all 654 NYUv2 test images, δ1 changes by only
**−0.046 percentage points** and AbsRel improves by **0.011 percentage points** relative to the
FP32 depth model under identical fixed-shape preprocessing and evaluation.
## 📦 Model Details
ZipDepth predicts affine-invariant inverse depth from one RGB image. The package accepts arbitrary
image dimensions up to 4096×4096, crops the largest top-left square, resizes it to the static
384×384 depth-model input, and restores the result to the original tensor shape with zero padding.
- **Developed by:** Fabio Tosi, Luca Bartolomei, Matteo Poggi, and Stefano Mattoccia
- **Model type:** Zero-shot monocular relative-depth estimation
- **License:** MIT
- **Base model:** [ZipDepth](https://github.com/fabiotosi92/ZipDepth)
- **Source revision:** [`a302e543`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470)
- **Source checkpoint:** `zipdepth_base_npu.pth`
- **ExecuTorch revision:** `d13c78971338b35a219e0a27f7045b29c7f4e8c1`
- **Parameters:** Approximately 6.1 million after inference-time fusion
- **Artifact size:** 23.47 MiB, 0.27% smaller than the 23.53 MiB FP32 artifact
## 🚀 Get Started with the Model
### 🔓 Compute Flow — Early Access
The inference engine for this model package is available through the Compute Flow Early Access
Program. Contact ai-early-access@arm.com to request access.
## 📊 Quality evaluation
Quality was evaluated on all 654 images in the NYUv2 test split. Both models received the same
top-left square crop resized to 384×384. Their predictions were resized to the ground-truth shape,
aligned independently in inverse-depth space using a least-squares scale and shift, and evaluated
over the NYUv2 Eigen crop and 0.001–10 m depth range. This isolates the static depth models; it is
not an evaluation of the dynamic prepare/restore transforms.
| Metric | FP32 | FP16 ExecuTorch | Change |
| --- | ---: | ---: | ---: |
| AbsRel ↓ | 15.796% | **15.785%** | 0.011 percentage points lower |
| δ1 ↑ | 77.055% | 77.009% | 0.046 percentage points lower |
The FP16 candidate uses a different official ZipDepth checkpoint because the mobile/NPU checkpoint
replaces convex unfold upsampling with a deployment-friendly learned blend. The comparison
therefore measures the complete mobile graph and FP16 conversion, not FP16 rounding in isolation.
On the distributed 384×384 sample, its output has 0.99645 Pearson correlation with the FP32 output.
## 🎯 Performance evaluation
The static depth models were profiled using the distributed `samples/im0.jpg` image, which is
already 384×384. The dynamic prepare and restore models were therefore bypassed.
- **CPU:** One Arm CPU thread.
- **Runs:** One unmeasured in-process inference to initialize each runtime, followed by 5 warmup
and 30 measured fresh-process runs with 1 second between processes.
- **Timing scope:** The setup call is not included.
- **Runtimes:** LiteRT 2.1.6 FP32 baseline and ExecuTorch `d13c7897` FP16 candidate, both with the
XNNPACK CPU backend.
- **Target:** Android™ Vivo X300 smartphone (`V2502A`) on Android™ 16, USB powered at 100% battery.
| Metric | Android™ Vivo X300: FP32 | Android™ Vivo X300: FP16 | Uplift |
| --- | ---: | ---: | ---: |
| p50 inference latency | 41.987 ms | **20.431 ms** | **2.06× faster** |
| p90 inference latency | 42.612 ms | **21.342 ms** | **2.00× faster** |
| p99 inference latency | 43.236 ms | **27.181 ms** | **1.59× faster** |
## 🛠️ Technical Specifications
### Runtime Architecture
| Component role | Framework / format |
| --- | --- |
| Input square crop and resize | ExecuTorch / PTE |
| ZipDepth Base inference | ExecuTorch / PTE |
| Output resize and zero padding | ExecuTorch / PTE |
### Precision
- Internal depth-model activations use FP16.
- Public input and output tensors remain FP32 to preserve the existing brick contract.
- Dynamic image-shape transforms use FP32.
- The fixed 24×24-to-12×12 adaptive average pool is represented as an equivalent 2×2,
stride-2 average pool supported by the deployed ExecuTorch kernel set.
### Input Specification
| Input | Data type | Shape | Value range |
| --- | --- | --- | --- |
| RGB image | FP32 | `3 × H × W`, with `1 ≤ H,W ≤ 4096` | `[0, 1]` |
### Output Specification
| Output | Data type | Shape | Description |
| --- | --- | --- | --- |
| Relative depth | FP32 | `H × W` | Affine-invariant inverse depth; padded regions are zero |
### Repository Contents
- `zipdepth_base_fp16.pte` — static FP16 ZipDepth Base mobile/NPU depth model.
- `zipdepth_prepare_image.pte` — dynamic input crop-and-resize transform.
- `zipdepth_restore_depth.pte` — dynamic output resize-and-padding transform.
- `zipdepth_manifest.json` — model tensor contracts and runtime configuration.
- `samples/im0.jpg` — 384×384 upstream sample used for profiling.
- `metadata.yaml` and `benchmarks/` — model and benchmark metadata.
- `LICENSE` — upstream MIT license.
- `SHA256SUMS` — reproducibility checksums.
## ⚠️ Known Limitations
- The square-crop package pipeline does not preserve the aspect-ratio preprocessing used for
ZipDepth's published benchmark results.
- The output is affine-invariant relative depth, not metric depth.
- The model has no temporal-consistency mechanism and may flicker when applied independently to
video frames.
## 🗂️ Model and Asset Origin
- [`zipdepth_base_npu.pth`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470/checkpoints)
is the source checkpoint for `zipdepth_base_fp16.pte`.
- [`samples/im0.jpg`](https://github.com/fabiotosi92/ZipDepth/blob/a302e5437bc58f15c4efd41d3e8222bf24f7d470/assets/examples/im0.jpg)
is an upstream example crop.
## 🔐 Checksums
Verify the checked-out files with:
```sh
shasum -a 256 -c SHA256SUMS
```
## Citation
```bibtex
@inproceedings{tosi2026zipdepth,
title = {{ZipDepth}: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device},
author = {Tosi, Fabio and Bartolomei, Luca and Poggi, Matteo and Mattoccia, Stefano},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}
```