--- library_name: executorch license: mit base_model: fabiotosi92/ZipDepth display_name: ZipDepth Base FP16 — ExecuTorch tags: - depth-estimation - monocular-depth-estimation - zipdepth - executorch - xnnpack - fp16 - arm - edge-ai pipeline_tag: depth-estimation --- # ZipDepth Base FP16 — ExecuTorch This repository packages the unfold-free ZipDepth Base mobile/NPU checkpoint as a static 384×384 ExecuTorch model. The depth network uses FP16 activations with FP32 serialized convolution weights and external tensors. ## ✨ Key Highlights - **2.06× faster on Android™ Vivo X300** — one-thread p50 depth-model latency is **20.431 ms**, compared with **41.987 ms** for FP32 under the same inference-only protocol. - **Negligible quality change** — across all 654 NYUv2 test images, δ1 changes by only **−0.046 percentage points** and AbsRel improves by **0.011 percentage points** relative to the FP32 depth model under identical fixed-shape preprocessing and evaluation. ## 📦 Model Details ZipDepth predicts affine-invariant inverse depth from one RGB image. The package accepts arbitrary image dimensions up to 4096×4096, crops the largest top-left square, resizes it to the static 384×384 depth-model input, and restores the result to the original tensor shape with zero padding. - **Developed by:** Fabio Tosi, Luca Bartolomei, Matteo Poggi, and Stefano Mattoccia - **Model type:** Zero-shot monocular relative-depth estimation - **License:** MIT - **Base model:** [ZipDepth](https://github.com/fabiotosi92/ZipDepth) - **Source revision:** [`a302e543`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470) - **Source checkpoint:** `zipdepth_base_npu.pth` - **ExecuTorch revision:** `d13c78971338b35a219e0a27f7045b29c7f4e8c1` - **Parameters:** Approximately 6.1 million after inference-time fusion - **Artifact size:** 23.47 MiB, 0.27% smaller than the 23.53 MiB FP32 artifact ## 🚀 Get Started with the Model ### 🔓 Compute Flow — Early Access The inference engine for this model package is available through the Compute Flow Early Access Program. Contact ai-early-access@arm.com to request access. ## 📊 Quality evaluation Quality was evaluated on all 654 images in the NYUv2 test split. Both models received the same top-left square crop resized to 384×384. Their predictions were resized to the ground-truth shape, aligned independently in inverse-depth space using a least-squares scale and shift, and evaluated over the NYUv2 Eigen crop and 0.001–10 m depth range. This isolates the static depth models; it is not an evaluation of the dynamic prepare/restore transforms. | Metric | FP32 | FP16 ExecuTorch | Change | | --- | ---: | ---: | ---: | | AbsRel ↓ | 15.796% | **15.785%** | 0.011 percentage points lower | | δ1 ↑ | 77.055% | 77.009% | 0.046 percentage points lower | The FP16 candidate uses a different official ZipDepth checkpoint because the mobile/NPU checkpoint replaces convex unfold upsampling with a deployment-friendly learned blend. The comparison therefore measures the complete mobile graph and FP16 conversion, not FP16 rounding in isolation. On the distributed 384×384 sample, its output has 0.99645 Pearson correlation with the FP32 output. ## 🎯 Performance evaluation The static depth models were profiled using the distributed `samples/im0.jpg` image, which is already 384×384. The dynamic prepare and restore models were therefore bypassed. - **CPU:** One Arm CPU thread. - **Runs:** One unmeasured in-process inference to initialize each runtime, followed by 5 warmup and 30 measured fresh-process runs with 1 second between processes. - **Timing scope:** The setup call is not included. - **Runtimes:** LiteRT 2.1.6 FP32 baseline and ExecuTorch `d13c7897` FP16 candidate, both with the XNNPACK CPU backend. - **Target:** Android™ Vivo X300 smartphone (`V2502A`) on Android™ 16, USB powered at 100% battery. | Metric | Android™ Vivo X300: FP32 | Android™ Vivo X300: FP16 | Uplift | | --- | ---: | ---: | ---: | | p50 inference latency | 41.987 ms | **20.431 ms** | **2.06× faster** | | p90 inference latency | 42.612 ms | **21.342 ms** | **2.00× faster** | | p99 inference latency | 43.236 ms | **27.181 ms** | **1.59× faster** | ## 🛠️ Technical Specifications ### Runtime Architecture | Component role | Framework / format | | --- | --- | | Input square crop and resize | ExecuTorch / PTE | | ZipDepth Base inference | ExecuTorch / PTE | | Output resize and zero padding | ExecuTorch / PTE | ### Precision - Internal depth-model activations use FP16. - Public input and output tensors remain FP32 to preserve the existing brick contract. - Dynamic image-shape transforms use FP32. - The fixed 24×24-to-12×12 adaptive average pool is represented as an equivalent 2×2, stride-2 average pool supported by the deployed ExecuTorch kernel set. ### Input Specification | Input | Data type | Shape | Value range | | --- | --- | --- | --- | | RGB image | FP32 | `3 × H × W`, with `1 ≤ H,W ≤ 4096` | `[0, 1]` | ### Output Specification | Output | Data type | Shape | Description | | --- | --- | --- | --- | | Relative depth | FP32 | `H × W` | Affine-invariant inverse depth; padded regions are zero | ### Repository Contents - `zipdepth_base_fp16.pte` — static FP16 ZipDepth Base mobile/NPU depth model. - `zipdepth_prepare_image.pte` — dynamic input crop-and-resize transform. - `zipdepth_restore_depth.pte` — dynamic output resize-and-padding transform. - `zipdepth_manifest.json` — model tensor contracts and runtime configuration. - `samples/im0.jpg` — 384×384 upstream sample used for profiling. - `metadata.yaml` and `benchmarks/` — model and benchmark metadata. - `LICENSE` — upstream MIT license. - `SHA256SUMS` — reproducibility checksums. ## ⚠️ Known Limitations - The square-crop package pipeline does not preserve the aspect-ratio preprocessing used for ZipDepth's published benchmark results. - The output is affine-invariant relative depth, not metric depth. - The model has no temporal-consistency mechanism and may flicker when applied independently to video frames. ## 🗂️ Model and Asset Origin - [`zipdepth_base_npu.pth`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470/checkpoints) is the source checkpoint for `zipdepth_base_fp16.pte`. - [`samples/im0.jpg`](https://github.com/fabiotosi92/ZipDepth/blob/a302e5437bc58f15c4efd41d3e8222bf24f7d470/assets/examples/im0.jpg) is an upstream example crop. ## 🔐 Checksums Verify the checked-out files with: ```sh shasum -a 256 -c SHA256SUMS ``` ## Citation ```bibtex @inproceedings{tosi2026zipdepth, title = {{ZipDepth}: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device}, author = {Tosi, Fabio and Bartolomei, Luca and Poggi, Matteo and Mattoccia, Stefano}, booktitle = {European Conference on Computer Vision (ECCV)}, year = {2026} } ```