|
Download README.md from Arm/zipdepth-fp16-executorch: direct link, hf CLI and curl.
- Browser
- Download file 7.17 kB
-
https://huggingface.co/Arm/zipdepth-fp16-executorch/resolve/main/README.md
- Command line
-
hf download hf://Arm/zipdepth-fp16-executorch/README.md
-
curl -L -o README.md https://huggingface.co/Arm/zipdepth-fp16-executorch/resolve/main/README.md
7.17 kB
| library_name: executorch | |
| license: mit | |
| base_model: fabiotosi92/ZipDepth | |
| display_name: ZipDepth Base FP16 — ExecuTorch | |
| tags: | |
| - depth-estimation | |
| - monocular-depth-estimation | |
| - zipdepth | |
| - executorch | |
| - xnnpack | |
| - fp16 | |
| - arm | |
| - edge-ai | |
| pipeline_tag: depth-estimation | |
| <!-- | |
| SPDX-FileCopyrightText: Copyright 2026 Arm Limited and/or its affiliates <open-source-office@arm.com> | |
| SPDX-License-Identifier: Apache-2.0 | |
| --> | |
| # ZipDepth Base FP16 — ExecuTorch | |
| This repository packages the unfold-free ZipDepth Base mobile/NPU checkpoint as a static 384×384 | |
| ExecuTorch model. The depth network uses FP16 activations with FP32 serialized convolution weights | |
| and external tensors. | |
| ## ✨ Key Highlights | |
| - **2.06× faster on Android™ Vivo X300** — one-thread p50 depth-model latency is **20.431 ms**, | |
| compared with **41.987 ms** for FP32 under the same inference-only protocol. | |
| - **Negligible quality change** — across all 654 NYUv2 test images, δ1 changes by only | |
| **−0.046 percentage points** and AbsRel improves by **0.011 percentage points** relative to the | |
| FP32 depth model under identical fixed-shape preprocessing and evaluation. | |
| ## 📦 Model Details | |
| ZipDepth predicts affine-invariant inverse depth from one RGB image. The package accepts arbitrary | |
| image dimensions up to 4096×4096, crops the largest top-left square, resizes it to the static | |
| 384×384 depth-model input, and restores the result to the original tensor shape with zero padding. | |
| - **Developed by:** Fabio Tosi, Luca Bartolomei, Matteo Poggi, and Stefano Mattoccia | |
| - **Model type:** Zero-shot monocular relative-depth estimation | |
| - **License:** MIT | |
| - **Base model:** [ZipDepth](https://github.com/fabiotosi92/ZipDepth) | |
| - **Source revision:** [`a302e543`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470) | |
| - **Source checkpoint:** `zipdepth_base_npu.pth` | |
| - **ExecuTorch revision:** `d13c78971338b35a219e0a27f7045b29c7f4e8c1` | |
| - **Parameters:** Approximately 6.1 million after inference-time fusion | |
| - **Artifact size:** 23.47 MiB, 0.27% smaller than the 23.53 MiB FP32 artifact | |
| ## 🚀 Get Started with the Model | |
| ### 🔓 Compute Flow — Early Access | |
| The inference engine for this model package is available through the Compute Flow Early Access | |
| Program. Contact ai-early-access@arm.com to request access. | |
| ## 📊 Quality evaluation | |
| Quality was evaluated on all 654 images in the NYUv2 test split. Both models received the same | |
| top-left square crop resized to 384×384. Their predictions were resized to the ground-truth shape, | |
| aligned independently in inverse-depth space using a least-squares scale and shift, and evaluated | |
| over the NYUv2 Eigen crop and 0.001–10 m depth range. This isolates the static depth models; it is | |
| not an evaluation of the dynamic prepare/restore transforms. | |
| | Metric | FP32 | FP16 ExecuTorch | Change | | |
| | --- | ---: | ---: | ---: | | |
| | AbsRel ↓ | 15.796% | **15.785%** | 0.011 percentage points lower | | |
| | δ1 ↑ | 77.055% | 77.009% | 0.046 percentage points lower | | |
| The FP16 candidate uses a different official ZipDepth checkpoint because the mobile/NPU checkpoint | |
| replaces convex unfold upsampling with a deployment-friendly learned blend. The comparison | |
| therefore measures the complete mobile graph and FP16 conversion, not FP16 rounding in isolation. | |
| On the distributed 384×384 sample, its output has 0.99645 Pearson correlation with the FP32 output. | |
| ## 🎯 Performance evaluation | |
| The static depth models were profiled using the distributed `samples/im0.jpg` image, which is | |
| already 384×384. The dynamic prepare and restore models were therefore bypassed. | |
| - **CPU:** One Arm CPU thread. | |
| - **Runs:** One unmeasured in-process inference to initialize each runtime, followed by 5 warmup | |
| and 30 measured fresh-process runs with 1 second between processes. | |
| - **Timing scope:** The setup call is not included. | |
| - **Runtimes:** LiteRT 2.1.6 FP32 baseline and ExecuTorch `d13c7897` FP16 candidate, both with the | |
| XNNPACK CPU backend. | |
| - **Target:** Android™ Vivo X300 smartphone (`V2502A`) on Android™ 16, USB powered at 100% battery. | |
| | Metric | Android™ Vivo X300: FP32 | Android™ Vivo X300: FP16 | Uplift | | |
| | --- | ---: | ---: | ---: | | |
| | p50 inference latency | 41.987 ms | **20.431 ms** | **2.06× faster** | | |
| | p90 inference latency | 42.612 ms | **21.342 ms** | **2.00× faster** | | |
| | p99 inference latency | 43.236 ms | **27.181 ms** | **1.59× faster** | | |
| ## 🛠️ Technical Specifications | |
| ### Runtime Architecture | |
| | Component role | Framework / format | | |
| | --- | --- | | |
| | Input square crop and resize | ExecuTorch / PTE | | |
| | ZipDepth Base inference | ExecuTorch / PTE | | |
| | Output resize and zero padding | ExecuTorch / PTE | | |
| ### Precision | |
| - Internal depth-model activations use FP16. | |
| - Public input and output tensors remain FP32 to preserve the existing brick contract. | |
| - Dynamic image-shape transforms use FP32. | |
| - The fixed 24×24-to-12×12 adaptive average pool is represented as an equivalent 2×2, | |
| stride-2 average pool supported by the deployed ExecuTorch kernel set. | |
| ### Input Specification | |
| | Input | Data type | Shape | Value range | | |
| | --- | --- | --- | --- | | |
| | RGB image | FP32 | `3 × H × W`, with `1 ≤ H,W ≤ 4096` | `[0, 1]` | | |
| ### Output Specification | |
| | Output | Data type | Shape | Description | | |
| | --- | --- | --- | --- | | |
| | Relative depth | FP32 | `H × W` | Affine-invariant inverse depth; padded regions are zero | | |
| ### Repository Contents | |
| - `zipdepth_base_fp16.pte` — static FP16 ZipDepth Base mobile/NPU depth model. | |
| - `zipdepth_prepare_image.pte` — dynamic input crop-and-resize transform. | |
| - `zipdepth_restore_depth.pte` — dynamic output resize-and-padding transform. | |
| - `zipdepth_manifest.json` — model tensor contracts and runtime configuration. | |
| - `samples/im0.jpg` — 384×384 upstream sample used for profiling. | |
| - `metadata.yaml` and `benchmarks/` — model and benchmark metadata. | |
| - `LICENSE` — upstream MIT license. | |
| - `SHA256SUMS` — reproducibility checksums. | |
| ## ⚠️ Known Limitations | |
| - The square-crop package pipeline does not preserve the aspect-ratio preprocessing used for | |
| ZipDepth's published benchmark results. | |
| - The output is affine-invariant relative depth, not metric depth. | |
| - The model has no temporal-consistency mechanism and may flicker when applied independently to | |
| video frames. | |
| ## 🗂️ Model and Asset Origin | |
| - [`zipdepth_base_npu.pth`](https://github.com/fabiotosi92/ZipDepth/tree/a302e5437bc58f15c4efd41d3e8222bf24f7d470/checkpoints) | |
| is the source checkpoint for `zipdepth_base_fp16.pte`. | |
| - [`samples/im0.jpg`](https://github.com/fabiotosi92/ZipDepth/blob/a302e5437bc58f15c4efd41d3e8222bf24f7d470/assets/examples/im0.jpg) | |
| is an upstream example crop. | |
| ## 🔐 Checksums | |
| Verify the checked-out files with: | |
| ```sh | |
| shasum -a 256 -c SHA256SUMS | |
| ``` | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{tosi2026zipdepth, | |
| title = {{ZipDepth}: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device}, | |
| author = {Tosi, Fabio and Bartolomei, Luca and Poggi, Matteo and Mattoccia, Stefano}, | |
| booktitle = {European Conference on Computer Vision (ECCV)}, | |
| year = {2026} | |
| } | |
| ``` | |