ZipDepth Base FP16 — ExecuTorch
This repository packages the unfold-free ZipDepth Base mobile/NPU checkpoint as a static 384×384 ExecuTorch model. The depth network uses FP16 activations with FP32 serialized convolution weights and external tensors.
✨ Key Highlights
- 2.06× faster on Android™ Vivo X300 — one-thread p50 depth-model latency is 20.431 ms, compared with 41.987 ms for FP32 under the same inference-only protocol.
- Negligible quality change — across all 654 NYUv2 test images, δ1 changes by only −0.046 percentage points and AbsRel improves by 0.011 percentage points relative to the FP32 depth model under identical fixed-shape preprocessing and evaluation.
📦 Model Details
ZipDepth predicts affine-invariant inverse depth from one RGB image. The package accepts arbitrary image dimensions up to 4096×4096, crops the largest top-left square, resizes it to the static 384×384 depth-model input, and restores the result to the original tensor shape with zero padding.
- Developed by: Fabio Tosi, Luca Bartolomei, Matteo Poggi, and Stefano Mattoccia
- Model type: Zero-shot monocular relative-depth estimation
- License: MIT
- Base model: ZipDepth
- Source revision:
a302e543 - Source checkpoint:
zipdepth_base_npu.pth - ExecuTorch revision:
d13c78971338b35a219e0a27f7045b29c7f4e8c1 - Parameters: Approximately 6.1 million after inference-time fusion
- Artifact size: 23.47 MiB, 0.27% smaller than the 23.53 MiB FP32 artifact
🚀 Get Started with the Model
🔓 Compute Flow — Early Access
The inference engine for this model package is available through the Compute Flow Early Access Program. Contact ai-early-access@arm.com to request access.
📊 Quality evaluation
Quality was evaluated on all 654 images in the NYUv2 test split. Both models received the same top-left square crop resized to 384×384. Their predictions were resized to the ground-truth shape, aligned independently in inverse-depth space using a least-squares scale and shift, and evaluated over the NYUv2 Eigen crop and 0.001–10 m depth range. This isolates the static depth models; it is not an evaluation of the dynamic prepare/restore transforms.
| Metric | FP32 | FP16 ExecuTorch | Change |
|---|---|---|---|
| AbsRel ↓ | 15.796% | 15.785% | 0.011 percentage points lower |
| δ1 ↑ | 77.055% | 77.009% | 0.046 percentage points lower |
The FP16 candidate uses a different official ZipDepth checkpoint because the mobile/NPU checkpoint replaces convex unfold upsampling with a deployment-friendly learned blend. The comparison therefore measures the complete mobile graph and FP16 conversion, not FP16 rounding in isolation. On the distributed 384×384 sample, its output has 0.99645 Pearson correlation with the FP32 output.
🎯 Performance evaluation
The static depth models were profiled using the distributed samples/im0.jpg image, which is
already 384×384. The dynamic prepare and restore models were therefore bypassed.
- CPU: One Arm CPU thread.
- Runs: One unmeasured in-process inference to initialize each runtime, followed by 5 warmup and 30 measured fresh-process runs with 1 second between processes.
- Timing scope: The setup call is not included.
- Runtimes: LiteRT 2.1.6 FP32 baseline and ExecuTorch
d13c7897FP16 candidate, both with the XNNPACK CPU backend. - Target: Android™ Vivo X300 smartphone (
V2502A) on Android™ 16, USB powered at 100% battery.
| Metric | Android™ Vivo X300: FP32 | Android™ Vivo X300: FP16 | Uplift |
|---|---|---|---|
| p50 inference latency | 41.987 ms | 20.431 ms | 2.06× faster |
| p90 inference latency | 42.612 ms | 21.342 ms | 2.00× faster |
| p99 inference latency | 43.236 ms | 27.181 ms | 1.59× faster |
🛠️ Technical Specifications
Runtime Architecture
| Component role | Framework / format |
|---|---|
| Input square crop and resize | ExecuTorch / PTE |
| ZipDepth Base inference | ExecuTorch / PTE |
| Output resize and zero padding | ExecuTorch / PTE |
Precision
- Internal depth-model activations use FP16.
- Public input and output tensors remain FP32 to preserve the existing brick contract.
- Dynamic image-shape transforms use FP32.
- The fixed 24×24-to-12×12 adaptive average pool is represented as an equivalent 2×2, stride-2 average pool supported by the deployed ExecuTorch kernel set.
Input Specification
| Input | Data type | Shape | Value range |
|---|---|---|---|
| RGB image | FP32 | 3 × H × W, with 1 ≤ H,W ≤ 4096 |
[0, 1] |
Output Specification
| Output | Data type | Shape | Description |
|---|---|---|---|
| Relative depth | FP32 | H × W |
Affine-invariant inverse depth; padded regions are zero |
Repository Contents
zipdepth_base_fp16.pte— static FP16 ZipDepth Base mobile/NPU depth model.zipdepth_prepare_image.pte— dynamic input crop-and-resize transform.zipdepth_restore_depth.pte— dynamic output resize-and-padding transform.zipdepth_manifest.json— model tensor contracts and runtime configuration.samples/im0.jpg— 384×384 upstream sample used for profiling.metadata.yamlandbenchmarks/— model and benchmark metadata.LICENSE— upstream MIT license.SHA256SUMS— reproducibility checksums.
⚠️ Known Limitations
- The square-crop package pipeline does not preserve the aspect-ratio preprocessing used for ZipDepth's published benchmark results.
- The output is affine-invariant relative depth, not metric depth.
- The model has no temporal-consistency mechanism and may flicker when applied independently to video frames.
🗂️ Model and Asset Origin
zipdepth_base_npu.pthis the source checkpoint forzipdepth_base_fp16.pte.samples/im0.jpgis an upstream example crop.
🔐 Checksums
Verify the checked-out files with:
shasum -a 256 -c SHA256SUMS
Citation
@inproceedings{tosi2026zipdepth,
title = {{ZipDepth}: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device},
author = {Tosi, Fabio and Bartolomei, Luca and Poggi, Matteo and Mattoccia, Stefano},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}
- Downloads last month
- 9