|
Download README.md from sm079/sharp-onnx-webgpu: direct link, hf CLI and curl.
- Browser
- Download file 2.5 kB
-
https://huggingface.co/sm079/sharp-onnx-webgpu/resolve/main/README.md
- Command line
-
hf download hf://sm079/sharp-onnx-webgpu/README.md
-
curl -L -o README.md https://huggingface.co/sm079/sharp-onnx-webgpu/resolve/main/README.md
2.5 kB
| license: other | |
| license_name: apple-ml-research-model-license | |
| license_link: LICENSE | |
| library_name: onnx | |
| pipeline_tag: image-to-3d | |
| tags: | |
| - onnx | |
| - onnxruntime-web | |
| - webgpu | |
| - gaussian-splatting | |
| - single-image-3d | |
| # SHARP (ONNX, WebGPU) | |
| An ONNX export of [SHARP](https://github.com/apple/ml-sharp) ("Sharp Monocular View Synthesis in Less Than a Second"), prepared for in-browser inference with ONNX Runtime Web on WebGPU. It is used by [SharpRig](https://github.com/sm079/sharp-rig), which turns a single photo into a camera-move video in the browser. | |
| > Apple Machine Learning Research Model is licensed under the Apple Machine Learning Research Model License Agreement. | |
| This is a **model derivative**, not an official Apple release, and it is not endorsed by Apple. Like the original, it may be used for **non-commercial research purposes only**; see [LICENSE](LICENSE). | |
| ## Files | |
| | File | Size | | | |
| |---|---|---| | |
| | `sharp.onnx` | 4 MB | graph, with the weights stored as external data (`sharp.onnx.data`) | | |
| | `sharp.int8.bin` | 0.66 GB | the external weights as int8 + per-output-channel scales, expanded back to fp16 on load | | |
| ## Modifications from the original checkpoint | |
| - **Export:** the network was traced in fp16 (fp32 inputs and outputs) from `sharp_2572gikvuh.pt` and exported to ONNX. | |
| - **Unprojection moved out of the graph:** the final unprojection from NDC to metric space, which needs an SVD that ONNX can't express, is left to the caller as a per-axis scale. | |
| - **Int8 packing:** the fp16 weights are packed as int8 with one float32 scale per output channel (round-to-nearest, weight-only). This halves the download. The pack is dequantised back to fp16 before inference, so the numerics are those of an fp16 model with int8-rounded weights. | |
| - **Accuracy:** measured against the fp32 checkpoint on a test photo, relative depth error is 0.19% (median) and 1.1% (p95). | |
| - **Code:** the export and packing are done by [`tools/export_sharp_onnx.py`](https://github.com/sm079/sharp-rig/blob/main/tools/export_sharp_onnx.py). No retraining or fine-tuning was done. | |
| ## Inputs and outputs | |
| - **Inputs:** `image` float32 `[1, 3, 1536, 1536]` (RGB in [0, 1]); `disparity_factor` float32 `[1]` (focal length in px / image width). | |
| - **Outputs** (N = 2 × 768 × 768): | |
| - `mean_vectors` `[1, N, 3]` | |
| - `singular_values` `[1, N, 3]` | |
| - `quaternions` `[1, N, 4]` (w, x, y, z) | |
| - `colors` `[1, N, 3]` (linear RGB) | |
| - `opacities` `[1, N]` | |
| All are Gaussians in SHARP's NDC space. | |