--- license: apache-2.0 tags: - onnx - segmentation - matting - sam2 - vitmatte - rife - frame-interpolation --- # Goblin models ONNX artifacts for [Goblin](https://github.com/rotoshake/Goblin), a GPU-accelerated node-based compositor. Everything Goblin's model installer downloads lives here, so a given Goblin version pins exactly the artifacts it was tested against rather than tracking upstream `main`. Four groups. ## `sam2-video/` — Goblin re-exports for SAM2 video propagation These are **not** upstream files. `facebook/sam2.1-hiera-tiny` has no published ONNX export of the video path, so these were produced for Goblin: | file | what it is | |---|---| | `memory_attention.onnx` | needed a real-valued RoPE rewrite — SAM2's complex tensors have no ONNX representation | | `memory_encoder.onnx` | memory encoder export | | `prompt_encoder_mask_decoder.onnx` | re-export carrying the `obj_ptr` output propagation needs, all four mask tokens, and the `input_masks` dense prompt | | `*.f32` | constant tensors dumped from the PyTorch checkpoint — not models | The decoder here **replaces** the upstream one. The upstream build has no `obj_ptr` output, so video propagation cannot run against it. ## `vitmatte/` — ViTMatte with Einsum rewritten as MatMul `model.onnx` is [`Xenova/vitmatte-small-composition-1k`](https://huggingface.co/Xenova/vitmatte-small-composition-1k) with its 24 `Einsum` nodes — 12 `bhwc,hkc->bhwk` and 12 `bhwc,wkc->bhwk`, ViT's decomposed relative position bias — rewritten as broadcasting `MatMul`. That single op was the reason the model could only run on CPU: * CoreML cannot execute `Einsum`, so it partitioned around all 24 (measured at ~51 partitions, no speedup over CPU). * DirectML has a documented `Einsum` wrong-results bug ([onnxruntime#19837](https://github.com/microsoft/onnxruntime/issues/19837)) that silently corrupts the matte. The rewrite is **bit-identical** on the CPU execution provider — max absolute difference 0.000e+00 over random inputs — and introduces no new operator and no opset change. Produced and verified by [`tools/ml/rewrite_vitmatte_einsum.py`](https://github.com/rotoshake/Goblin/blob/main/tools/ml/rewrite_vitmatte_einsum.py), which refuses to emit a model whose output drifts by more than 1e-4. Goblin selects the execution provider from the graph rather than from a flag: a model still containing `Einsum` runs on CPU regardless of settings, because a wrong matte looks like a slightly different matte rather than an error. ## `rife/` — RIFE 4.25 exported as fields, for Time's Motion (ML) retime `interp_model.onnx` is [Practical-RIFE](https://github.com/hzwer/Practical-RIFE) 4.25 exported so that it returns the two fields its output is built from, **not** a finished frame: | output | shape | meaning | |---|---|---| | `flow` | `[1,4,H,W]` | pixel offsets: frame 0 → t in `xy`, frame 1 → t in `zw` | | `mask` | `[1,1,H,W]` | weight of the frame-0 side, sigmoid applied | Inputs are `img0`, `img1` (`[1,3,H,W]`, sRGB 0..1, H and W multiples of 64) and `timestep` (`[1,1,1,1]`). Height and width stay dynamic. RIFE 4.x has no refinement network, so `warp(img0, flow.xy) * mask + warp(img1, flow.zw) * (1 - mask)` **is** the model's output. Goblin applies that warp itself, to the original linear float frames at full resolution, which is how HDR highlights and alpha survive a model trained on 8-bit sRGB. The export verifies the fields rebuild stock IFNet's frame bit-identically (max difference 0.000e+00) and checks the graph against PyTorch at sizes other than the traced one. Produced by [`tools/ml/export_rife_fields.py`](https://github.com/rotoshake/Goblin/blob/main/tools/ml/export_rife_fields.py) from the official 4.25 weights (`flownet_v4.25.pkl`). ## `sam2-upstream/` — mirror of the SAM2 image-mode files Byte-for-byte copies of the ONNX from [`onnx-community/sam2.1-hiera-tiny-ONNX`](https://huggingface.co/onnx-community/sam2.1-hiera-tiny-ONNX), mirrored so an install is self-contained and cannot change under a force-push upstream. `prompt_encoder_mask_decoder.onnx` here is the original community build; `sam2-video/` holds the re-export that supersedes it. ## Licensing SAM 2 is Apache 2.0 (Meta). ViTMatte weights derive from `hustvl/vitmatte-small-composition-1k`. RIFE is MIT (hzwer, Practical-RIFE); its license is in `rife/LICENSE-Practical-RIFE`. Derived artifacts here carry their upstream licenses; the export and rewrite tooling is part of Goblin.