goblin-models / README.md
glowskeleton's picture
Add RIFE 4.25 exported as flow + mask fields (Time Motion (ML) retime)
4e93798 verified
|
Raw History Blame Contribute Delete
4.52 kB
---
license: apache-2.0
tags:
- onnx
- segmentation
- matting
- sam2
- vitmatte
- rife
- frame-interpolation
---
# Goblin models
ONNX artifacts for [Goblin](https://github.com/rotoshake/Goblin), a
GPU-accelerated node-based compositor. Everything Goblin's model installer
downloads lives here, so a given Goblin version pins exactly the artifacts it
was tested against rather than tracking upstream `main`.
Four groups.
## `sam2-video/` β€” Goblin re-exports for SAM2 video propagation
These are **not** upstream files. `facebook/sam2.1-hiera-tiny` has no published
ONNX export of the video path, so these were produced for Goblin:
| file | what it is |
|---|---|
| `memory_attention.onnx` | needed a real-valued RoPE rewrite β€” SAM2's complex tensors have no ONNX representation |
| `memory_encoder.onnx` | memory encoder export |
| `prompt_encoder_mask_decoder.onnx` | re-export carrying the `obj_ptr` output propagation needs, all four mask tokens, and the `input_masks` dense prompt |
| `*.f32` | constant tensors dumped from the PyTorch checkpoint β€” not models |
The decoder here **replaces** the upstream one. The upstream build has no
`obj_ptr` output, so video propagation cannot run against it.
## `vitmatte/` β€” ViTMatte with Einsum rewritten as MatMul
`model.onnx` is [`Xenova/vitmatte-small-composition-1k`](https://huggingface.co/Xenova/vitmatte-small-composition-1k)
with its 24 `Einsum` nodes β€” 12 `bhwc,hkc->bhwk` and 12 `bhwc,wkc->bhwk`, ViT's
decomposed relative position bias β€” rewritten as broadcasting `MatMul`.
That single op was the reason the model could only run on CPU:
* CoreML cannot execute `Einsum`, so it partitioned around all 24 (measured at
~51 partitions, no speedup over CPU).
* DirectML has a documented `Einsum` wrong-results bug
([onnxruntime#19837](https://github.com/microsoft/onnxruntime/issues/19837))
that silently corrupts the matte.
The rewrite is **bit-identical** on the CPU execution provider β€” max absolute
difference 0.000e+00 over random inputs β€” and introduces no new operator and no
opset change. Produced and verified by
[`tools/ml/rewrite_vitmatte_einsum.py`](https://github.com/rotoshake/Goblin/blob/main/tools/ml/rewrite_vitmatte_einsum.py),
which refuses to emit a model whose output drifts by more than 1e-4.
Goblin selects the execution provider from the graph rather than from a flag: a
model still containing `Einsum` runs on CPU regardless of settings, because a
wrong matte looks like a slightly different matte rather than an error.
## `rife/` β€” RIFE 4.25 exported as fields, for Time's Motion (ML) retime
`interp_model.onnx` is [Practical-RIFE](https://github.com/hzwer/Practical-RIFE)
4.25 exported so that it returns the two fields its output is built from,
**not** a finished frame:
| output | shape | meaning |
|---|---|---|
| `flow` | `[1,4,H,W]` | pixel offsets: frame 0 β†’ t in `xy`, frame 1 β†’ t in `zw` |
| `mask` | `[1,1,H,W]` | weight of the frame-0 side, sigmoid applied |
Inputs are `img0`, `img1` (`[1,3,H,W]`, sRGB 0..1, H and W multiples of 64)
and `timestep` (`[1,1,1,1]`). Height and width stay dynamic.
RIFE 4.x has no refinement network, so
`warp(img0, flow.xy) * mask + warp(img1, flow.zw) * (1 - mask)` **is** the
model's output. Goblin applies that warp itself, to the original linear float
frames at full resolution, which is how HDR highlights and alpha survive a
model trained on 8-bit sRGB. The export verifies the fields rebuild stock
IFNet's frame bit-identically (max difference 0.000e+00) and checks the graph
against PyTorch at sizes other than the traced one. Produced by
[`tools/ml/export_rife_fields.py`](https://github.com/rotoshake/Goblin/blob/main/tools/ml/export_rife_fields.py)
from the official 4.25 weights (`flownet_v4.25.pkl`).
## `sam2-upstream/` β€” mirror of the SAM2 image-mode files
Byte-for-byte copies of the ONNX from
[`onnx-community/sam2.1-hiera-tiny-ONNX`](https://huggingface.co/onnx-community/sam2.1-hiera-tiny-ONNX),
mirrored so an install is self-contained and cannot change under a force-push
upstream. `prompt_encoder_mask_decoder.onnx` here is the original community
build; `sam2-video/` holds the re-export that supersedes it.
## Licensing
SAM 2 is Apache 2.0 (Meta). ViTMatte weights derive from
`hustvl/vitmatte-small-composition-1k`. RIFE is MIT (hzwer, Practical-RIFE); its
license is in `rife/LICENSE-Practical-RIFE`. Derived artifacts here carry their
upstream licenses; the export and rewrite tooling is part of Goblin.