MVID β multi-view intrinsic decomposition
Weights for the MVID demo. Decomposes images into I = A Β· S + R:
- A β diffuse albedo (reflectance, not base colour; metals go dark)
- S β three-channel coloured irradiance
- R β non-diffuse residual: speculars, light sources, interreflections
VGGT-Omega backbone (1.35B params) with alternating frame-wise and global attention, so a set of views is decomposed jointly rather than per-image.
Files
| file | dtype | size |
|---|---|---|
mvid_ep186_bf16.safetensors |
bfloat16 | 2.5 GB |
Load into MVIDModel with strict=False. Inference runs under
torch.autocast(bfloat16); against the fp32 weights the outputs differ by
<0.002 mean on albedo/shading/residual/depth.
Trained at 512 px. Inputs must be a multiple of patch size 16 β the aggregator silently drops the remainder otherwise.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support