MVID β€” multi-view intrinsic decomposition

Weights for the MVID demo. Decomposes images into I = A Β· S + R:

  • A β€” diffuse albedo (reflectance, not base colour; metals go dark)
  • S β€” three-channel coloured irradiance
  • R β€” non-diffuse residual: speculars, light sources, interreflections

VGGT-Omega backbone (1.35B params) with alternating frame-wise and global attention, so a set of views is decomposed jointly rather than per-image.

Files

file dtype size
mvid_ep186_bf16.safetensors bfloat16 2.5 GB

Load into MVIDModel with strict=False. Inference runs under torch.autocast(bfloat16); against the fp32 weights the outputs differ by <0.002 mean on albedo/shading/residual/depth.

Trained at 512 px. Inputs must be a multiple of patch size 16 β€” the aggregator silently drops the remainder otherwise.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support