FreeMatching

Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching

Luping Liu, Bingyi Kang, Yifan Wang, Dong Xu

Project repository

FreeMatching combines generative and semantic foundation representations, heterogeneous correspondence supervision, and teacher-guided iterative refinement for dense correspondence, including image editing and reference-guided generation.

Released checkpoints

Directory Snapshot Saved step File size
stage1/ Supervised model 275000 7,762,299,928 bytes
stage2/ EMA-refined model 10000 7,762,299,928 bytes

Each directory contains model.safetensors and config.json. The saved-step counters identify the original checkpoints. Each checkpoint contains 257 tensors for the transformer, semantic projection, and correspondence decoder. Optimizer state and experiment metadata are omitted.

from huggingface_hub import snapshot_download
snapshot_download(
    "luping-liu/FreeMatching",
    allow_patterns=["stage2/*"],
    local_dir="checkpoints",
)

Replace stage2/* with stage1/* to obtain the supervised model. File hashes are listed in SHA256SUMS and the per-stage configuration files.

Model requirements and output convention

Use the FreeMatching implementation with its modified FLUX transformer/pipeline and learned flow decoder. These files are not directly loadable via Transformers AutoModel.from_pretrained.

  • Generative backbone: FLUX.2-klein-base-4B.
  • Semantic encoder: DINOv3 ViT-L/16 (facebook/dinov3-vitl16-pretrain-lvd1689m).
  • Default inference size: 512×288; weights use BF16.
  • Default timestep: 620, configurable, preserving the original checkpoint evaluation setting.
  • Output direction: target-to-source, normalized absolute (x,y) coordinates, align_corners=True.
  • Stage 2 uses the EMA model. Its fixed teacher is used during refinement and is not needed for inference.

The frozen VAE/text encoder and DINOv3 encoder must be obtained separately from their official repositories. DINOv3 access is governed by its own license and access conditions. Third-party dependencies retain their respective terms.

Data availability

Due to the size of the video data and copyright restrictions on the source videos, we do not distribute the video corpus or precomputed video trajectory annotations. The code release provides loading/conversion code and preparation instructions for videos users obtain and are authorized to use.

Validation

Both snapshots were exported as tensor-only safetensors. Every tensor name and shape was checked against the release architecture, and uploaded file sizes and SHA-256 hashes were verified. CPU checks cover the custom transformer, correspondence warping, checkpoint loading, and trajectory conversion; this packaging pass does not include a full GPU inference or retraining validation.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for luping-liu/FreeMatching

Finetuned
(53)
this model