FreeMatching
Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
Luping Liu, Bingyi Kang, Yifan Wang, Dong Xu
FreeMatching combines generative and semantic foundation representations, heterogeneous correspondence supervision, and teacher-guided iterative refinement for dense correspondence, including image editing and reference-guided generation.
Released checkpoints
| Directory | Snapshot | Saved step | File size |
|---|---|---|---|
stage1/ |
Supervised model | 275000 | 7,762,299,928 bytes |
stage2/ |
EMA-refined model | 10000 | 7,762,299,928 bytes |
Each directory contains model.safetensors and config.json. The saved-step counters identify the original checkpoints. Each checkpoint contains 257 tensors for the transformer, semantic projection, and correspondence decoder. Optimizer state and experiment metadata are omitted.
from huggingface_hub import snapshot_download
snapshot_download(
"luping-liu/FreeMatching",
allow_patterns=["stage2/*"],
local_dir="checkpoints",
)
Replace stage2/* with stage1/* to obtain the supervised model. File hashes are listed in SHA256SUMS and the per-stage configuration files.
Model requirements and output convention
Use the FreeMatching implementation with its modified FLUX transformer/pipeline and learned flow decoder. These files are not directly loadable via Transformers AutoModel.from_pretrained.
- Generative backbone: FLUX.2-klein-base-4B.
- Semantic encoder: DINOv3 ViT-L/16 (
facebook/dinov3-vitl16-pretrain-lvd1689m). - Default inference size: 512×288; weights use BF16.
- Default timestep: 620, configurable, preserving the original checkpoint evaluation setting.
- Output direction: target-to-source, normalized absolute
(x,y)coordinates,align_corners=True. - Stage 2 uses the EMA model. Its fixed teacher is used during refinement and is not needed for inference.
The frozen VAE/text encoder and DINOv3 encoder must be obtained separately from their official repositories. DINOv3 access is governed by its own license and access conditions. Third-party dependencies retain their respective terms.
Data availability
Due to the size of the video data and copyright restrictions on the source videos, we do not distribute the video corpus or precomputed video trajectory annotations. The code release provides loading/conversion code and preparation instructions for videos users obtain and are authorized to use.
Validation
Both snapshots were exported as tensor-only safetensors. Every tensor name and shape was checked against the release architecture, and uploaded file sizes and SHA-256 hashes were verified. CPU checks cover the custom transformer, correspondence warping, checkpoint loading, and trajectory conversion; this packaging pass does not include a full GPU inference or retraining validation.
Model tree for luping-liu/FreeMatching
Base model
black-forest-labs/FLUX.2-klein-base-4B