Request access to Jeff-Tracker weights

Access to the trained checkpoint is granted per request, so the maintainer knows who is using it and in what context. Requests are reviewed manually.

Log in or Sign Up to review the conditions and access this model content.

Jeff-Tracker

Track any point through a video.

Give it a video and a set of points; it returns where each point is on every frame, whether it can see it, and how sure it is. It runs at 0.0065 s/frame and returns a calibrated per-frame confidence alongside every track.

Code, benchmarks and method: https://github.com/antonyjaguar19-dotcom/Jeff-Tracker

It is LocoTrack-B with cross-track attention added β€” the tracks attend to each other, so a point that is hidden can be inferred from the neighbours that are still visible β€” then fine-tuned for occlusion β€” on MOVi-E, and in the newer checkpoint on Meta's Apache-2.0 Kubric release.

Access

This repository is gated. Click Request access; the maintainer approves manually. Once approved:

hf auth login
python tools/fetch_weights.py --all

Files

file what it is
c8_kubric.ckpt trained on Meta's CoTracker3_Kubric. Wins on the synthetic benches and has not been shown to help on real footage -- read "Which checkpoint" below, 66 MB
inf_s4000.ckpt the original release checkpoint. Step 4000, occ_pos_weight=0.2, 66 MB

The LocoTrack baseline weights are not mirrored here; they come from the LocoTrack authors' own release, hamacojr/LocoTrack-pytorch-weights.

Use

import torch
tracker = torch.hub.load("antonyjaguar19-dotcom/Jeff-Tracker", "jefftracker").cuda()

# frames: (T, H, W, 3) uint8 BGR      queries: (N, 3) as [frame, x, y]
tracks, vis, conf = tracker.track_queries_conf(frames, queries)
# tracks (T, N, 2) xy, y DOWN   vis (T, N) bool   conf (T, N) float

Keep conf. At 256Γ—256 it detects the model's own >5 px frames at AUC 0.955, and gating at 0.5 capped the worst sample in a whole run at 9 px.

Results

Two checkpoints, and which one is better depends on what you are tracking. Both numbers below are measured here on the same day with the same binary.

inf_s4000 c8_kubric
TAP-Vid DAVIS strided (AJ / Ξ΄_avg / OA) 62.6 / 74.9 / 86.7 see below
Localisation, exact synthetic ground truth 1.03 px @ 256Γ—256 ~3% tighter
Pixel locking +0.0005 px β€” none, on an NCC-free path same
Confidence as a bad-frame detector AUC 0.955 @ 256Γ—256 same path
Speed 0.0065 s/frame Β· 6.85 GB peak on 4K same

Which checkpoint β€” and why the honest answer is the older one

c8_kubric was compared against c5_occl1 with three late checkpoints per arm, and on the synthetic occlusion benches it separated better on 5 of 5 for visible accuracy (3%) and re-acquisition (4.5%).

That reads stronger than it is, and the caveat is the important part: all five of those benches derive from one synthetic shot β€” a textured flat plane moved by a homography, with cards or depth planes composited over it. Five benches, one plate. A gain there says the model improved on textured planes under clean motion.

Every measurement on real footage says something else:

test better
synthetic occlusion benches (5, one source plate) c8_kubric
TAP-Vid DAVIS, 30 real clips c5_occl1 β€” 68.3 / 79.8 / 89.9 against 67.2 / 79.3 / 88.8
real 4K plate, frames the model will commit to c5_occl1 β€” 53.1% against 48.7%
real plate against a hand-built reference tie, inside the reference's own 0.60 px noise
real plate, forward-and-back closure c5_occl1 β€” mean 0.67 px against 0.72, worst 7.2 against 12.6
a second real plate, closure mixed β€” c8_kubric the better tail, c5_occl1 the better median

So c5_occl1 remains the recommendation. c8_kubric is published because the training is reproducible and its tail is better on one plate, not because it is an upgrade. If your footage resembles clean synthetic motion it may help; on handheld plates with grain and defocus the evidence does not support the swap.

This section is worded the way it is because the first version of this card recommended c8_kubric on the bench result alone, before the real-footage numbers existed.

Unchanged either way: accuracy while a point is hidden. 2.90 px against CoTracker3's 1.69 on the same bench. Neither checkpoint moved that column β€” see Limitations.

Limitations β€” read before deploying

It gaps an occlusion rather than crossing one, and c8_kubric gaps MORE. Willingness to commit a position on a ground-truth-occluded frame, at confidence gate 0.5:

occluded frames given a position
c5_occl1 ~5.5%
c8_kubric ~3.7%

The newer checkpoint is more cautious, not more capable. That is the whole explanation of its result: fewer committed guesses give tighter visible tracks and better re-acquisition, and no progress through an occlusion. Whether this is an improvement depends on your pipeline. For matchmove it generally is β€” a clean hole is visible and fillable, while a confidently wrong position inside an occluder looks correct in the viewport and quietly poisons a camera solve. If your downstream stage requires a position on every frame, it is a regression.

Five separate attempts at the occlusion gap have now been measured and none moved it: more neighbours per chunk (0.002 px at 9Γ— the neighbours), neighbours carrying temporal state (a wash), a fourth correlation level (worse on three benches), better training data (this checkpoint β€” localisation yes, occlusion no), and doubling the training window from 24 to 48 frames (no effect; the coverage figure above was identical at both). Three of those change the architecture and two change what it is fed. Treat the occluded column as a known structural limitation of this architecture rather than a tuning problem.

Use 256Γ—256. At 384Γ—680 the median improves to 0.697 px while the mean is 40.6 px, and no confidence threshold recovers it β€” at conf β‰₯ 0.99 the mean is still 39.3 px.

It is a seed-and-track stage, not a finished pipeline. The localisation figures are raw neural output with no refinement stage.

Training

inf_s4000 β€” fine-tuned from locotrack_base.ckpt on movi_e/256x256 from gs://kubric-public/tfds (published by Google). Only the 5.8M cross-track parameters were trained. 4000 steps, ~2 s/step at 2.3 GB on an RTX A4000.

c8_kubric β€” continues that line and changes the data. Trained on facebook/CoTracker3_Kubric, the synthetic set Meta released for CoTracker3 stage 1, which is Apache-2.0 β€” note that CoTracker3's own code and weights are CC-BY-NC and are neither used nor derived from here. 1,954 shots of 120 frames at 512Γ—512 with per-frame depth and a ground-truth camera. 20,000 steps at 256Γ—256, 24-frame windows, 256 tracks, cosine schedule, EMA 0.995, temporal mixer unfrozen, L1 on the occluded population. 2.6 h at 3.0 GB on an RTX A4000.

Three properties of that dataset are counter-intuitive and each will silently invert a training run if read the obvious way: the array named visibility holds occlusion (True = hidden), coordinates are (x, y) column-first, and 75% of its "occluded" samples are simply off screen rather than hidden behind anything. All three were established photometrically against the frames rather than assumed.

License

Apache-2.0. See LICENSE and NOTICE, which must travel with any copy.

Cross-track attention and the occluded-position loss weighting are implemented from the description in Karaev et al., arXiv:2410.11831.

Citing

@inproceedings{cho2024locotrack,
  title     = {Local All-Pair Correspondence for Point Tracking},
  author    = {Cho, Seokju and Huang, Jiahui and Nam, Seungryong and
               Min, Dongbo and Lee, Joon-Young},
  booktitle = {ECCV},
  year      = {2024}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for JeffyAntony/Jeff-Tracker