Jeff-Tracker
Track any point through a video.
Give it a video and a set of points; it returns where each point is on every frame, whether it can see it, and how sure it is. It runs at 0.0065 s/frame and returns a calibrated per-frame confidence alongside every track.
Code, benchmarks and method: https://github.com/antonyjaguar19-dotcom/Jeff-Tracker
It is LocoTrack-B with cross-track attention added β the tracks attend to each other, so a point that is hidden can be inferred from the neighbours that are still visible β then fine-tuned for occlusion β on MOVi-E, and in the newer checkpoint on Meta's Apache-2.0 Kubric release.
Access
This repository is gated. Click Request access; the maintainer approves manually. Once approved:
hf auth login
python tools/fetch_weights.py --all
Files
| file | what it is |
|---|---|
c8_kubric.ckpt |
trained on Meta's CoTracker3_Kubric. Wins on the synthetic benches and has not been shown to help on real footage -- read "Which checkpoint" below, 66 MB |
inf_s4000.ckpt |
the original release checkpoint. Step 4000, occ_pos_weight=0.2, 66 MB |
The LocoTrack baseline weights are not mirrored here; they come from the LocoTrack authors'
own release, hamacojr/LocoTrack-pytorch-weights.
Use
import torch
tracker = torch.hub.load("antonyjaguar19-dotcom/Jeff-Tracker", "jefftracker").cuda()
# frames: (T, H, W, 3) uint8 BGR queries: (N, 3) as [frame, x, y]
tracks, vis, conf = tracker.track_queries_conf(frames, queries)
# tracks (T, N, 2) xy, y DOWN vis (T, N) bool conf (T, N) float
Keep conf. At 256Γ256 it detects the model's own >5 px frames at AUC 0.955, and
gating at 0.5 capped the worst sample in a whole run at 9 px.
Results
Two checkpoints, and which one is better depends on what you are tracking. Both numbers below are measured here on the same day with the same binary.
inf_s4000 |
c8_kubric |
|
|---|---|---|
| TAP-Vid DAVIS strided (AJ / Ξ΄_avg / OA) | 62.6 / 74.9 / 86.7 | see below |
| Localisation, exact synthetic ground truth | 1.03 px @ 256Γ256 | ~3% tighter |
| Pixel locking | +0.0005 px β none, on an NCC-free path | same |
| Confidence as a bad-frame detector | AUC 0.955 @ 256Γ256 | same path |
| Speed | 0.0065 s/frame Β· 6.85 GB peak on 4K | same |
Which checkpoint β and why the honest answer is the older one
c8_kubric was compared against c5_occl1 with three late checkpoints per arm, and on
the synthetic occlusion benches it separated better on 5 of 5 for visible accuracy
(3%) and re-acquisition (4.5%).
That reads stronger than it is, and the caveat is the important part: all five of those benches derive from one synthetic shot β a textured flat plane moved by a homography, with cards or depth planes composited over it. Five benches, one plate. A gain there says the model improved on textured planes under clean motion.
Every measurement on real footage says something else:
| test | better |
|---|---|
| synthetic occlusion benches (5, one source plate) | c8_kubric |
| TAP-Vid DAVIS, 30 real clips | c5_occl1 β 68.3 / 79.8 / 89.9 against 67.2 / 79.3 / 88.8 |
| real 4K plate, frames the model will commit to | c5_occl1 β 53.1% against 48.7% |
| real plate against a hand-built reference | tie, inside the reference's own 0.60 px noise |
| real plate, forward-and-back closure | c5_occl1 β mean 0.67 px against 0.72, worst 7.2 against 12.6 |
| a second real plate, closure | mixed β c8_kubric the better tail, c5_occl1 the better median |
So c5_occl1 remains the recommendation. c8_kubric is published because the training
is reproducible and its tail is better on one plate, not because it is an upgrade. If your
footage resembles clean synthetic motion it may help; on handheld plates with grain and
defocus the evidence does not support the swap.
This section is worded the way it is because the first version of this card recommended
c8_kubric on the bench result alone, before the real-footage numbers existed.
Unchanged either way: accuracy while a point is hidden. 2.90 px against CoTracker3's 1.69 on the same bench. Neither checkpoint moved that column β see Limitations.
Limitations β read before deploying
It gaps an occlusion rather than crossing one, and c8_kubric gaps MORE. Willingness to
commit a position on a ground-truth-occluded frame, at confidence gate 0.5:
| occluded frames given a position | |
|---|---|
c5_occl1 |
~5.5% |
c8_kubric |
~3.7% |
The newer checkpoint is more cautious, not more capable. That is the whole explanation of its result: fewer committed guesses give tighter visible tracks and better re-acquisition, and no progress through an occlusion. Whether this is an improvement depends on your pipeline. For matchmove it generally is β a clean hole is visible and fillable, while a confidently wrong position inside an occluder looks correct in the viewport and quietly poisons a camera solve. If your downstream stage requires a position on every frame, it is a regression.
Five separate attempts at the occlusion gap have now been measured and none moved it: more neighbours per chunk (0.002 px at 9Γ the neighbours), neighbours carrying temporal state (a wash), a fourth correlation level (worse on three benches), better training data (this checkpoint β localisation yes, occlusion no), and doubling the training window from 24 to 48 frames (no effect; the coverage figure above was identical at both). Three of those change the architecture and two change what it is fed. Treat the occluded column as a known structural limitation of this architecture rather than a tuning problem.
Use 256Γ256. At 384Γ680 the median improves to 0.697 px while the mean is 40.6 px, and no confidence threshold recovers it β at conf β₯ 0.99 the mean is still 39.3 px.
It is a seed-and-track stage, not a finished pipeline. The localisation figures are raw neural output with no refinement stage.
Training
inf_s4000 β fine-tuned from locotrack_base.ckpt on movi_e/256x256 from
gs://kubric-public/tfds (published by Google). Only the 5.8M cross-track parameters were
trained. 4000 steps, ~2 s/step at 2.3 GB on an RTX A4000.
c8_kubric β continues that line and changes the data. Trained on
facebook/CoTracker3_Kubric,
the synthetic set Meta released for CoTracker3 stage 1, which is Apache-2.0 β note that
CoTracker3's own code and weights are CC-BY-NC and are neither used nor derived from here.
1,954 shots of 120 frames at 512Γ512 with per-frame depth and a ground-truth camera. 20,000
steps at 256Γ256, 24-frame windows, 256 tracks, cosine schedule, EMA 0.995, temporal mixer
unfrozen, L1 on the occluded population. 2.6 h at 3.0 GB on an RTX A4000.
Three properties of that dataset are counter-intuitive and each will silently invert a
training run if read the obvious way: the array named visibility holds occlusion
(True = hidden), coordinates are (x, y) column-first, and 75% of its "occluded" samples
are simply off screen rather than hidden behind anything. All three were established
photometrically against the frames rather than assumed.
License
Apache-2.0. See LICENSE and NOTICE, which must travel with any copy.
Cross-track attention and the occluded-position loss weighting are implemented from the description in Karaev et al., arXiv:2410.11831.
Citing
@inproceedings{cho2024locotrack,
title = {Local All-Pair Correspondence for Point Tracking},
author = {Cho, Seokju and Huang, Jiahui and Nam, Seungryong and
Min, Dongbo and Lee, Joon-Young},
booktitle = {ECCV},
year = {2024}
}