Papers
arxiv:2609.33419

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

Published on Sep 27
ยท Submitted by
Shih-Ying Yeh
on Sep 29
ยท nvidia NVIDIA
Authors:
,
,
,
,

Abstract

Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched 4 times 6 = 24 architecture-objective study at roughly 170M ~ 190M encoder scale on sim1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.

Community

Paper author Paper submitter
โ€ข
edited about 3 hours ago

We introduce TT-VidT, a video pretraining framework that treats spatial and temporal separately instead of stacking per-frame embeddings or running one global 3D model: it outputs a per-frame "motion token," and a decoder rebuilds each frame from a key frame plus that token, reaching strong results with only small-scale training. When we flip or time-reverse SSv2 mirror-class videos, TT-VidT keeps its original answer on โ‰ค1% of clips while baselines stay on 9โ€“43%, so it really reads the motion. We think this is a promising direction with lots of room to grow, and we'd love to see you build on it!

project page:
https://kohakublueleaf.github.io/TTVidT/

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33419
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33419 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33419 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33419 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.