File size: 7,256 Bytes
742c169
 
 
 
 
 
c9eef3f
 
742c169
c9eef3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
742c169
 
c9eef3f
742c169
c9eef3f
742c169
c9eef3f
742c169
c9eef3f
742c169
c9eef3f
 
 
 
 
 
 
 
 
3c7a789
 
 
 
c9eef3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
742c169
 
 
 
 
dd4e57f
c9eef3f
 
742c169
c9eef3f
742c169
 
c9eef3f
 
742c169
 
 
 
dd4e57f
742c169
 
 
c9eef3f
 
742c169
 
 
c9eef3f
 
 
 
 
 
742c169
c9eef3f
 
742c169
24879fd
 
 
 
 
 
c9eef3f
 
24879fd
c9eef3f
 
 
 
 
 
 
 
24879fd
c9eef3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24879fd
742c169
 
c9eef3f
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
---
license: apache-2.0
library_name: transformers
pipeline_tag: video-classification
tags:
  - video
  - video-representation-learning
  - self-supervised-learning
  - motion
  - temporal-modeling
  - dinov3
  - vision-transformer
  - custom_code
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
base_model_relation: finetune
datasets:
  - nkp37/OpenVid-1M
metrics:
  - accuracy
model-index:
  - name: TT-VidT (TT3D)
    results:
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: hmdb51
          name: HMDB51
        metrics:
          - type: accuracy
            value: 25.2
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: arid
          name: ARID
        metrics:
          - type: accuracy
            value: 36.1
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: iard
          name: IARD
        metrics:
          - type: accuracy
            value: 86.0
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: jester
          name: Jester
        metrics:
          - type: accuracy
            value: 72.9
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: something-something-v2
          name: Something-Something v2
        metrics:
          - type: accuracy
            value: 25.4
            name: Top-1 accuracy (mean of 3 seeds)
---

<div align="center">

# TT-VidT

### Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

**NeurIPS 2026 (main track)**

[![Project Page](https://img.shields.io/badge/Project-Page-green)](https://kohakublueleaf.github.io/TTVidT/)
[![Paper](https://img.shields.io/badge/🤗%20Paper-2609.33419-yellow)](https://huggingface.co/papers/2609.33419)
[![arXiv](https://img.shields.io/badge/arXiv-2609.33419-b31b1b)](https://arxiv.org/abs/2609.33419)
[![Code](https://img.shields.io/badge/GitHub-KohakuBlueleaf%2FTTVidT-181717?logo=github)](https://github.com/KohakuBlueleaf/TTVidT)
[![Decoders](https://img.shields.io/badge/🤗%20Decoders-TTVidT--decoders-orange)](https://huggingface.co/KBlueLeaf/TTVidT-decoders)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue)](#license)

</div>

<!-- ![TT-VidT overview](assets/overview.png) -->


![image](https://cdn-uploads.huggingface.co/production/uploads/630593e2fca1d8d92b81d2a1/JKhgNd-K2wnoaAp1fYwR2.png)

**TT-VidT** is a self-supervised video encoder built for *motion*. A DINOv3 ViT-B/16
processes every frame independently (the appearance path), while a compact
**Temporal Transfer** pathway turns each frame into a few motion tokens that
exchange information across time. This repository holds the pretrained **TT3D**
encoder (195.5M parameters).

## Model

![TT-VidT architecture](assets/architecture.png)

- **Encoder**: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D)
  layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion
  tokens together with its 4x-downsampled spatial tokens, and writes the result back
  to the spatial stream.
- **Pretraining objective**: Diff Compression. A DiT decoder reconstructs every
  later frame from the *first* frame's features plus that frame's motion tokens, so
  the motion tokens carry what the first frame cannot explain.

| | |
|---|---|
| Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) |
| Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining |
| Output | `motion_output`: one 768-d motion embedding per frame, `[B, T, 1, 768]` |
| Weights | fp32 `safetensors`, encoder only |

## Quick start

**With `transformers` only** (the model code ships in this repository):

```python
import torch
from transformers import AutoModel

model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()

video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1    # [B, T, C, H, W] in [-1, 1]
with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
    motion = model(video).motion_output                           # [B, T, 1, 768]
```

**With the [TT-VidT codebase](https://github.com/KohakuBlueleaf/TTVidT)** (training,
evaluation, feature extraction):

```python
from ttvidt.hub import load_model

model = load_model("KBlueLeaf/TTVidT", device="cuda")
motion = model.encoder(video).motion_output
```

Loading needs no access to the (gated) DINOv3 base weights: every weight is in this
repository.

## Results

Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds,
8 frames (evaluation protocol of the paper):

| HMDB51 | ARID | IARD | Jester | SSv2 |
|:---:|:---:|:---:|:---:|:---:|
| 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |

For the full study (24 architecture–objective pairs at matched scale, fine-tuning and
diagnostics), see the [paper](https://huggingface.co/papers/2609.33419).

## Possible downstream uses

The encoder gives a compact sequence of per-frame motion tokens alongside the
DINOv3 appearance features. Some directions we think are worth trying:

- **From image models to video models**: pair an existing image model with the
  motion token sequence to get a video model for understanding in the broad sense:
  classification, retrieval, captioning, question answering, or any other task.
- **Generation and motion transfer**: use the motion tokens as a conditioning
  signal for video generation, or take them from one clip and apply them to another
  subject or scene.

> [!TIP]
> Further exploration and feedback are very welcome, and so are attempts at larger
> scale (bigger backbones, more data, longer training). Please open an issue or a
> discussion on [GitHub](https://github.com/KohakuBlueleaf/TTVidT) or in the
> Community tab.

## Related resources

| Resource | Link |
|---|---|
| Project page | [kohakublueleaf.github.io/TTVidT](https://kohakublueleaf.github.io/TTVidT/) |
| Paper | [huggingface.co/papers/2609.33419](https://huggingface.co/papers/2609.33419) · [arXiv:2609.33419](https://arxiv.org/abs/2609.33419) |
| Source code (training, evaluation, all paper configs) | [github.com/KohakuBlueleaf/TTVidT](https://github.com/KohakuBlueleaf/TTVidT) |
| Pretrained DiT decoders (Diff Compression and the other objectives) | [KBlueLeaf/TTVidT-decoders](https://huggingface.co/KBlueLeaf/TTVidT-decoders) |

## Files

| File | Content |
|---|---|
| `config.json` | architecture and loader configuration (`transformers` + TT-VidT codebase) |
| `model.safetensors` | fp32 encoder weights |
| `*.py` | encoder code for `trust_remote_code` (needs only `torch` and `transformers`) |
| `assets/` | figures of this card |

## License

Apache-2.0. The spatial path is initialised from
[DINOv3](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m), which is
released under its own license.