Video Classification
Transformers
Safetensors
ttvidt
feature-extraction
video
video-representation-learning
self-supervised-learning
motion
temporal-modeling
dinov3
vision-transformer
custom_code
Eval Results (legacy)
Instructions to use KBlueLeaf/TTVidT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KBlueLeaf/TTVidT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("video-classification", model="KBlueLeaf/TTVidT", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 7,256 Bytes
742c169 c9eef3f 742c169 c9eef3f 742c169 c9eef3f 742c169 c9eef3f 742c169 c9eef3f 742c169 c9eef3f 742c169 c9eef3f 3c7a789 c9eef3f 742c169 dd4e57f c9eef3f 742c169 c9eef3f 742c169 c9eef3f 742c169 dd4e57f 742c169 c9eef3f 742c169 c9eef3f 742c169 c9eef3f 742c169 24879fd c9eef3f 24879fd c9eef3f 24879fd c9eef3f 24879fd 742c169 c9eef3f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 | ---
license: apache-2.0
library_name: transformers
pipeline_tag: video-classification
tags:
- video
- video-representation-learning
- self-supervised-learning
- motion
- temporal-modeling
- dinov3
- vision-transformer
- custom_code
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
base_model_relation: finetune
datasets:
- nkp37/OpenVid-1M
metrics:
- accuracy
model-index:
- name: TT-VidT (TT3D)
results:
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: hmdb51
name: HMDB51
metrics:
- type: accuracy
value: 25.2
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: arid
name: ARID
metrics:
- type: accuracy
value: 36.1
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: iard
name: IARD
metrics:
- type: accuracy
value: 86.0
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: jester
name: Jester
metrics:
- type: accuracy
value: 72.9
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: something-something-v2
name: Something-Something v2
metrics:
- type: accuracy
value: 25.4
name: Top-1 accuracy (mean of 3 seeds)
---
<div align="center">
# TT-VidT
### Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
**NeurIPS 2026 (main track)**
[](https://kohakublueleaf.github.io/TTVidT/)
[](https://huggingface.co/papers/2609.33419)
[](https://arxiv.org/abs/2609.33419)
[](https://github.com/KohakuBlueleaf/TTVidT)
[](https://huggingface.co/KBlueLeaf/TTVidT-decoders)
[](#license)
</div>
<!--  -->

**TT-VidT** is a self-supervised video encoder built for *motion*. A DINOv3 ViT-B/16
processes every frame independently (the appearance path), while a compact
**Temporal Transfer** pathway turns each frame into a few motion tokens that
exchange information across time. This repository holds the pretrained **TT3D**
encoder (195.5M parameters).
## Model

- **Encoder**: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D)
layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion
tokens together with its 4x-downsampled spatial tokens, and writes the result back
to the spatial stream.
- **Pretraining objective**: Diff Compression. A DiT decoder reconstructs every
later frame from the *first* frame's features plus that frame's motion tokens, so
the motion tokens carry what the first frame cannot explain.
| | |
|---|---|
| Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) |
| Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining |
| Output | `motion_output`: one 768-d motion embedding per frame, `[B, T, 1, 768]` |
| Weights | fp32 `safetensors`, encoder only |
## Quick start
**With `transformers` only** (the model code ships in this repository):
```python
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()
video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W] in [-1, 1]
with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
motion = model(video).motion_output # [B, T, 1, 768]
```
**With the [TT-VidT codebase](https://github.com/KohakuBlueleaf/TTVidT)** (training,
evaluation, feature extraction):
```python
from ttvidt.hub import load_model
model = load_model("KBlueLeaf/TTVidT", device="cuda")
motion = model.encoder(video).motion_output
```
Loading needs no access to the (gated) DINOv3 base weights: every weight is in this
repository.
## Results
Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds,
8 frames (evaluation protocol of the paper):
| HMDB51 | ARID | IARD | Jester | SSv2 |
|:---:|:---:|:---:|:---:|:---:|
| 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |
For the full study (24 architecture–objective pairs at matched scale, fine-tuning and
diagnostics), see the [paper](https://huggingface.co/papers/2609.33419).
## Possible downstream uses
The encoder gives a compact sequence of per-frame motion tokens alongside the
DINOv3 appearance features. Some directions we think are worth trying:
- **From image models to video models**: pair an existing image model with the
motion token sequence to get a video model for understanding in the broad sense:
classification, retrieval, captioning, question answering, or any other task.
- **Generation and motion transfer**: use the motion tokens as a conditioning
signal for video generation, or take them from one clip and apply them to another
subject or scene.
> [!TIP]
> Further exploration and feedback are very welcome, and so are attempts at larger
> scale (bigger backbones, more data, longer training). Please open an issue or a
> discussion on [GitHub](https://github.com/KohakuBlueleaf/TTVidT) or in the
> Community tab.
## Related resources
| Resource | Link |
|---|---|
| Project page | [kohakublueleaf.github.io/TTVidT](https://kohakublueleaf.github.io/TTVidT/) |
| Paper | [huggingface.co/papers/2609.33419](https://huggingface.co/papers/2609.33419) · [arXiv:2609.33419](https://arxiv.org/abs/2609.33419) |
| Source code (training, evaluation, all paper configs) | [github.com/KohakuBlueleaf/TTVidT](https://github.com/KohakuBlueleaf/TTVidT) |
| Pretrained DiT decoders (Diff Compression and the other objectives) | [KBlueLeaf/TTVidT-decoders](https://huggingface.co/KBlueLeaf/TTVidT-decoders) |
## Files
| File | Content |
|---|---|
| `config.json` | architecture and loader configuration (`transformers` + TT-VidT codebase) |
| `model.safetensors` | fp32 encoder weights |
| `*.py` | encoder code for `trust_remote_code` (needs only `torch` and `transformers`) |
| `assets/` | figures of this card |
## License
Apache-2.0. The spatial path is initialised from
[DINOv3](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m), which is
released under its own license.
|