Instructions to use KBlueLeaf/TTVidT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KBlueLeaf/TTVidT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("video-classification", model="KBlueLeaf/TTVidT", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from KBlueLeaf/TTVidT: direct link, hf CLI and curl.
- Browser
- Download file 7.26 kB
-
https://huggingface.co/KBlueLeaf/TTVidT/resolve/main/README.md
- Command line
-
hf download hf://KBlueLeaf/TTVidT/README.md
-
curl -L -o README.md https://huggingface.co/KBlueLeaf/TTVidT/resolve/main/README.md
license: apache-2.0
library_name: transformers
pipeline_tag: video-classification
tags:
- video
- video-representation-learning
- self-supervised-learning
- motion
- temporal-modeling
- dinov3
- vision-transformer
- custom_code
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
base_model_relation: finetune
datasets:
- nkp37/OpenVid-1M
metrics:
- accuracy
model-index:
- name: TT-VidT (TT3D)
results:
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: hmdb51
name: HMDB51
metrics:
- type: accuracy
value: 25.2
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: arid
name: ARID
metrics:
- type: accuracy
value: 36.1
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: iard
name: IARD
metrics:
- type: accuracy
value: 86
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: jester
name: Jester
metrics:
- type: accuracy
value: 72.9
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: something-something-v2
name: Something-Something v2
metrics:
- type: accuracy
value: 25.4
name: Top-1 accuracy (mean of 3 seeds)
TT-VidT
Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
NeurIPS 2026 (main track)
TT-VidT is a self-supervised video encoder built for motion. A DINOv3 ViT-B/16 processes every frame independently (the appearance path), while a compact Temporal Transfer pathway turns each frame into a few motion tokens that exchange information across time. This repository holds the pretrained TT3D encoder (195.5M parameters).
Model
- Encoder: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D) layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion tokens together with its 4x-downsampled spatial tokens, and writes the result back to the spatial stream.
- Pretraining objective: Diff Compression. A DiT decoder reconstructs every later frame from the first frame's features plus that frame's motion tokens, so the motion tokens carry what the first frame cannot explain.
| Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) |
| Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining |
| Output | motion_output: one 768-d motion embedding per frame, [B, T, 1, 768] |
| Weights | fp32 safetensors, encoder only |
Quick start
With transformers only (the model code ships in this repository):
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()
video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W] in [-1, 1]
with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
motion = model(video).motion_output # [B, T, 1, 768]
With the TT-VidT codebase (training, evaluation, feature extraction):
from ttvidt.hub import load_model
model = load_model("KBlueLeaf/TTVidT", device="cuda")
motion = model.encoder(video).motion_output
Loading needs no access to the (gated) DINOv3 base weights: every weight is in this repository.
Results
Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds, 8 frames (evaluation protocol of the paper):
| HMDB51 | ARID | IARD | Jester | SSv2 |
|---|---|---|---|---|
| 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |
For the full study (24 architecture–objective pairs at matched scale, fine-tuning and diagnostics), see the paper.
Possible downstream uses
The encoder gives a compact sequence of per-frame motion tokens alongside the DINOv3 appearance features. Some directions we think are worth trying:
- From image models to video models: pair an existing image model with the motion token sequence to get a video model for understanding in the broad sense: classification, retrieval, captioning, question answering, or any other task.
- Generation and motion transfer: use the motion tokens as a conditioning signal for video generation, or take them from one clip and apply them to another subject or scene.
Further exploration and feedback are very welcome, and so are attempts at larger scale (bigger backbones, more data, longer training). Please open an issue or a discussion on GitHub or in the Community tab.
Related resources
| Resource | Link |
|---|---|
| Project page | kohakublueleaf.github.io/TTVidT |
| Paper | huggingface.co/papers/2609.33419 · arXiv:2609.33419 |
| Source code (training, evaluation, all paper configs) | github.com/KohakuBlueleaf/TTVidT |
| Pretrained DiT decoders (Diff Compression and the other objectives) | KBlueLeaf/TTVidT-decoders |
Files
| File | Content |
|---|---|
config.json |
architecture and loader configuration (transformers + TT-VidT codebase) |
model.safetensors |
fp32 encoder weights |
*.py |
encoder code for trust_remote_code (needs only torch and transformers) |
assets/ |
figures of this card |
License
Apache-2.0. The spatial path is initialised from DINOv3, which is released under its own license.

