SemanTok / README.md
mishok43's picture
card: combined teaser GIF, Community License text
d1b5e9d verified
|
Raw History Blame Contribute Delete
6.57 kB
---
license: other
pipeline_tag: text-to-video
tags:
- video-tokenizer
- text-to-video
- class-to-video
- autoregressive
- checkpoint
---
# SemanTok
Predictable Semantic Tokens for Efficient Autoregressive Video Generation
<p align="center"><img src="./assets/teaser.gif" width="100%" alt="Text-to-video on uCO3D and class-to-video on Kinetics-600 at token budgets k=4 to 256, VideoFlexTok vs. SemanTok"></p>
SemanTok is a flexible (coarse-to-fine) video tokenizer whose first tokens carry the clip's semantics.
Every nested token prefix is trained to reconstruct DINOv2 features, so an autoregressive (AR) model
can stop after any number of tokens *k* per frame and still produce a semantically faithful video. A
201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4× its size.
This repository holds the paper's tokenizers and AR models for both SemanTok and the VideoFlexTok
baseline, on Kinetics-600 (class-to-video) and uCO3D (text-to-video).
![SemanTok overview: the tokenizer (left) and AR generation (right)](./assets/overview.png)
Please note: For individuals or organizations generating annual revenue of US $1,000,000 (or local currency equivalent) or more, regardless of the source of that revenue, you must obtain an enterprise commercial license directly from Stability AI before commercially using SemanTok, derivative works of SemanTok, or outputs from SemanTok. See https://stability.ai/license and https://stability.ai/enterprise.
## Model Description
- Developed by: [Stability AI](https://stability.ai/)
- Authors: Mikhail Dereviannykh, Vikram Voleti, Simon Donné, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss
- Model type: Flexible-length video tokenizer and class- / text-conditioned autoregressive video generation models
- Model details: the tokenizer encodes each of 5 latent frames (17 frames, 128×128) into 256 ordered FSQ tokens (64k codes); any prefix of *k* tokens per frame decodes to a video through a rectified-flow decoder. SemanTok adds DINOv2 features to the encoder input, the DINO class token to the first register token, and dense and class DINO heads on every kept prefix during training. The AR models are LLaMA-style transformers that predict tokens coarse to fine up to the chosen budget *k*.
- Base stack: [VideoFlexTok](https://github.com/apple/ml-videoflextok) (tokenizer architecture), [VidTok](https://github.com/microsoft/VidTok) VAE, [DINOv2](https://github.com/facebookresearch/dinov2)-L, umT5 text encoder (via Wan 2.1) for uCO3D captions
## Model Sources
- Repository: https://github.com/Stability-AI/SemanTok
- Project page: https://semantoken.github.io/
- Paper: coming soon
For installation, evaluation protocols, and data preparation, please use the [GitHub repository](https://github.com/Stability-AI/SemanTok).
## Usage
```bash
git clone https://github.com/Stability-AI/SemanTok && cd SemanTok
pip install -e . # Python >= 3.10, CUDA GPU
huggingface-cli download StabilityLabs/SemanTok --local-dir ckpts
CKPT=ckpts
# Reconstruct a clip from its first k tokens per frame
python scripts/reconstruct.py --ckpt-root $CKPT --tokenizer k600-semantok --video clip.mp4 --ks 1 4 16 64 256
# Class-to-video (Kinetics-600)
python scripts/generate.py --ckpt-root $CKPT --ar k600-semantok-d16 --class-name "yoga" --ks 4 16 64 256
# Text-to-video (uCO3D)
python scripts/generate.py --ckpt-root $CKPT --ar uco3d-semantok-d16 \
--prompt "A small orange basketball on a plaid tablecloth" --ks 4 16 64
```
From Python:
```python
from semantok import Tokenizer
from semantok.data.video import load_kinetics_clip
tok = Tokenizer(f"{CKPT}/tokenizers/k600-semantok")
clip = load_kinetics_clip("clip.mp4")[None] # [1, 3, 17, 128, 128] in [-1, 1]
tokens = tok.encode(clip) # [1, 5, 256] FSQ ids, coarse to fine
video = tok.decode(tokens, k=16) # decode from the first 16 tokens per frame
```
## Files
Every model is a directory with `config.json` and `model.safetensors` (bf16):
```
tokenizers/{k600,uco3d}-{videoflextok,semantok}/
ar/{k600,uco3d}-{videoflextok,semantok}-d{10,12,16,20,24}/
```
| tokenizer | data | training |
|---|---|---|
| `k600-videoflextok`, `k600-semantok` | Kinetics-600 | 200k steps (131B tokens) |
| `uco3d-videoflextok`, `uco3d-semantok` | uCO3D | 100k steps (66B tokens) |
| AR depth | d10 | d12 | d16 | d20 | d24 |
|---|---|---|---|---|---|
| parameters | 49M | 85M | 201M | 393M | 679M |
- Kinetics-600 AR models are class-conditioned (597 classes); uCO3D AR models are conditioned on umT5 caption embeddings.
- All AR models here were trained for 20k steps. The paper's larger d30 (1.33B) and d36 (2.29B) models are not included.
- Weights are stored in bf16; the paper evaluated the same weights in fp32 under bf16 autocast.
- The tokenizers do not include the VidTok VAE when it is identical to the public one; the loader fetches it from [`EPFL-VILAB/videoflextok_d18_d18_k600`](https://huggingface.co/EPFL-VILAB/videoflextok_d18_d18_k600).
- `SHA256SUMS` lists the checksum of every `model.safetensors`.
## License
- Community License: Free for research, non-commercial, and commercial use by organizations and individuals generating annual revenue of US $1,000,000 (or local currency equivalent) or less, regardless of the source of that revenue.
- If your annual revenue exceeds US $1M, any commercial use of this model or derivative works requires an Enterprise License directly from Stability AI.
The full text is in [LICENSE.md](./LICENSE.md); see also https://stability.ai/license.
## Intended Uses
Intended uses include research on video tokenization and autoregressive video generation, in particular on generation at a variable token budget, and reproduction of the paper's evaluations.
## Out-of-Scope Uses
SemanTok is not intended to generate factual representations of people, events, products, or places. Use must comply with Stability AI's Acceptable Use Policy and license terms.
## Citation
```bibtex
@misc{dereviannykh2026semantok,
title = {SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation},
author = {Dereviannykh, Mikhail and Voleti, Vikram and Donn{\'e}, Simon and
Reddy, Mallikarjun Byrasandra Ramalinga and Vainer, Shimon and Boss, Mark},
year = {2026}
}
```
## Safety
Users should evaluate generated outputs for their own application constraints and apply additional mitigations where needed. Report safety issues to safety@stability.ai and security issues to security@stability.ai.