Model Architecture

The Video Tokenizer architecture extends standard image-based VQ-GANs into the time dimension using a 3D Causal architecture combined with Lookup-Free Quantization (LFQ). It consists of three main components: the Encoder, the Quantizer, and the Decoder.

1. 3D Causal Encoder

The encoder projects the input video tensor $X \in \mathbb{R}^{B \times C \times T \times H \times W}$ into a continuous latent space $Z \in \mathbb{R}^{B \times D \times T' \times H' \times W'}$.

  • Spatial Downsampling: standard 2D/3D convolutions reduce the spatial dimensions by a factor of 8 ($H/8$, $W/8$).
  • Temporal Compression: In this specific configuration, temporal downsampling is set to 1x ($T/1$). The model processes video continuously without aggressive time compression to maximize frame-by-frame fidelity.
  • Residual Blocks & Attention: Spatial-temporal ResNet blocks combined with Axial Attention ensure that both local textures and global temporal consistency are captured.

2. Lookup-Free Quantization (LFQ)

The continuous latents are mapped to a discrete codebook without the standard embedding lookup bottleneck.

  • Factorized Codes: For every spatial coordinate (e.g., in a 16x16 grid), the model utilizes multiple codebooks (8 codebook entries per grid location).
  • Latent Shape: Given a $128 \times 128$ video with 16 frames, the latent representation takes the shape [Batch, 16, 16, 16, 8]. This equates to [Batch, Time, Height, Width, Codebooks].
  • Codebook Utilization: LFQ mitigates index collapse naturally without the need for complex commitment loss schedules, maximizing the usage of the latent space.

3. Decoder & Adversarial Training

The decoder reconstructs the video from the quantized latents $\hat{Z}$.

  • Upsampling: Symmetrical 3D causal transposed convolutions upsample the latents back to the original $T \times H \times W$ dimensions.
  • Discriminator: A 3D PatchGAN Discriminator is used during training to enforce perceptual realism, penalizing blurry outputs and temporal flickering.
  • Loss Objectives: L1 + L2 Reconstruction, LPIPS Perceptual Loss, and 3D Adversarial (GAN) Loss.

Videos

video video video

Model code

Look -> https://github.com/firdavsus/Video_Tokenizer Github

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support