Video generation router (generic)

Given a video generation request — a text prompt for text-to-video (T2V), or a prompt plus a source image for image-to-video (I2V) — this model scores every candidate video generator and returns them ranked best first. Use it to pick which generator to send a request to without generating any videos.

Built on Qwen3-VL-Embedding-2B. Each generator is a special token <|vm:NAME|>; a (request, generator) pair is encoded as one sequence and mapped to a scalar score. Higher is better. Scores are only meaningful relative to other candidates for the same request.

Requirements

pip install "transformers>=5.16" torch pillow

The model uses custom code, so load it with trust_remote_code=True.

Quick start

import torch
from transformers import AutoModel, AutoProcessor

repo = "theblackcat102/video_routing_test_generic"
model = AutoModel.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).cuda().eval()
processor = AutoProcessor.from_pretrained(repo, trust_remote_code=True, padding_side="right")

# Text-to-video: rank all supported generators
ranking = model.route(processor, "A corgi surfing a wave at sunset, slow motion", "t2v")
for name, score in ranking[:5]:
    print(f"{name:32s} {score:+.4f}")

best_generator = ranking[0][0]

Image-to-video

Pass the source image as a path or a PIL.Image:

from PIL import Image

ranking = model.route(
    processor,
    "The camera slowly pushes in as rain starts to fall",
    "i2v",
    image=Image.open("first_frame.png"),
)

Restricting the candidates

Rank only the generators you can actually call (for example, those within your budget). Alias names such as veo-3.1-generate-001 are accepted and mapped to their canonical names:

ranking = model.route(
    processor,
    "A product shot of a perfume bottle rotating on a marble table",
    "t2v",
    candidates=["veo-3.1-fast", "kling-v2.5-turbo-pro", "seedance-2.0-fast"],
)

A generator that is not in the list below raises KeyError.

Supported generators

print(model.config.video_models)   # canonical names
print(model.config.model_aliases)  # alias -> canonical name

flux-3, gemini-omni-flash, grok-imagine, grok-imagine-video-v1.5, grok-imagine-video-v1.5-lite, happy-horse-v1.1, kling-v2.5-turbo-pro, kling-v2.5-turbo-standard, longcat-video-distilled-720p, ltx-2.3, ltx-2.3-fast, ltx-2.5-fast, ltx-video-13b-distilled, luma-ray-v3.2, minimax-h3, minimax-h3-max, minimax-h3-max-turbo, pika-v2.2, pixverse-c1, pixverse-v6, seedance-2.0-fast, seedance-2.0-mini, seedance-2.5-us, veo-3.1, veo-3.1-fast, veo-3.1-lite, vidu-q3-image-to-video-turbo, vidu-q3-text-to-video-turbo, wan-3.0, wan-v2.2-5b, wan-v2.2-a14b

Notes

  • task must be "t2v" or "i2v"; I2V requests need an image, and an image passed with a T2V request is ignored.
  • Prompts longer than 512 tokens are truncated.
  • Input images are resized to at most 256 × 32² pixels, preserving aspect ratio.
  • Each call scores all candidates in one batch (one sequence per candidate).
  • Inference needs about 6.5 GB of GPU memory in bfloat16 when scoring all generators at once.
Downloads last month
21
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for theblackcat102/video_routing_test_generic

Finetuned
(14)
this model

Dataset used to train theblackcat102/video_routing_test_generic