TGIF text-to-video research model
Current Qwen-conditioned VideoCrafter2 package
The qwen-videocrafter2/ package contains a
58.4M-parameter adapter connecting Qwen3.5-0.8B to a pretrained VideoCrafter2
video U-Net. Qwen is the runtime prompt encoder. Qwen, the U-Net and VAE remain
frozen; this release contains no fine-tuned video-denoiser LoRA.
The adapter completed 5,000 broader-caption updates after an earlier 5,000-update stage. The broader stage uses 90,000 caption examples and 3,000 validation texts, including authored compositional texts without paired videos. The package records checkpoint ancestry, exact settings and data counts.
In subjective assistant development reviews, 20/24 subjects and 13/24 actions were clear in one prompt bank. A second bank using new wording produced 44/48 clear subjects and 15/48 clear actions across two seeds. These are small development checks using known vocabulary, not independent benchmark scores. Fine actions, hands, counts, object contact and anatomical consistency remain unreliable. All reviewed prompts and failures are included in the package evaluation.
hf download virtualkevin/tgif-video-gen --include 'qwen-videocrafter2/*' --local-dir tgif-video-gen
cd tgif-video-gen/qwen-videocrafter2
uv sync --locked
uv run python -m scripts.download_qwen_video_assets
uv run python -m scripts.sample_qwen_videocrafter --checkpoint adapter.pt --prompt "A turtle crawls across a sandy path." --seed 31234 --output turtle.gif
Defaults are 512×320, 16 frames at 8 fps, BF16 denoising, 25 DDIM steps and guidance 12. Grayscale conversion follows RGB generation. Qwen accepts up to 1,024 tokens; long-prompt reliability has not been established. The standalone package passed a clean dependency installation and actual GPU CLI generation; published files were downloaded at their immutable commit and verified by SHA256.
Inference weights, full adapter training state, runnable code and a dependency lock are included. Pretrained dependencies are downloaded separately at pinned revisions. Source GIFs, feature caches and credentials are excluded. Exact training resumption requires the original matching private cache contract. VideoCrafter's upstream research/non-commercial license restrictions continue to apply; see the package licenses and README.
Final audit of the retained model (2026-10-04)
The retained checkpoint is still broad5000. Two later v5 LoRA experiments (temporal attention, and temporal plus spatial text cross-attention) each completed 1,000 updates, but neither convincingly improved action following in the development comparison. Neither adapter was promoted or included in this release.
After that decision was frozen, the retained model generated 24 previously unused final prompt wordings at two seeds. Every frame of all 48 clips was reviewed:
| Seed | Clear subjects | Clear requested actions | Both clear |
|---|---|---|---|
| 91234 | 18/24 | 6/24 | 6/24 |
| 91235 | 16/24 | 2/24 | 2/24 |
Only 2/24 prompts had both a clear subject and action at both seeds: a horse walking through shallow water, and a child waving both hands. The final prompts include fine actions and event transitions; these results are not directly comparable to the different development banks above. Subjects are often recognizable, but dependable action control has not been achieved.
A separate control used six still/action caption pairs at two further seeds, with identical initial noise and sampling RNG within each pair. Direct visual review found a clear requested behavioral contrast in 0/6 pairs at seed 91236 and 1/6 at seed 91237; partial contrast occurred in 3/6 pairs at each seed. The remaining pairs did not demonstrate the requested contrast. Matching noise does not ensure identical scene composition after changing the caption.
These are subjective assistant judgments, not a human benchmark. One assistant reviewed each seed; reviewer and seed effects are therefore confounded. Prompt concepts overlap with earlier development vocabulary. All final clips use the unchanged 512×320, 16-frame, 8-fps, 25-step DDIM, CFG-12 profile. No model selection, training or tuning used these final outputs. The complete audit report preserves every caption, rating, frame-linked observation, failure and uncertainty. The protocol was frozen before final-bank inspection. This audit evaluates only the retained model and does not establish a v5 improvement.
Earlier Qwen-conditioned AnimateDiff package
The qwen-animatediff/ folder contains an interim,
GPU-tested improvement: a trained Qwen3.5-0.8B conditioning adapter driving a
pretrained SD1.5 U-Net with AnimateDiff motion modules. Qwen is the only runtime
text encoder. CLIP supplies training targets and a fixed empty-prompt constant.
The U-Net, motion modules, VAE, and Qwen remain frozen in this training stage.
This earlier checkpoint trains a 58.2M-parameter adapter for 5,000 updates on 70,000 texts (50,000 source captions and 20,000 authored compositional texts). On a fresh 24-prompt development set at one fixed seed, 19 subjects and 10 actions were judged clear. The preceding checkpoint had 8 clear subjects on the same prompts and seed. This qualitative comparison changes model capacity, training data and update count together; it is not an unbiased benchmark or a test of one isolated change. Identity/count changes, anatomical defects, and missed fine actions remain. Broader experiments are continuing. Default output is a 512×512, 16-frame grayscale GIF; grayscale is an output conversion of the RGB generator.
Download only the improved package, then follow its pinned dependency setup:
hf download virtualkevin/tgif-video-gen --include 'qwen-animatediff/*' --local-dir tgif-video-gen
cd tgif-video-gen/qwen-animatediff
uv sync --locked
uv run python -m scripts.download_qwen_video_assets
uv run python -m tinygif.qwen_video_sample --checkpoint adapter.pt --prompt "A dog runs across a field." --output dog.gif
The folder includes inference weights, resumable training state, runnable code, checksums, the qualitative evaluation, and evidence from a standalone GPU smoke test. Pretrained dependencies are downloaded separately at pinned revisions. Source GIFs, feature caches, and credentials are not included.
Original 32×32 research checkpoint
The files at the repository root retain the earlier model described below.
Checkpoint at 10,000 optimizer updates. Training stage: full-data text-conditioning training. Frozen Qwen3.5-0.8B caption features condition a trainable spatiotemporal U-Net. Generates eight 32×32 grayscale frames at 8 fps by default. This early research checkpoint is not a claim of reliable prompt following or production video quality.
model.pt contains EMA inference weights; training-state.pt retains model,
EMA, optimizer and per-rank state for resuming the original distributed run.
Both files represent this exact step. Previous published versions remain in
Hub commit history. Do not resume training from the inference-only file.
The U-Net is initialized from our retained unconditional TGIF model. It uses pixel-space velocity prediction, caption dropout, and classifier-free guidance. Captions describe whole GIFs while the cached training clips are short temporal crops: these are weak text labels, not precisely aligned event supervision. Existing source-group train/validation/test splits are preserved. The machine captions contain errors. No source GIF media is included in this repository.
Generate from a prompt
Download this repository to a local directory, then run from that directory:
uv sync --locked
uv run python -m tinygif.text_sample --checkpoint model.pt --prompt "A person waves a hand at the camera." --output samples/waving --num-samples 4 --steps 100 --guidance-scale 3
The first prompt loads the frozen Qwen encoder at revision
2fc06364715b967f1860aea9cf38778875588b17. Its weights are downloaded separately. The encoder
settings are recorded in config.json; generation uses the same normalization,
pooling and token limit as training. GPU execution and Python 3.12 are expected.
This is a custom PyTorch model, not a Diffusers pipeline or hosted inference API.
Larger guidance values are not a guarantee of better motion or prompt accuracy.
Provenance
Training repository: virtualkevin/tgif-video-gen. Original GIFs come from TGIF; follow TGIF's
dataset usage terms. Qwen retains its upstream license. The vendored U-Net's MIT license is included under tinygif/vendor.
Artifact checksums are in
manifest.json. Training precision and distributed settings are in config.json.
- Downloads last month
- 139