Text-to-Video
Diffusers
Safetensors
WanPipeline
video-generation
wan2.1
dmd
quantization-aware-training
nvfp4
sparse-attention
Instructions to use memset0/vsqa-preview-14b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use memset0/vsqa-preview-14b with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("memset0/vsqa-preview-14b", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Download source.yaml from memset0/vsqa-preview-14b: direct link, hf CLI and curl.
- Browser
- Download file 3.47 kB
-
https://huggingface.co/memset0/vsqa-preview-14b/resolve/main/source.yaml
- Command line
-
hf download hf://memset0/vsqa-preview-14b/source.yaml
-
curl -L -o source.yaml https://huggingface.co/memset0/vsqa-preview-14b/resolve/main/source.yaml
3.47 kB
| models: | |
| student: | |
| _target_: fastvideo.train.models.wan.WanModel | |
| init_from: /mnt/lustre/vlm-miz055/vsqa-14b/runs/fastvideo-v0361-14b-c256-ode500-260922-123321/export-step500 | |
| trainable: true | |
| attention_backend: VSA_QAT_TRAIN_C256 | |
| enable_gradient_checkpointing_type: full | |
| flow_shift: 8.0 | |
| teacher: | |
| _target_: fastvideo.train.models.wan.WanModel | |
| init_from: Wan-AI/Wan2.1-T2V-14B-Diffusers | |
| trainable: false | |
| attention_backend: TORCH_SDPA | |
| disable_custom_init_weights: true | |
| flow_shift: 8.0 | |
| critic: | |
| _target_: fastvideo.train.models.wan.WanModel | |
| init_from: Wan-AI/Wan2.1-T2V-14B-Diffusers | |
| trainable: true | |
| attention_backend: TORCH_SDPA | |
| disable_custom_init_weights: true | |
| enable_gradient_checkpointing_type: full | |
| flow_shift: 8.0 | |
| method: | |
| _target_: fastvideo.train.methods.distribution_matching.dmd2.DMD2Method | |
| rollout_mode: simulate | |
| generator_update_interval: 1 | |
| real_score_guidance_scale: 6.0 | |
| dmd_denoising_steps: | |
| - 1000.0 | |
| - 941.1763916015625 | |
| - 800.0 | |
| min_timestep_ratio: 0.02 | |
| max_timestep_ratio: 0.98 | |
| fake_score_learning_rate: 4.0e-07 | |
| fake_score_betas: | |
| - 0.0 | |
| - 0.999 | |
| fake_score_lr_scheduler: constant | |
| shifted_ode_schedule: | |
| version: 1 | |
| teacher_steps: 48 | |
| flow_shift: 8.0 | |
| student_steps: 3 | |
| indices: | |
| - 0 | |
| - 16 | |
| - 32 | |
| timesteps: | |
| - 1000.0 | |
| - 941.1763916015625 | |
| - 800.0 | |
| sigmas: | |
| - 1.0 | |
| - 0.9411764144897461 | |
| - 0.800000011920929 | |
| warp_denoising_step: false | |
| training: | |
| distributed: | |
| num_gpus: 16 | |
| sp_size: 4 | |
| tp_size: 1 | |
| hsdp_replicate_dim: 1 | |
| hsdp_shard_dim: 16 | |
| data: | |
| data_path: /mnt/lustre/vlm-miz055/vsqa-14b/data/vidprom-14b-conditioning-full-canonical-r38ec498/training_dataset | |
| dataloader_num_workers: 0 | |
| train_batch_size: 1 | |
| training_cfg_rate: 0.0 | |
| seed: 1000 | |
| num_latent_t: 20 | |
| num_height: 768 | |
| num_width: 1280 | |
| num_frames: 77 | |
| preprocessed_data_type: text_only | |
| optimizer: | |
| learning_rate: 2.0e-06 | |
| betas: | |
| - 0.0 | |
| - 0.999 | |
| weight_decay: 0.01 | |
| lr_scheduler: constant | |
| lr_warmup_steps: 0 | |
| loop: | |
| max_train_steps: 1000 | |
| gradient_accumulation_steps: 4 | |
| checkpoint: | |
| output_dir: /mnt/lustre/vlm-miz055/vsqa-14b/runs/fastvideo-v0362-14b-c256-dmd1000-260922-123321/checkpoints | |
| training_state_checkpointing_steps: 50 | |
| checkpoints_total_limit: 0 | |
| resume_from_checkpoint: '' | |
| tracker: | |
| trackers: | |
| - wandb | |
| project_name: fastvideo | |
| run_name: fastvideo-v0362-14b-c256-dmd1000-260922-123321 | |
| entity: alexzms-ucsd | |
| vsa: | |
| sparsity: 0.9 | |
| cache_tile_buf: false | |
| dit_precision: fp32 | |
| callbacks: | |
| grad_clip: | |
| max_grad_norm: 1.0 | |
| validation: | |
| pipeline_target: fastvideo.pipelines.basic.wan.wan_shifted_dmd_pipeline.WanShiftedDMDPipeline | |
| dataset_file: /mnt/lustre/vlm-miz055/vsqa-14b/runs/fastvideo-v0362-14b-c256-dmd1000-260922-123321/config/validation_16.json | |
| every_steps: 50 | |
| run_at_start: false | |
| sampling_steps: | |
| - 3 | |
| guidance_scale: 1.0 | |
| height: 768 | |
| width: 1280 | |
| num_frames: 77 | |
| attn_qat_infer: false | |
| _target_: fastvideo.train.methods.ode_init.dmd_validation.ShiftedDMDValidationCallback | |
| use_validation_media_conditioning: false | |
| offload_training_state: true | |
| unload_pipeline_after_validation: true | |
| pipeline: | |
| flow_shift: 8.0 | |
| dit_config: | |
| quant_config: nvfp4_qat_train | |
| dmd_denoising_steps: | |
| - 1000.0 | |
| - 941.1763916015625 | |
| - 800.0 | |
| warp_denoising_step: false | |