Video-4M / scripts /README_EXAMPLES.md
Muhammad Uzair Khattak
Claude Opus 5.5
Rename A2A-Video to Video-4M
5cc082e
|
Raw History Blame Contribute Delete
8.18 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade

Adding / refreshing curated examples for the Any-to-Any tab

Phases 1-3 curate the example inputs; phase 4 pre-generates the model outputs the demo shows without a GPU. Run them in order, all from inside hf_space_demo/ on the cluster (where both the tokenized dataset and the checkpoints actually live).

Phase 1 -- pick clips + upload tokens

python scripts/upload_examples_to_hub.py \
    --source_dir /datasets/uzair/weights_from_clariden/test_cvpr_final_set_13_mod \
    --repo_id EPFL-VILAB/Video-4M-examples \
    --num_examples 6 --seed 0

This randomly picks --num_examples clips that have a file in every modality subfolder, uploads all of their tokenized data to the EPFL-VILAB/Video-4M-examples dataset repo, and saves the picked clip IDs locally to examples_manifest.json (in whatever directory you ran the command from -- keep this file, phases 2 and 3 both need it).

Phase 2 -- render human-viewable previews

cd ..   # repo root, since the script reads hf_space_demo/examples_manifest.json
python visualize_multimodal_pretraining_data_13_modalities.py \
    --config cfgs/default/4m/models/main/4m_L_video_all_modalities_final_config.yaml

This reads hf_space_demo/examples_manifest.json and renders a .mp4 preview per modality for exactly those clips (not a random sample), saved under args.output_dir_videos (default: /datasets/uzair/weights_from_clariden/cvpr_generations/GT_visualizations_post_neurips_opticalflow_fixed), in <domain>_detokenized/<clip_basename>.mp4 subfolders.

Phase 3 -- upload the previews

cd hf_space_demo
python scripts/upload_examples_to_hub.py \
    --repo_id EPFL-VILAB/Video-4M-examples \
    --detokenized_dir /datasets/uzair/weights_from_clariden/cvpr_generations/GT_visualizations_post_neurips_opticalflow_fixed \
    --previews_only

--previews_only skips re-uploading the (large, slow) tokens. Without --force_repick, this automatically reuses the exact clips already recorded in examples_manifest.json -- it will not pick a new random set.

Once this is done, the Any-to-Any tab's example/input-modality dropdowns will show a live preview before you even hit Generate.

Phase 4 -- precompute the demo's default results

Phases 1-3 give the demo example inputs. This phase generates the outputs for the default configuration of each tab, so the Space can open on real predictions with no GPU at all.

Why it matters: ZeroGPU gives free-tier visitors a few minutes of GPU per day, and load_pipeline's ~37.5 GB checkpoint load happens inside that allocation -- so a single click can burn the whole budget and a visitor never gets to see both tabs. With this in place the page opens on finished predictions, and clicking Generate without changing anything is free. Anything else (a different clip, a custom chain, an upload, a typed caption) still runs live on the GPU.

source <(grep '^export' run_demo.sh)   # FOURM_LOCAL_* checkpoint paths
export HF_TOKEN=...

# 1. inspect the two configurations -- loads no checkpoints, safe on a login node
python scripts/precompute_demo_outputs.py --dry_run

# 2. the real run: ~11 mp4s, ~10 min of GPU including the one-time model load
python scripts/precompute_demo_outputs.py --out_dir /tmp/precomputed

# 3. optional local check with no Hub round trip -- copy /tmp/precomputed to
#    <examples dir>/precomputed/ first
FOURM_LOCAL_EXAMPLES_DIR=<examples dir> python app.py

# 4. publish
python scripts/upload_precomputed_to_hub.py --precomputed_dir /tmp/precomputed

The default recipe is exactly two runs, both at seed 0 with untouched sliders -- i.e. precisely what a visitor who changes nothing would ask for:

id configuration
a2a_01 Any-to-Any tab defaults: first example clip, RGB input, "Recommended chain" (9 targets)
fp_01 Future Prediction tab defaults: first future-example clip, RGB observed, first 5 frames, unconditional, every other modality predicted

To widen the cache, add entries to build_recipe() in scripts/precompute_demo_outputs.py and re-run -- no app-side change is needed. Entries already in the manifest are skipped unless you pass --overwrite, so a re-run only generates what's new.

Notes

  • Re-run this whenever the main checkpoint changes, or whenever you edit CONFIGS / HYPERPARAM_PRESETS / COARSE_TO_FINE_BY_INPUT / FUTURE_CHAIN / a tab's default control values. Preset edits are detected -- the app prints a code_fingerprint warning at startup and every lookup simply misses, so the demo falls back to live generation rather than serving stale results. A new checkpoint cannot be detected automatically; the manifest records model_provenance so you can at least tell after the fact.
  • Everything degrades safely. No precomputed/ on the Hub, an unreadable manifest, or media that fails to download all mean "no cache": the demo boots exactly as it did before and every click runs live.
  • Honesty. Cards served from the cache carry a cached pill and the status line reads precomputed · seed N instead of done.
  • Testing without a GPU. --fake writes real mp4s of placeholder art, so the whole manifest/lookup/prefill path can be exercised on a login node:
    python scripts/precompute_demo_outputs.py --fake --out_dir /tmp/pc_fake
    
    The uploader refuses to publish a --fake manifest.

A second curated set (e.g. for the Future Prediction tab)

The Future Prediction tab needs its own hand-picked clips (e.g. ones with clear, consistent motion) rather than a random sample, and shouldn't overwrite the any-to-any tab's set. Same repo, same 3 phases, just with --stems_file instead of --num_examples/--seed, and a distinct --manifest_out/--manifest_repo_filename so the two sets don't collide:

# Phase 1 -- explicit stems (one per line, e.g. "vol_12/clip_000123",
# matching the tok_video_rgb@128 subfolder layout)
python scripts/upload_examples_to_hub.py \
    --source_dir /datasets/uzair/weights_from_clariden/test_cvpr_final_set_13_mod \
    --repo_id EPFL-VILAB/Video-4M-examples \
    --stems_file my_future_pred_stems.txt \
    --manifest_out future_examples_manifest.json \
    --manifest_repo_filename future_examples.json
# Phase 2 -- point the visualization script at this manifest via env var
cd ..
EXAMPLES_MANIFEST_PATH=hf_space_demo/future_examples_manifest.json \
    python visualize_multimodal_pretraining_data_13_modalities.py \
    --config cfgs/default/4m/models/main/4m_L_video_all_modalities_final_config.yaml
# Phase 3 -- upload previews, same distinct manifest names as phase 1
cd hf_space_demo
python scripts/upload_examples_to_hub.py \
    --repo_id EPFL-VILAB/Video-4M-examples \
    --detokenized_dir /datasets/uzair/weights_from_clariden/cvpr_generations/GT_visualizations_post_neurips_opticalflow_fixed \
    --manifest_out future_examples_manifest.json \
    --manifest_repo_filename future_examples.json \
    --previews_only

Every stem in --stems_file must have a file in every modality subfolder -- the script errors out listing exactly which ones are missing, rather than silently dropping them (unlike phase 1's random-selection path, which just skips incomplete candidates).

Notes

  • Lost your local examples_manifest.json? It was also uploaded to the repo as examples.json in phase 1 -- pull it back down instead of re-picking clips:
    python -c "
    from huggingface_hub import hf_hub_download
    import shutil
    path = hf_hub_download(repo_id='EPFL-VILAB/Video-4M-examples', filename='examples.json', repo_type='dataset')
    shutil.copy(path, 'examples_manifest.json')
    "
    
    (For the future-prediction set, substitute future_examples.json / future_examples_manifest.json.)
  • Want a completely different set of clips? Re-run phase 1 with --force_repick (optionally a different --seed/--num_examples, or a different --stems_file), then repeat phases 2 and 3.
  • Just fixing/re-rendering previews for the same clips? Skip phase 1, start at phase 2 -- the manifest already has the clip IDs.