Spaces:
Running on Zero
Download scripts/README_EXAMPLES.md from EPFL-VILAB/Video-4M: direct link, hf CLI and curl.
- Browser
- Download file 8.18 kB
-
https://huggingface.co/spaces/EPFL-VILAB/Video-4M/resolve/main/scripts/README_EXAMPLES.md
- Command line
-
hf download hf://spaces/EPFL-VILAB/Video-4M/scripts/README_EXAMPLES.md
-
curl -L -o README_EXAMPLES.md https://huggingface.co/spaces/EPFL-VILAB/Video-4M/resolve/main/scripts/README_EXAMPLES.md
A newer version of the Gradio SDK is available: 6.30.0
Adding / refreshing curated examples for the Any-to-Any tab
Phases 1-3 curate the example inputs; phase 4 pre-generates the model
outputs the demo shows without a GPU. Run them in order, all from inside
hf_space_demo/ on the cluster (where both the tokenized dataset and the
checkpoints actually live).
Phase 1 -- pick clips + upload tokens
python scripts/upload_examples_to_hub.py \
--source_dir /datasets/uzair/weights_from_clariden/test_cvpr_final_set_13_mod \
--repo_id EPFL-VILAB/Video-4M-examples \
--num_examples 6 --seed 0
This randomly picks --num_examples clips that have a file in every modality
subfolder, uploads all of their tokenized data to the EPFL-VILAB/Video-4M-examples
dataset repo, and saves the picked clip IDs locally to examples_manifest.json
(in whatever directory you ran the command from -- keep this file, phases 2
and 3 both need it).
Phase 2 -- render human-viewable previews
cd .. # repo root, since the script reads hf_space_demo/examples_manifest.json
python visualize_multimodal_pretraining_data_13_modalities.py \
--config cfgs/default/4m/models/main/4m_L_video_all_modalities_final_config.yaml
This reads hf_space_demo/examples_manifest.json and renders a
.mp4 preview per modality for exactly those clips (not a random sample),
saved under args.output_dir_videos (default:
/datasets/uzair/weights_from_clariden/cvpr_generations/GT_visualizations_post_neurips_opticalflow_fixed),
in <domain>_detokenized/<clip_basename>.mp4 subfolders.
Phase 3 -- upload the previews
cd hf_space_demo
python scripts/upload_examples_to_hub.py \
--repo_id EPFL-VILAB/Video-4M-examples \
--detokenized_dir /datasets/uzair/weights_from_clariden/cvpr_generations/GT_visualizations_post_neurips_opticalflow_fixed \
--previews_only
--previews_only skips re-uploading the (large, slow) tokens. Without
--force_repick, this automatically reuses the exact clips already recorded
in examples_manifest.json -- it will not pick a new random set.
Once this is done, the Any-to-Any tab's example/input-modality dropdowns will show a live preview before you even hit Generate.
Phase 4 -- precompute the demo's default results
Phases 1-3 give the demo example inputs. This phase generates the outputs for the default configuration of each tab, so the Space can open on real predictions with no GPU at all.
Why it matters: ZeroGPU gives free-tier visitors a few minutes of GPU per
day, and load_pipeline's ~37.5 GB checkpoint load happens inside that
allocation -- so a single click can burn the whole budget and a visitor never
gets to see both tabs. With this in place the page opens on finished
predictions, and clicking Generate without changing anything is free. Anything
else (a different clip, a custom chain, an upload, a typed caption) still runs
live on the GPU.
source <(grep '^export' run_demo.sh) # FOURM_LOCAL_* checkpoint paths
export HF_TOKEN=...
# 1. inspect the two configurations -- loads no checkpoints, safe on a login node
python scripts/precompute_demo_outputs.py --dry_run
# 2. the real run: ~11 mp4s, ~10 min of GPU including the one-time model load
python scripts/precompute_demo_outputs.py --out_dir /tmp/precomputed
# 3. optional local check with no Hub round trip -- copy /tmp/precomputed to
# <examples dir>/precomputed/ first
FOURM_LOCAL_EXAMPLES_DIR=<examples dir> python app.py
# 4. publish
python scripts/upload_precomputed_to_hub.py --precomputed_dir /tmp/precomputed
The default recipe is exactly two runs, both at seed 0 with untouched sliders -- i.e. precisely what a visitor who changes nothing would ask for:
| id | configuration |
|---|---|
a2a_01 |
Any-to-Any tab defaults: first example clip, RGB input, "Recommended chain" (9 targets) |
fp_01 |
Future Prediction tab defaults: first future-example clip, RGB observed, first 5 frames, unconditional, every other modality predicted |
To widen the cache, add entries to build_recipe() in
scripts/precompute_demo_outputs.py and re-run -- no app-side change is
needed. Entries already in the manifest are skipped unless you pass
--overwrite, so a re-run only generates what's new.
Notes
- Re-run this whenever the main checkpoint changes, or whenever you edit
CONFIGS/HYPERPARAM_PRESETS/COARSE_TO_FINE_BY_INPUT/FUTURE_CHAIN/ a tab's default control values. Preset edits are detected -- the app prints acode_fingerprintwarning at startup and every lookup simply misses, so the demo falls back to live generation rather than serving stale results. A new checkpoint cannot be detected automatically; the manifest recordsmodel_provenanceso you can at least tell after the fact. - Everything degrades safely. No
precomputed/on the Hub, an unreadable manifest, or media that fails to download all mean "no cache": the demo boots exactly as it did before and every click runs live. - Honesty. Cards served from the cache carry a
cachedpill and the status line readsprecomputed · seed Ninstead ofdone. - Testing without a GPU.
--fakewrites real mp4s of placeholder art, so the whole manifest/lookup/prefill path can be exercised on a login node:
The uploader refuses to publish apython scripts/precompute_demo_outputs.py --fake --out_dir /tmp/pc_fake--fakemanifest.
A second curated set (e.g. for the Future Prediction tab)
The Future Prediction tab needs its own hand-picked clips (e.g. ones with
clear, consistent motion) rather than a random sample, and shouldn't
overwrite the any-to-any tab's set. Same repo, same 3 phases, just with
--stems_file instead of --num_examples/--seed, and a distinct
--manifest_out/--manifest_repo_filename so the two sets don't collide:
# Phase 1 -- explicit stems (one per line, e.g. "vol_12/clip_000123",
# matching the tok_video_rgb@128 subfolder layout)
python scripts/upload_examples_to_hub.py \
--source_dir /datasets/uzair/weights_from_clariden/test_cvpr_final_set_13_mod \
--repo_id EPFL-VILAB/Video-4M-examples \
--stems_file my_future_pred_stems.txt \
--manifest_out future_examples_manifest.json \
--manifest_repo_filename future_examples.json
# Phase 2 -- point the visualization script at this manifest via env var
cd ..
EXAMPLES_MANIFEST_PATH=hf_space_demo/future_examples_manifest.json \
python visualize_multimodal_pretraining_data_13_modalities.py \
--config cfgs/default/4m/models/main/4m_L_video_all_modalities_final_config.yaml
# Phase 3 -- upload previews, same distinct manifest names as phase 1
cd hf_space_demo
python scripts/upload_examples_to_hub.py \
--repo_id EPFL-VILAB/Video-4M-examples \
--detokenized_dir /datasets/uzair/weights_from_clariden/cvpr_generations/GT_visualizations_post_neurips_opticalflow_fixed \
--manifest_out future_examples_manifest.json \
--manifest_repo_filename future_examples.json \
--previews_only
Every stem in --stems_file must have a file in every modality subfolder --
the script errors out listing exactly which ones are missing, rather than
silently dropping them (unlike phase 1's random-selection path, which just
skips incomplete candidates).
Notes
- Lost your local
examples_manifest.json? It was also uploaded to the repo asexamples.jsonin phase 1 -- pull it back down instead of re-picking clips:(For the future-prediction set, substitutepython -c " from huggingface_hub import hf_hub_download import shutil path = hf_hub_download(repo_id='EPFL-VILAB/Video-4M-examples', filename='examples.json', repo_type='dataset') shutil.copy(path, 'examples_manifest.json') "future_examples.json/future_examples_manifest.json.) - Want a completely different set of clips? Re-run phase 1 with
--force_repick(optionally a different--seed/--num_examples, or a different--stems_file), then repeat phases 2 and 3. - Just fixing/re-rendering previews for the same clips? Skip phase 1, start at phase 2 -- the manifest already has the clip IDs.