Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| VBench | 4,139 items | ||
| VideoAlign | 95 items | ||
| __pycache__ | 5 items | ||
| eval_camera | 10 items | ||
| eval_camera_new | 48 items | ||
| eval_consistency | 8 items | ||
| eval_depth | 10 items | ||
| eval_inpainting | 4 items | ||
| eval_pose | 6 items | ||
| README.md | 18.1 kB xet | c1a3dade | |
| combine_results.py | 11 kB xet | 5b74df1c | |
| derange.py | 2.33 kB xet | 246ae48c | |
| eval_vbench.sh | 6.04 kB xet | bff8d3bc | |
| evaluate.py | 11.3 kB xet | d2208aaa | |
| merge_configs.py | 3.54 kB xet | 43b8c0f6 | |
| merge_seeds.py | 3.52 kB xet | 55e08648 | |
| run_eval.sh | 24.2 kB xet | 8057ef14 | |
| setup_envs.sh | 5.87 kB xet | e2175c24 |
FlexBench Evaluation
Evaluators for generated videos: general video quality (VBench), prompt following, and one condition-fidelity metric group per conditioning type (subject / depth / pose / inpainting).
Everything is driven by run_eval.sh, which picks the applicable
evaluators, runs them, and merges the per-sample scores into a single
combined_results.csv.
Quick start
# one-time: create the two conda envs and check the model checkpoints
evaluation/setup_envs.sh
# evaluate a generation run (dataset dir + condition types come from manifest.json)
evaluation/run_eval.sh --generated-dir results/flexbench/v0/ltx-2/subject_depth
Output: <generated-dir>/eval/combined_results.csv (or
<generated-dir>/seed_<N>/eval/combined_results.csv for a multi-seed run).
If --generated-dir has no manifest.json, pass the dataset dir and the
condition selectors explicitly:
evaluation/run_eval.sh \
--generated-dir results/flexbench/v0/ltx-2/subject_depth \
--dataset-dir datasets/v0/subject_depth \
--subject-consistency --depth
Which metrics run for which condition
| Condition | Runs | Ground truth it needs | Metrics |
|---|---|---|---|
| (always) | VBench dimensions | captions.json |
SuC, BC, AQ, IQ, OvC, MS, DD |
| (always) | text consistency | captions.json |
ViCLIP text–video cosine similarity ↑ |
subject |
subject consistency | <dataset>/subject/<id>.jpg|png |
CLIP ↑, DINO ↑ |
depth |
depth fidelity (background only when combined with pose) | <dataset>/depth/<id>.mp4 |
MAE ↓, RMSE ↓ |
pose |
pose fidelity | <dataset>/pose/raw/<id>[_pose].npy (+ <dataset>/pose/<id>.mp4 for resolution) |
AKD ↓, PCK@0.05 ↑, PCK@0.10 ↑ |
inpainting |
unmasked-region fidelity | <dataset>/inpainting/<id>.mp4 + <dataset>/inpainting/bbox/<id>.mp4 |
PSNR ↑, SSIM ↑, LPIPS ↓, MSE ↓, MAE ↓ |
camera |
camera-trajectory adherence | <dataset>/camera/<id>.mp4 |
RotErr ↓, TransErr ↓, CamMC ↓ |
The first two only need the captions file, so they always run. The other five run
only when both the selector is enabled (via manifest.json's run.conditions, or
an explicit --subject-consistency / --depth / --pose / --inpainting /
--camera flag)
and the corresponding input path exists — otherwise they are skipped with a
message, and the pipeline continues.
A combined config like subject_depth simply enables both of its groups, so the
combined CSV carries the VBench + text columns plus one column block per condition.
The metrics in detail
VBench dimensions — VBench/evaluate.py
Run in custom_input mode over the generated .mp4s with captions.json as the
prompt file. Default dimension set (all higher-is-better, reported on a 0–100 scale):
| Column | Dimension | What it measures |
|---|---|---|
SuC |
subject_consistency |
DINO similarity of the same subject across frames (temporal, not vs. a reference image) |
BC |
background_consistency |
CLIP similarity of the background across frames |
AQ |
aesthetic_quality |
LAION aesthetic predictor, per frame |
IQ |
imaging_quality |
MUSIQ distortion score (blur, noise, over-exposure) |
OvC |
overall_consistency |
ViCLIP video–text alignment, VBench's own prompt-following score |
MS |
motion_smoothness |
AMT frame-interpolation error — how physically plausible the motion is |
DD |
dynamic_degree |
RAFT-based: fraction of videos with non-trivial motion (guards against static-video gaming of the consistency metrics) |
Override with --vbench-dimensions "d1 d2 ...".
Note the i2v dimensions (i2v_subject etc., which need evaluate_i2v.py and a
reference image folder) are not part of the automated pipeline — run_eval.sh
uses the t2v path only. See eval_vbench.sh for hand-written i2v
invocations.
Text consistency — eval_consistency/text_consistency.py
Raw ViCLIP cosine similarity between the prompt and the video (8 middle frames).
Same feature extractor as VBench's overall_consistency, but reads a plain
{"<id>.mp4": "prompt"} captions file and reports the raw similarity rather than
VBench's rendered percentage. Higher is better.
Subject consistency — eval_consistency/subject_consistency.py
CLIP (ViT-B/32) and DINO (ViT-B/16) cosine similarity between the reference subject image and every frame of the generated video, averaged over frames. Higher is better.
Distinct from VBench's i2v_subject (which weights 0.4*max + 0.3*mean + 0.3*min
over frame-to-frame and frame-to-reference similarities) — this one is the plain
per-frame average against the reference, unclamped.
Depth — eval_depth/eval.py, eval_depth/metric.py
- Load the condition depth video (
<dataset>/depth/<id>.mp4, per-video normalized grayscale), resize-to-fill + center-crop to the generated resolution — the same transformic_lora.pyapplies to conditioning videos, so the compared region is the one the model actually saw. - Re-extract depth from the generated video with Video-Depth-Anything (
vitl), normalized per video to the same 0–255 scale. - For a combined
pose + depthrun, prompt SAM 3.1 withpersononce and use its multiplex video detector/tracker to segment every returned person instance. All person masks are merged per frame, so background scoring excludes bystanders as well as the pose-conditioned subject. If SAM3 returns no person on a frame, that frame gets an empty exclusion mask and its full valid depth area is evaluated. Enclosed holes are filled, then the mask is dilated by 2% of the shorter frame side to remove silhouette-boundary leakage and small held-object gaps. - Least-squares scale + shift alignment of the extracted depth to the condition depth, which cancels the affine ambiguity that per-video normalization introduces. All generated-person pixels are excluded from this fit as well as from scoring; otherwise the foreground would still bias the background alignment.
- Score over all frames, excluding all generated people and pixels where condition depth is 0.
Person masking is controlled by --depth-person-mask auto|always|never (default
auto, which enables it exactly when depth and pose are both selected). Tune
--person-mask-dilation if needed. With --verbose, reusable masks, per-frame
provenance/area diagnostics, and generated-overlay/binary-mask videos are kept under
<output>/depth_person_masks/; videos are in its video/ subfolder.
Reported metrics (both lower-is-better, in 0–255 grayscale units — relative disparity, not metric depth):
mae— mean absolute errorrmse_linear— root mean squared error
metric.py also implements the rest of the standard monocular-depth suite
(abs_relative_difference, rmse_log, log10, delta1/2/3_acc, i_rmse,
silog_rmse) — they aren't in the default eval_metrics list because scale-
and shift-aligned relative disparity makes the log- and ratio-based ones hard to
interpret. Add them to eval_metrics in eval_depth/eval.py if you want them.
Pose — eval_pose/eval.py, eval_pose/metric.py
- Load condition keypoints from
<id>_pose.npy(DWPose wholebody, 134 keypoints in OpenPose order: 0–17 body, 18–23 foot, 24–91 face, 92–133 hands). - Re-extract pose from the generated video with DWPose.
- Normalize both to
[0, 1]by their own image dimensions (condition resolution comes from--cond-video-dir, falling back to the generated video's size). - Match persons per frame, greedily by normalized body-centroid proximity.
- Score the first 18 (body) keypoints only, keeping keypoints with confidence
> 0.3in both sequences.
Reported metrics:
akd↓ — Average Keypoint Distance: mean Euclidean distance in normalized[0, 1]space. Resolution-independent, so 0.05 means "5% of the frame diagonal-ish, off".pck@0.05,pck@0.10↑ — Percentage of Correct Keypoints: fraction within that normalized distance of GT.@0.05is the strict threshold;@0.10is more forgiving and useful when the model gets the pose roughly right but the body scale drifts.
Camera — eval_camera_new/eval.py, eval_camera_new/metric.py
CamCloneMaster's camera-accuracy protocol (arXiv 2506.03140, Sec. 5.1), with the three errors defined by CamI2V.
- Estimate per-frame camera poses with MegaSaM for the generated clip and for its camera reference. Both are cached by video content hash, so the duplicated per-subject reference files are only reconstructed once.
- Express each trajectory relative to its own first frame, so the two reconstructions' arbitrary world frames cancel.
- Scale-normalize translations, each trajectory by its own max ‖t‖ — monocular reconstruction cannot recover absolute scale.
- Uniformly resample both to the shorter length, then score.
| Metric | Direction | Notes |
|---|---|---|
RotErr |
↓ | Σ arccos((tr(R_refᵀ R_gen) − 1)/2) — geodesic rotation error in radians, summed over frames |
TransErr |
↓ | Σ‖t_gen − t_ref‖₂ on scale-normalized translations, so it measures the shape of the translation path, not its magnitude |
CamMC |
↓ | `Σ‖[R |
Sums match the papers' convention; *_per_frame means are emitted too and are what to
compare when clips differ in length.
Every metric also carries <metric>_chance: the static-camera control, an identity pose
at every frame scored against the same reference. References differ enormously in motion
magnitude, so a raw error partly reports which camera the sample drew — read
<metric>_rel = <metric> / <metric>_chance instead, where below 1 means the shot beat a
locked-off camera. Because that control is analytic, --shuffle-refs is a no-op here.
Caveats: MegaSaM is a monocular reconstruction, so a low-parallax or heavily dynamic clip
can yield a near-degenerate trajectory — those are flagged (degenerate, ref_max_trans,
gen_max_trans) rather than reported as sound. This needs a GPU, roughly 100 s per clip on
an A100, in its own flexbench-megasam env. See
eval_camera_new/README.md.
Inpainting — eval_inpainting/eval.py, eval_inpainting/metric.py
Measures how well the generated video preserves the region it was not asked to regenerate — i.e. reconstruction fidelity outside the bbox.
- Reference =
<dataset>/inpainting/<id>.mp4(the original video with the bbox filled black; outside the box it is the original, re-encoded). - Mask =
<dataset>/inpainting/bbox/<id>.mp4(white = the region to inpaint). Thresholded at 127/255 since the masks are h264-compressed; after resampling, any pixel the box even partially touches counts as masked. - Align reference and mask to the generated video: truncate to the shortest clip, then resize-to-fill + center-crop to the generated resolution.
- Score over the unmasked pixels only.
| Metric | Direction | Notes |
|---|---|---|
psnr |
↑ | dB, per frame over valid pixels, averaged across frames |
ssim |
↑ | [0, 1]. 11×11 Gaussian window, valid-mode convolution, and the mask is eroded by the same window — so no averaged window ever straddles the inpainted box |
lpips |
↓ | [0, 1]. LPIPS is a deep patch metric and can't be restricted to a pixel set, so the masked region is filled with the same mid-gray in both videos and contributes ~0 distance (the standard convention). Backbone via --lpips-net (default alex) |
mse |
↓ | on the 0–255 pixel scale, not [0, 1] — a [0,1]-scale MSE of ~3e-4 would collapse to 0.0003 under the CSV's 4-decimal formatting |
mae |
↓ | likewise 0–255 |
Caveat: the reference is a re-encode of the original, which puts a ceiling of roughly ~36 dB on PSNR. These scores are comparable across configs and seeds, not against numbers in the literature.
Outputs
Each evaluator writes results.json ({"mean": ..., "per_sample": {...}}) and
results.csv into its own subfolder. combine_results.py
merges whichever subfolders exist into one wide table:
text subject depth vbench
sample_id consistency ↑ clip ↑ dino ↑ mae ↓ rmse_linear ↓ SuC ↑ BC ↑ ...
average 0.2314 0.7821 0.6543 12.3401 18.9922 94.12 96.03 ...
001 ...
Two header rows: a group label (written once per group so it reads as a spanning
header in a spreadsheet) and the metric name carrying ↑/↓ for its direction.
The average row comes right after the header, before the per-sample rows.
By default only combined_results.csv is kept — the per-evaluator subfolders go to
a scratch dir and are deleted. Pass --verbose to keep them under --output-dir.
Multi-seed and cross-config aggregation
# run_eval.sh already evaluates each seed_<N>/ subfolder when manifest.json lists seeds
evaluation/run_eval.sh --generated-dir results/flexbench/v0_scale_up/lora-switch/subject_depth
# collapse the per-seed averages into "mean (std)" per metric
python evaluation/merge_seeds.py --generated-dir results/flexbench/v0_scale_up/lora-switch/subject_depth
# -> <generated-dir>/seed_summary.csv
# compare that summary row across configs, marking the best per metric with "*"
python evaluation/merge_configs.py --subfolder subject_depth \
--configs lora-switch lora-direct-merge lora-norm-consistency lora-union
# -> <base-dir>/subject_depth_comparison.csv
merge_configs.py knows which metrics are lower-is-better (pose_akd, depth_mae,
depth_rmse_linear, inpaint_lpips, inpaint_mse, inpaint_mae) — extend
LOWER_IS_BETTER there if you add an error metric.
Environments and checkpoints
setup_envs.sh creates three conda envs (re-runnable; it skips existing envs):
| Env | Used by | Notes |
|---|---|---|
flexbench-vbench |
VBench dimensions, text consistency, subject consistency | torch 2.7.1 + cu118, editable install of evaluation/VBench |
flexbench-cond |
depth, pose, inpainting | torch 2.1.1 + cu118, numpy<2, onnxruntime-gpu + pip cuDNN 9 (the cu118 torch build only bundles cuDNN 8), lpips |
sam3 |
all-person masks for pose+depth | SAM 3.1 multiplex, torch 2.10 + cu128, optional compile mode |
Override with --vbench-env / --cond-env / --sam3-env. Select the GPU
with --gpu N.
Checkpoints that must be present (setup_envs.sh checks and prints the download command for any that are missing):
tools/Video-Depth-Anything/checkpoints/video_depth_anything_vitl.pthtools/DWPose/ControlNet-v1-1-nightly/annotator/ckpts/yolox_l.onnxtools/DWPose/ControlNet-v1-1-nightly/annotator/ckpts/dw-ll_ucoco_384.onnxtools/sam3/sam3.1/sam3.1_multiplex.pt- ViCLIP:
~/.cache/vbench/ViCLIP/ViClip-InternVid-10M-FLT.pth(plus its BPE vocab) - CLIP / DINO / LPIPS backbones download themselves on first use
Running a single evaluator
Each evaluator is a standalone CLI, useful for debugging one condition. Paths below are relative to the repo root; the working directory doesn't matter.
# Optional first step for standalone background-only depth scoring. run_eval.sh
# does this automatically for pose+depth.
conda run -n sam3 python evaluation/eval_depth/generate_person_masks.py \
--generated-dir <videos> \
--sam3-repo-path tools/sam3 \
--checkpoint tools/sam3/sam3.1/sam3.1_multiplex.pt \
--output-dir /tmp/person_masks \
--save-video
conda run -n flexbench-cond python evaluation/eval_depth/eval.py \
--generated-dir <videos> \
--dataset-dir <dataset>/depth \
--depth-repo-path tools/Video-Depth-Anything \
--exclude-mask-dir /tmp/person_masks \
--output-dir /tmp/depth_eval \
--show-video # generated RGB | condition depth | aligned extracted depth
conda run -n flexbench-cond python evaluation/eval_pose/eval.py \
--generated-dir <videos> \
--dataset-dir <dataset>/pose/raw \
--dwpose-repo-path tools/DWPose \
--cond-video-dir <dataset>/pose \
--output-dir /tmp/pose_eval \
--show-video # generated RGB | condition pose | extracted pose
conda run -n flexbench-cond python evaluation/eval_inpainting/eval.py \
--generated-dir <videos> \
--dataset-dir <dataset>/inpainting \
--output-dir /tmp/inpaint_eval \
--show-video # generated | reference | generated with the scored region kept
conda run -n flexbench-vbench python evaluation/eval_consistency/subject_consistency.py \
--generated-dir <videos> --subject-dir <dataset>/subject --output-dir /tmp/subj_eval
--show-video (depth / pose / inpainting) writes side-by-side comparison clips —
the fastest way to check that alignment, cropping, and masking are behaving before
trusting the numbers.
Notes
- Sample matching is by file stem everywhere:
<id>.mp4in the generated dir must match<id>.mp4/<id>.jpg/<id>_pose.npyin the dataset dir. Non-matching samples are silently dropped from that evaluator's average, so check the reported sample count if a number looks off. - Frame-count mismatches are handled by truncating to the shortest clip; resolution
mismatches by resize-to-fill + center-crop to the generated resolution (matching
ic_lora.py's conditioning transform). VideoAlign/is vendored here but is not wired intorun_eval.sh.
- Total size
- 69.5 GB
- Files
- 49,395
- Last updated
- Aug 20
- Pre-warmed CDN
- US EU US EU