69.5 GB
49,395 files
Updated about 2 months ago
Name
Size
VBench
VideoAlign
__pycache__
eval_camera
eval_camera_new
eval_consistency
eval_depth
eval_inpainting
eval_pose
README.md18.1 kB
xet
combine_results.py11 kB
xet
derange.py2.33 kB
xet
eval_vbench.sh6.04 kB
xet
evaluate.py11.3 kB
xet
merge_configs.py3.54 kB
xet
merge_seeds.py3.52 kB
xet
run_eval.sh24.2 kB
xet
setup_envs.sh5.87 kB
xet
README.md

FlexBench Evaluation

Evaluators for generated videos: general video quality (VBench), prompt following, and one condition-fidelity metric group per conditioning type (subject / depth / pose / inpainting).

Everything is driven by run_eval.sh, which picks the applicable evaluators, runs them, and merges the per-sample scores into a single combined_results.csv.


Quick start

# one-time: create the two conda envs and check the model checkpoints
evaluation/setup_envs.sh

# evaluate a generation run (dataset dir + condition types come from manifest.json)
evaluation/run_eval.sh --generated-dir results/flexbench/v0/ltx-2/subject_depth

Output: <generated-dir>/eval/combined_results.csv (or <generated-dir>/seed_<N>/eval/combined_results.csv for a multi-seed run).

If --generated-dir has no manifest.json, pass the dataset dir and the condition selectors explicitly:

evaluation/run_eval.sh \
    --generated-dir results/flexbench/v0/ltx-2/subject_depth \
    --dataset-dir   datasets/v0/subject_depth \
    --subject-consistency --depth

Which metrics run for which condition

Condition Runs Ground truth it needs Metrics
(always) VBench dimensions captions.json SuC, BC, AQ, IQ, OvC, MS, DD
(always) text consistency captions.json ViCLIP text–video cosine similarity ↑
subject subject consistency <dataset>/subject/<id>.jpg|png CLIP ↑, DINO ↑
depth depth fidelity (background only when combined with pose) <dataset>/depth/<id>.mp4 MAE ↓, RMSE ↓
pose pose fidelity <dataset>/pose/raw/<id>[_pose].npy (+ <dataset>/pose/<id>.mp4 for resolution) AKD ↓, PCK@0.05 ↑, PCK@0.10 ↑
inpainting unmasked-region fidelity <dataset>/inpainting/<id>.mp4 + <dataset>/inpainting/bbox/<id>.mp4 PSNR ↑, SSIM ↑, LPIPS ↓, MSE ↓, MAE ↓
camera camera-trajectory adherence <dataset>/camera/<id>.mp4 RotErr ↓, TransErr ↓, CamMC ↓

The first two only need the captions file, so they always run. The other five run only when both the selector is enabled (via manifest.json's run.conditions, or an explicit --subject-consistency / --depth / --pose / --inpainting / --camera flag) and the corresponding input path exists — otherwise they are skipped with a message, and the pipeline continues.

A combined config like subject_depth simply enables both of its groups, so the combined CSV carries the VBench + text columns plus one column block per condition.


The metrics in detail

VBench dimensions — VBench/evaluate.py

Run in custom_input mode over the generated .mp4s with captions.json as the prompt file. Default dimension set (all higher-is-better, reported on a 0–100 scale):

Column Dimension What it measures
SuC subject_consistency DINO similarity of the same subject across frames (temporal, not vs. a reference image)
BC background_consistency CLIP similarity of the background across frames
AQ aesthetic_quality LAION aesthetic predictor, per frame
IQ imaging_quality MUSIQ distortion score (blur, noise, over-exposure)
OvC overall_consistency ViCLIP video–text alignment, VBench's own prompt-following score
MS motion_smoothness AMT frame-interpolation error — how physically plausible the motion is
DD dynamic_degree RAFT-based: fraction of videos with non-trivial motion (guards against static-video gaming of the consistency metrics)

Override with --vbench-dimensions "d1 d2 ...".

Note the i2v dimensions (i2v_subject etc., which need evaluate_i2v.py and a reference image folder) are not part of the automated pipeline — run_eval.sh uses the t2v path only. See eval_vbench.sh for hand-written i2v invocations.

Text consistency — eval_consistency/text_consistency.py

Raw ViCLIP cosine similarity between the prompt and the video (8 middle frames). Same feature extractor as VBench's overall_consistency, but reads a plain {"<id>.mp4": "prompt"} captions file and reports the raw similarity rather than VBench's rendered percentage. Higher is better.

Subject consistency — eval_consistency/subject_consistency.py

CLIP (ViT-B/32) and DINO (ViT-B/16) cosine similarity between the reference subject image and every frame of the generated video, averaged over frames. Higher is better.

Distinct from VBench's i2v_subject (which weights 0.4*max + 0.3*mean + 0.3*min over frame-to-frame and frame-to-reference similarities) — this one is the plain per-frame average against the reference, unclamped.

Depth — eval_depth/eval.py, eval_depth/metric.py

  1. Load the condition depth video (<dataset>/depth/<id>.mp4, per-video normalized grayscale), resize-to-fill + center-crop to the generated resolution — the same transform ic_lora.py applies to conditioning videos, so the compared region is the one the model actually saw.
  2. Re-extract depth from the generated video with Video-Depth-Anything (vitl), normalized per video to the same 0–255 scale.
  3. For a combined pose + depth run, prompt SAM 3.1 with person once and use its multiplex video detector/tracker to segment every returned person instance. All person masks are merged per frame, so background scoring excludes bystanders as well as the pose-conditioned subject. If SAM3 returns no person on a frame, that frame gets an empty exclusion mask and its full valid depth area is evaluated. Enclosed holes are filled, then the mask is dilated by 2% of the shorter frame side to remove silhouette-boundary leakage and small held-object gaps.
  4. Least-squares scale + shift alignment of the extracted depth to the condition depth, which cancels the affine ambiguity that per-video normalization introduces. All generated-person pixels are excluded from this fit as well as from scoring; otherwise the foreground would still bias the background alignment.
  5. Score over all frames, excluding all generated people and pixels where condition depth is 0.

Person masking is controlled by --depth-person-mask auto|always|never (default auto, which enables it exactly when depth and pose are both selected). Tune --person-mask-dilation if needed. With --verbose, reusable masks, per-frame provenance/area diagnostics, and generated-overlay/binary-mask videos are kept under <output>/depth_person_masks/; videos are in its video/ subfolder.

Reported metrics (both lower-is-better, in 0–255 grayscale units — relative disparity, not metric depth):

  • mae — mean absolute error
  • rmse_linear — root mean squared error

metric.py also implements the rest of the standard monocular-depth suite (abs_relative_difference, rmse_log, log10, delta1/2/3_acc, i_rmse, silog_rmse) — they aren't in the default eval_metrics list because scale- and shift-aligned relative disparity makes the log- and ratio-based ones hard to interpret. Add them to eval_metrics in eval_depth/eval.py if you want them.

Pose — eval_pose/eval.py, eval_pose/metric.py

  1. Load condition keypoints from <id>_pose.npy (DWPose wholebody, 134 keypoints in OpenPose order: 0–17 body, 18–23 foot, 24–91 face, 92–133 hands).
  2. Re-extract pose from the generated video with DWPose.
  3. Normalize both to [0, 1] by their own image dimensions (condition resolution comes from --cond-video-dir, falling back to the generated video's size).
  4. Match persons per frame, greedily by normalized body-centroid proximity.
  5. Score the first 18 (body) keypoints only, keeping keypoints with confidence > 0.3 in both sequences.

Reported metrics:

  • akd ↓ — Average Keypoint Distance: mean Euclidean distance in normalized [0, 1] space. Resolution-independent, so 0.05 means "5% of the frame diagonal-ish, off".
  • pck@0.05, pck@0.10 ↑ — Percentage of Correct Keypoints: fraction within that normalized distance of GT. @0.05 is the strict threshold; @0.10 is more forgiving and useful when the model gets the pose roughly right but the body scale drifts.

Camera — eval_camera_new/eval.py, eval_camera_new/metric.py

CamCloneMaster's camera-accuracy protocol (arXiv 2506.03140, Sec. 5.1), with the three errors defined by CamI2V.

  1. Estimate per-frame camera poses with MegaSaM for the generated clip and for its camera reference. Both are cached by video content hash, so the duplicated per-subject reference files are only reconstructed once.
  2. Express each trajectory relative to its own first frame, so the two reconstructions' arbitrary world frames cancel.
  3. Scale-normalize translations, each trajectory by its own max ‖t‖ — monocular reconstruction cannot recover absolute scale.
  4. Uniformly resample both to the shorter length, then score.
Metric Direction Notes
RotErr ↓ Σ arccos((tr(R_refᵀ R_gen) − 1)/2) — geodesic rotation error in radians, summed over frames
TransErr ↓ Σ‖t_gen − t_ref‖₂ on scale-normalized translations, so it measures the shape of the translation path, not its magnitude
CamMC ↓ `Σ‖[R

Sums match the papers' convention; *_per_frame means are emitted too and are what to compare when clips differ in length.

Every metric also carries <metric>_chance: the static-camera control, an identity pose at every frame scored against the same reference. References differ enormously in motion magnitude, so a raw error partly reports which camera the sample drew — read <metric>_rel = <metric> / <metric>_chance instead, where below 1 means the shot beat a locked-off camera. Because that control is analytic, --shuffle-refs is a no-op here.

Caveats: MegaSaM is a monocular reconstruction, so a low-parallax or heavily dynamic clip can yield a near-degenerate trajectory — those are flagged (degenerate, ref_max_trans, gen_max_trans) rather than reported as sound. This needs a GPU, roughly 100 s per clip on an A100, in its own flexbench-megasam env. See eval_camera_new/README.md.

Inpainting — eval_inpainting/eval.py, eval_inpainting/metric.py

Measures how well the generated video preserves the region it was not asked to regenerate — i.e. reconstruction fidelity outside the bbox.

  1. Reference = <dataset>/inpainting/<id>.mp4 (the original video with the bbox filled black; outside the box it is the original, re-encoded).
  2. Mask = <dataset>/inpainting/bbox/<id>.mp4 (white = the region to inpaint). Thresholded at 127/255 since the masks are h264-compressed; after resampling, any pixel the box even partially touches counts as masked.
  3. Align reference and mask to the generated video: truncate to the shortest clip, then resize-to-fill + center-crop to the generated resolution.
  4. Score over the unmasked pixels only.
Metric Direction Notes
psnr ↑ dB, per frame over valid pixels, averaged across frames
ssim ↑ [0, 1]. 11×11 Gaussian window, valid-mode convolution, and the mask is eroded by the same window — so no averaged window ever straddles the inpainted box
lpips ↓ [0, 1]. LPIPS is a deep patch metric and can't be restricted to a pixel set, so the masked region is filled with the same mid-gray in both videos and contributes ~0 distance (the standard convention). Backbone via --lpips-net (default alex)
mse ↓ on the 0–255 pixel scale, not [0, 1] — a [0,1]-scale MSE of ~3e-4 would collapse to 0.0003 under the CSV's 4-decimal formatting
mae ↓ likewise 0–255

Caveat: the reference is a re-encode of the original, which puts a ceiling of roughly ~36 dB on PSNR. These scores are comparable across configs and seeds, not against numbers in the literature.


Outputs

Each evaluator writes results.json ({"mean": ..., "per_sample": {...}}) and results.csv into its own subfolder. combine_results.py merges whichever subfolders exist into one wide table:

                text     subject          depth          vbench
sample_id       consistency ↑  clip ↑  dino ↑  mae ↓  rmse_linear ↓  SuC ↑  BC ↑  ...
average         0.2314   0.7821  0.6543  12.3401  18.9922        94.12  96.03 ...
001             ...

Two header rows: a group label (written once per group so it reads as a spanning header in a spreadsheet) and the metric name carrying ↑/↓ for its direction. The average row comes right after the header, before the per-sample rows.

By default only combined_results.csv is kept — the per-evaluator subfolders go to a scratch dir and are deleted. Pass --verbose to keep them under --output-dir.

Multi-seed and cross-config aggregation

# run_eval.sh already evaluates each seed_<N>/ subfolder when manifest.json lists seeds
evaluation/run_eval.sh --generated-dir results/flexbench/v0_scale_up/lora-switch/subject_depth

# collapse the per-seed averages into "mean (std)" per metric
python evaluation/merge_seeds.py --generated-dir results/flexbench/v0_scale_up/lora-switch/subject_depth
#   -> <generated-dir>/seed_summary.csv

# compare that summary row across configs, marking the best per metric with "*"
python evaluation/merge_configs.py --subfolder subject_depth \
    --configs lora-switch lora-direct-merge lora-norm-consistency lora-union
#   -> <base-dir>/subject_depth_comparison.csv

merge_configs.py knows which metrics are lower-is-better (pose_akd, depth_mae, depth_rmse_linear, inpaint_lpips, inpaint_mse, inpaint_mae) — extend LOWER_IS_BETTER there if you add an error metric.


Environments and checkpoints

setup_envs.sh creates three conda envs (re-runnable; it skips existing envs):

Env Used by Notes
flexbench-vbench VBench dimensions, text consistency, subject consistency torch 2.7.1 + cu118, editable install of evaluation/VBench
flexbench-cond depth, pose, inpainting torch 2.1.1 + cu118, numpy<2, onnxruntime-gpu + pip cuDNN 9 (the cu118 torch build only bundles cuDNN 8), lpips
sam3 all-person masks for pose+depth SAM 3.1 multiplex, torch 2.10 + cu128, optional compile mode

Override with --vbench-env / --cond-env / --sam3-env. Select the GPU with --gpu N.

Checkpoints that must be present (setup_envs.sh checks and prints the download command for any that are missing):

  • tools/Video-Depth-Anything/checkpoints/video_depth_anything_vitl.pth
  • tools/DWPose/ControlNet-v1-1-nightly/annotator/ckpts/yolox_l.onnx
  • tools/DWPose/ControlNet-v1-1-nightly/annotator/ckpts/dw-ll_ucoco_384.onnx
  • tools/sam3/sam3.1/sam3.1_multiplex.pt
  • ViCLIP: ~/.cache/vbench/ViCLIP/ViClip-InternVid-10M-FLT.pth (plus its BPE vocab)
  • CLIP / DINO / LPIPS backbones download themselves on first use

Running a single evaluator

Each evaluator is a standalone CLI, useful for debugging one condition. Paths below are relative to the repo root; the working directory doesn't matter.

# Optional first step for standalone background-only depth scoring. run_eval.sh
# does this automatically for pose+depth.
conda run -n sam3 python evaluation/eval_depth/generate_person_masks.py \
    --generated-dir <videos> \
    --sam3-repo-path tools/sam3 \
    --checkpoint tools/sam3/sam3.1/sam3.1_multiplex.pt \
    --output-dir /tmp/person_masks \
    --save-video

conda run -n flexbench-cond python evaluation/eval_depth/eval.py \
    --generated-dir <videos> \
    --dataset-dir   <dataset>/depth \
    --depth-repo-path tools/Video-Depth-Anything \
    --exclude-mask-dir /tmp/person_masks \
    --output-dir /tmp/depth_eval \
    --show-video          # generated RGB | condition depth | aligned extracted depth

conda run -n flexbench-cond python evaluation/eval_pose/eval.py \
    --generated-dir <videos> \
    --dataset-dir   <dataset>/pose/raw \
    --dwpose-repo-path tools/DWPose \
    --cond-video-dir <dataset>/pose \
    --output-dir /tmp/pose_eval \
    --show-video          # generated RGB | condition pose | extracted pose

conda run -n flexbench-cond python evaluation/eval_inpainting/eval.py \
    --generated-dir <videos> \
    --dataset-dir   <dataset>/inpainting \
    --output-dir /tmp/inpaint_eval \
    --show-video          # generated | reference | generated with the scored region kept

conda run -n flexbench-vbench python evaluation/eval_consistency/subject_consistency.py \
    --generated-dir <videos> --subject-dir <dataset>/subject --output-dir /tmp/subj_eval

--show-video (depth / pose / inpainting) writes side-by-side comparison clips — the fastest way to check that alignment, cropping, and masking are behaving before trusting the numbers.


Notes

  • Sample matching is by file stem everywhere: <id>.mp4 in the generated dir must match <id>.mp4 / <id>.jpg / <id>_pose.npy in the dataset dir. Non-matching samples are silently dropped from that evaluator's average, so check the reported sample count if a number looks off.
  • Frame-count mismatches are handled by truncating to the shortest clip; resolution mismatches by resize-to-fill + center-crop to the generated resolution (matching ic_lora.py's conditioning transform).
  • VideoAlign/ is vendored here but is not wired into run_eval.sh.
Total size
69.5 GB
Files
49,395
Last updated
Aug 20
Pre-warmed CDN
US EU US EU

Contributors