Instructions to use xju-arlab/vbench-model with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use xju-arlab/vbench-model with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
VBench Repair: selected trained models
One selected model per dimension: spatial_relationship/ (v8, step 600),
scene/ (v8, step 300), human_action/ (v9, step 300),
multiple_objects/ (v6, step 900), object_class/ and color/ (paper-final step 300).
For Spatial, Scene, Action and Objects, selection uses the best retained checkpoint by the original development exact-match
probe, with latest-step tie breaking. Each folder includes selection.json, the
training/development record, tokenizer files, portable PEFT configuration, and
SHA-256 of the byte-preserved adapter. The development probe has only 24 examples.
Action's higher-scoring step 150 was already pruned by the original training run;
the published step 300 ties the best surviving 200/250/300 checkpoints. No test
scores are used to select a checkpoint.
The frozen base is Qwen/Qwen3-8B, revision
b968826d9c46dd6066d109eabc6255188de91218. It is not duplicated here.
These are LoRA adapters, not complete standalone 8B model weights.
from huggingface_hub import snapshot_download
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
root = snapshot_download("xju-arlab/vbench-model", allow_patterns=["scene/*"])
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen3-8B", revision="b968826d9c46dd6066d109eabc6255188de91218",
torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(base, root + "/scene")
tokenizer = AutoTokenizer.from_pretrained(root + "/scene")
The full repaired metric also needs the task schemas, deterministic interfaces
and video backends. The training/inference source snapshot is included under
code/vbench_prompts_compile/; this is not a claim that raw LoRA text output
alone reproduces the reported metric. Frozen configs retain source paths as
provenance; users should set their own data paths.
Dynamic Degree: dynamic_degree/aligned.pt is the user-accepted paper model
(fixed step 300, not dev-selected). The paired frozen V-JEPA 2.1 ViT-B backbone
is in dynamic_degree/backbone/. Selection, hashes, original failed dev gates
and completed 450-source evaluation are recorded alongside the weights.
Together with the other repaired methods, the paper now covers nine dimensions.
Data: xju-arlab/vbench-repair. Third-party base-model and dataset terms continue to apply.
Object Class and Color publish the paper-final step 300 at the user's explicit request. These two models were not selected by development evaluation. Their original runs used a fixed 300-step schedule and have no dev comparison between the retained steps 200 and 300. The checkpoint hashes, training results and selection rationale are preserved in each dimension directory. Prompt-only inference source and the frozen vocabulary are in code/object_color/.
Dynamic external generalization: LASIESTA and BMC (2026-09-23)
The unchanged aligned-v1 checkpoint was evaluated on two external real-video datasets; frozen results and limitations. LASIESTA contributes 43 clips/9 recordings and BMC 84 clips/7 recordings. No retraining or posthoc score mapping. BMC mean repair score 0.282374โ0.292002 versus Origin 0.261905โ1.000000 under 8px local texture jitter. BMC labels are exploratory agent review and LASIESTA native timing is unverified. Individual failures, confidence intervals and original dev failures are retained. This update changes documentation only; pinned weight revisions remain valid.
- Downloads last month
- -