ProcObject-Qwen3-VL-4B

Qwen3-VL-4B-Instruct fine-tuned with object-centric SFT on ProcObject-10K — the fine-tuned model of ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos (NeurIPS 2026, Evaluations & Datasets Track). Paper · Code · Dataset

Given frames of an instructional video clip and a question about an object, the model answers and localizes the supporting moments:

{"answer": "The tortilla starts whole on the table, is torn into two pieces, and finally placed into a bowl.", "evidence": [[0, 8], [14, 16]]}

evidence intervals are in seconds from the start of the clip.

Training

Two stages on the 9,472 ProcObject-10K training QA (8 GPUs, bf16):

  1. Evidence-prediction SFT — LoRA (r=16, α=32) on the language model's linear layers, trained on the JSON answer + evidence target; 3 epochs, lr 5e-5; 2 FPS, ≤48 frames.
  2. Object-centric SFT — continues the same LoRA, adds LoRA on the attention layers of the top third of the vision blocks, tunes the vision merger, and trains two auxiliary heads: a spatial head supervised by soft patch masks from Grounding DINO boxes of LLM-extracted object phrases, and a temporal head supervised by frame-in-evidence labels. L = L_gen + 0.05·L_spl + 0.10·L_tmp, 3 epochs.

The auxiliary heads are training-only and are not part of this checkpoint: it is a standard merged Qwen3-VL model with no extra inference cost. Full recipe and code: github.com/WenliangGuo/ProcObject-10K.

Results on the ProcObject-10K test set

With the released evaluation harness (48 frames; J. = 0–5 LLM judge, mean of Qwen3-4B and Llama-3.2-3B):

S. B. J. mIoU mIoP mIoG R@0.3
Qwen3-VL-4B-Instruct (zero-shot, paper) 73.3 89.2 – 38.6 63.1 54.1 –
this model 80.8 92.0 3.25 45.3 66.8 56.4 63.5

mIoU is 51.9 on Multi-hop Reasoning and 35.9 on Needle-in-a-Haystack questions.

Usage

The reported numbers use the benchmark protocol: 48 frames sampled uniformly from the clip, resized to 512×512, each sent as an image preceded by [Timestamp: <t>s], after the system prompt in benchmark/prompts/system_prompt.txt, greedy decoding. The easiest way to reproduce them is the repository's harness:

git clone https://github.com/WenliangGuo/ProcObject-10K && cd ProcObject-10K
pip install -r requirements/eval.txt
# build data/clips/ first (preprocess/README.md), then:
MODEL=BrightGuo/ProcObject-Qwen3-VL-4B NAME=procobject_qwen3vl_4b NUM_FRAMES=48 \
    bash benchmark/scripts/run_sharded.sh 0
bash benchmark/scripts/evaluate.sh benchmark/pred_results/procobject_qwen3vl_4b_predictions.json

Or serve it with vLLM (vllm serve BrightGuo/ProcObject-Qwen3-VL-4B --limit-mm-per-prompt '{"image": 64}') and send the same message layout through the OpenAI-compatible API (benchmark/models/runners.py, VLLMRunner).

Tested with vLLM 0.17.1 / transformers 4.57.6.

License

CC BY-NC 4.0, following the ProcObject-10K annotations it was trained on; the base model Qwen3-VL-4B-Instruct is Apache-2.0.

Citation

@inproceedings{guo2026procobject,
  title     = {{ProcObject-10K}: Benchmarking Object-Centric Procedural Understanding in Instructional Videos},
  author    = {Guo, Wenliang and Kong, Yu},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track},
  year      = {2026},
  url       = {https://arxiv.org/abs/2512.03479}
}
Downloads last month
21
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BrightGuo/ProcObject-Qwen3-VL-4B

Finetuned
(469)
this model
Quantizations
1 model

Dataset used to train BrightGuo/ProcObject-Qwen3-VL-4B

Paper for BrightGuo/ProcObject-Qwen3-VL-4B