NVRM higher-ranked and lower-ranked long-video continuations

NVRM-World

Next Video Reward Model for long-horizon video world models

WoRL · Backbone

Access status: canonical NVRM release artifact for WoRL. Repository visibility controls access only and does not change the Apache-2.0 terms.

NVRM-World is a multimodal scalar reward model for ranking candidate video continuations during reinforcement learning of long-horizon world models. It is designed to judge a new continuation in the context of the generated video history and provide a relative preference signal to the WoRL policy update.

This repository is an inference release built on Qwen3-VL-4B-Instruct. The training adapter and reviewed visual overrides are merged into the backbone; the one-dimensional reward head remains a small, separately serialized tensor.

At a glance

Field Value
Model type Multimodal scalar video reward model
Backbone Qwen/Qwen3-VL-4B-Instruct
Pinned base revision ebb281ec70b05090aa6165b016eac8ec08e71b17
Architecture 36-layer Qwen3-VL backbone + scalar reward head
Training adapter PEFT LoRA rank 64, alpha 128; merged for inference
Precision BF16 model and raw reward logit
Output One unnormalized last-token scalar
Thinking / hidden-state normalization Disabled / disabled

NVRM scores are relative and distribution-dependent, not calibrated quality probabilities. Within a comparable candidate group, a higher raw logit indicates stronger preference under the model's training distribution.

How NVRM is trained

NVRM pairwise preference training

NVRM receives text, preceding video context, and a current candidate continuation. Pairwise preferences supervise a Bradley–Terry objective: the higher-ranked continuation should receive the larger scalar reward. WoRL uses graph-cleaned preference pairs so local video comparisons can be organized into consistent training signals without requiring an expert score for every clip.

The qualitative examples at the top of this card illustrate the intended ranking behavior: continuations that better preserve objects, scene geometry, and long-horizon consistency are ranked above alternatives with accumulated visual or physical failures.

Released inference contract

The release bundle pins the exact preprocessing and output semantics used by the validated WoRL runtime:

Stage Contract
Video duration One 5.0-second continuation per request
Temporal sampling 1 FPS
Spatial preprocessing Resize to height 360
Frame encoding JPEG quality 85
Image-token bounds 1,024–76,800 pixels
Message formatting Fixed single-video chat grammar; thinking disabled
Output One raw BF16 logit before groupwise transformation

Changing preprocessing, chat formatting, dtype, or the base revision changes the evaluated model contract.

Repository contents

Path Description
model-*.safetensors Three indexed shards containing the merged Qwen3-VL backbone
nvrm_head.safetensors Scalar head containing rm_head.weight
model.safetensors.index.json Weight-to-shard index
Tokenizer / processor files Local text, image, video, and chat-template assets
worl-bundle.json SHA-256 inventory and pinned inference metadata
LICENSE.txt, NOTICE.txt Release terms and attribution

The source adapter, optimizer, scheduler, RNG state, trainer state, and conversion workspace are not part of the inference release.

Download and runtime

Download an immutable private snapshot with the qyoo-authorized token and pin the resolved Hub revision in the caller's artifact lock:

hf download qyoo/NVRM-world \
  --revision <full-hub-commit> \
  --local-dir artifacts/NVRM-world

Verify worl-bundle.json before allocating model weights. The validated software profile uses Transformers 4.57.6, Accelerate 1.13.0, FlashAttention 2.7.4.post1, qwen-vl-utils 0.0.14, Pillow 12.3.0, and opencv-python-headless 4.11.0.86. Do not enable trust_remote_code; the repository contains no custom remote-code mapping.

Validation

The source trainer reported the following measurements on its internal held-out reward_model_v3_3d_filtered slice. They characterize that validation slice and are not a public benchmark.

Metric Value
Examples 2,204
Three-class accuracy 0.5285843920
Binary accuracy 0.7055353902
Evaluation loss 1.48802918
Mean chosen score 0.0406154
Mean rejected score -2.6530874
Mean chosen-minus-rejected margin 2.6937027

Release conversion was also checked with a protected content-addressed inference fixture. The WoRL runtime matched every processed-input tensor and the raw BF16 output tensor under schema worl.reward-inference-golden.v1. Golden digest: d8aada44a467b9cdefa7d7363240edbeef9d54e38abbd62e2f2f4f518867f507. The validated CUDA run peaked at approximately 9.33 GB allocated memory. This fixture validates conversion and runtime wiring, not general reward quality.

Intended use

  • Ranking camera-conditioned candidate continuations inside the WoRL reward pipeline.
  • Reproducing the scalar video-reward component under the pinned input and software contract.
  • Research on relative video preference modeling and long-horizon consistency.

NVRM-World is not a content-safety classifier, factuality detector, calibrated quality score, or substitute for human review. It is not intended for medical, legal, employment, credit, surveillance, biometric, autonomous-control, or other safety- or rights-critical decisions.

Limitations, bias, and safety

  • Scores can reward superficial correlations, miss temporal or physical failures, and be vulnerable to reward hacking or adversarial inputs.
  • Behavior outside the documented 5-second sampling profile is not characterized by this release.
  • The model inherits biases from Qwen3-VL, candidate generators, preference data, filtering, and annotation. No demographic fairness audit is reported.
  • A scalar logit is not an explanation of which frames or objects caused a ranking decision.
  • The release has not been evaluated for memorization, membership inference, or sensitive-information extraction. Do not submit unauthorized personal or confidential media.

All weights use safetensors and the repository excludes executable code and pickle checkpoints. Consumers should still pin a commit, verify the bundle, apply least-privilege access, bound media inputs and GPU allocations, inspect ranked rollouts, and retain human oversight.

License

The model files and documentation are licensed under the Apache License, Version 2.0. Retain NOTICE.txt when redistributing this release or a derivative. The merged Qwen/Qwen3-VL-4B-Instruct backbone is also distributed under Apache-2.0.

Downloads last month
15
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for qyoo/NVRM-world

Finetuned
(457)
this model