Instructions to use qyoo/NVRM-world with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use qyoo/NVRM-world with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="qyoo/NVRM-world")# Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("qyoo/NVRM-world") model = AutoModel.from_pretrained("qyoo/NVRM-world", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Access status: canonical NVRM release artifact for WoRL. Repository visibility controls access only and does not change the Apache-2.0 terms.
NVRM-World is a multimodal scalar reward model for ranking candidate video continuations during reinforcement learning of long-horizon world models. It is designed to judge a new continuation in the context of the generated video history and provide a relative preference signal to the WoRL policy update.
This repository is an inference release built on Qwen3-VL-4B-Instruct. The training adapter and reviewed visual overrides are merged into the backbone; the one-dimensional reward head remains a small, separately serialized tensor.
At a glance
| Field | Value |
|---|---|
| Model type | Multimodal scalar video reward model |
| Backbone | Qwen/Qwen3-VL-4B-Instruct |
| Pinned base revision | ebb281ec70b05090aa6165b016eac8ec08e71b17 |
| Architecture | 36-layer Qwen3-VL backbone + scalar reward head |
| Training adapter | PEFT LoRA rank 64, alpha 128; merged for inference |
| Precision | BF16 model and raw reward logit |
| Output | One unnormalized last-token scalar |
| Thinking / hidden-state normalization | Disabled / disabled |
NVRM scores are relative and distribution-dependent, not calibrated quality probabilities. Within a comparable candidate group, a higher raw logit indicates stronger preference under the model's training distribution.
How NVRM is trained
NVRM receives text, preceding video context, and a current candidate continuation. Pairwise preferences supervise a Bradley–Terry objective: the higher-ranked continuation should receive the larger scalar reward. WoRL uses graph-cleaned preference pairs so local video comparisons can be organized into consistent training signals without requiring an expert score for every clip.
The qualitative examples at the top of this card illustrate the intended ranking behavior: continuations that better preserve objects, scene geometry, and long-horizon consistency are ranked above alternatives with accumulated visual or physical failures.
Released inference contract
The release bundle pins the exact preprocessing and output semantics used by the validated WoRL runtime:
| Stage | Contract |
|---|---|
| Video duration | One 5.0-second continuation per request |
| Temporal sampling | 1 FPS |
| Spatial preprocessing | Resize to height 360 |
| Frame encoding | JPEG quality 85 |
| Image-token bounds | 1,024–76,800 pixels |
| Message formatting | Fixed single-video chat grammar; thinking disabled |
| Output | One raw BF16 logit before groupwise transformation |
Changing preprocessing, chat formatting, dtype, or the base revision changes the evaluated model contract.
Repository contents
| Path | Description |
|---|---|
model-*.safetensors |
Three indexed shards containing the merged Qwen3-VL backbone |
nvrm_head.safetensors |
Scalar head containing rm_head.weight |
model.safetensors.index.json |
Weight-to-shard index |
| Tokenizer / processor files | Local text, image, video, and chat-template assets |
worl-bundle.json |
SHA-256 inventory and pinned inference metadata |
LICENSE.txt, NOTICE.txt |
Release terms and attribution |
The source adapter, optimizer, scheduler, RNG state, trainer state, and conversion workspace are not part of the inference release.
Download and runtime
Download an immutable private snapshot with the qyoo-authorized token and pin the resolved Hub revision in the caller's artifact lock:
hf download qyoo/NVRM-world \
--revision <full-hub-commit> \
--local-dir artifacts/NVRM-world
Verify worl-bundle.json before allocating model weights. The validated
software profile uses Transformers 4.57.6, Accelerate 1.13.0,
FlashAttention 2.7.4.post1, qwen-vl-utils 0.0.14, Pillow 12.3.0, and
opencv-python-headless 4.11.0.86. Do not enable trust_remote_code; the
repository contains no custom remote-code mapping.
Validation
The source trainer reported the following measurements on its internal held-out
reward_model_v3_3d_filtered slice. They characterize that validation slice
and are not a public benchmark.
| Metric | Value |
|---|---|
| Examples | 2,204 |
| Three-class accuracy | 0.5285843920 |
| Binary accuracy | 0.7055353902 |
| Evaluation loss | 1.48802918 |
| Mean chosen score | 0.0406154 |
| Mean rejected score | -2.6530874 |
| Mean chosen-minus-rejected margin | 2.6937027 |
Release conversion was also checked with a protected content-addressed
inference fixture. The WoRL runtime matched every processed-input tensor and
the raw BF16 output tensor under schema worl.reward-inference-golden.v1.
Golden digest:
d8aada44a467b9cdefa7d7363240edbeef9d54e38abbd62e2f2f4f518867f507.
The validated CUDA run peaked at approximately 9.33 GB allocated memory. This
fixture validates conversion and runtime wiring, not general reward quality.
Intended use
- Ranking camera-conditioned candidate continuations inside the WoRL reward pipeline.
- Reproducing the scalar video-reward component under the pinned input and software contract.
- Research on relative video preference modeling and long-horizon consistency.
NVRM-World is not a content-safety classifier, factuality detector, calibrated quality score, or substitute for human review. It is not intended for medical, legal, employment, credit, surveillance, biometric, autonomous-control, or other safety- or rights-critical decisions.
Limitations, bias, and safety
- Scores can reward superficial correlations, miss temporal or physical failures, and be vulnerable to reward hacking or adversarial inputs.
- Behavior outside the documented 5-second sampling profile is not characterized by this release.
- The model inherits biases from Qwen3-VL, candidate generators, preference data, filtering, and annotation. No demographic fairness audit is reported.
- A scalar logit is not an explanation of which frames or objects caused a ranking decision.
- The release has not been evaluated for memorization, membership inference, or sensitive-information extraction. Do not submit unauthorized personal or confidential media.
All weights use safetensors and the repository excludes executable code and pickle checkpoints. Consumers should still pin a commit, verify the bundle, apply least-privilege access, bound media inputs and GPU allocations, inspect ranked rollouts, and retain human oversight.
License
The model files and documentation are licensed under the
Apache License, Version 2.0. Retain NOTICE.txt when
redistributing this release or a derivative. The merged
Qwen/Qwen3-VL-4B-Instruct backbone is also distributed under Apache-2.0.
- Downloads last month
- 15
Model tree for qyoo/NVRM-world
Base model
Qwen/Qwen3-VL-4B-Instruct