dLLM PRM Gap · Outcome reward model

News

  • 2026-09: Accepted at NeurIPS 2026.

Public adapter release for dLLM PRM Gap. bidirectional final-state scoring for the matched-compute ORM Rerank baseline

This repository contains only trainable adapter parameters and the reward head. It does not include the base model. The release configuration is recorded in config.json; the arXiv paper defines the release scope and citation.

Files

  • adapter.safetensors — compact adapter weights
  • config.json — public base-model id, architecture, and provenance

Load

from huggingface_hub import snapshot_download
from prm.checkpointing import load_diffusion_prm

path = snapshot_download("YanZhanPKU/dLLM-PRM-Gap-orm-dream7b")
model, report = load_diffusion_prm(
    checkpoint=path,
    model_path="Dream-org/Dream-v0-Instruct-7B",
    local_files_only=False,
)

Role: final-state outcome scorer for ORM Rerank@N.

Paper: https://arxiv.org/abs/2609.35472.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YanZhanPKU/dLLM-PRM-Gap-orm-dream7b

Adapter
(29)
this model

Collection including YanZhanPKU/dLLM-PRM-Gap-orm-dream7b

Paper for YanZhanPKU/dLLM-PRM-Gap-orm-dream7b