OpenJev-4B

Learning Open-Vocabulary Probabilistic Decision Models from Pretrained Language Models

GitHub repository

1. Introduction

OpenJev-4B is a decision model built from the text backbone of Qwen3.5-4B. Given a state, a question, and a request-defined set of natural-language options, it directly predicts a probability distribution over those options. It supports selection, binary judgments, and rubric-based scoring through the same probabilistic interface.

The model combines a language-model backbone with a lightweight Option Set Interactor and a shared decision head. Its output dimension follows the candidate set supplied with each request: option names and optional criteria are inputs, rather than classes fixed in the model parameters. Decision probabilities are computed without autoregressive answer generation.

OpenJev-4B is post-trained on OpenJevData-140k in two stages: supervised fine-tuning (SFT) followed by REINFORCE-Analysis (REINFORCE-A). SFT establishes decision behavior using predominantly hard labels. RL increases the share of soft-target contexts and refines probability quality through sampled outcome feedback.

2. Model Summary

OpenJev architecture: independent option branches, a two-layer Option Set Interactor, and question-option compatibility scoring

OpenJev-4B brings together the following architecture and training features:

  • Open-Vocabulary Decisions: Accepts request-defined options and criteria, returning explicit probabilities for selection, binary judgments and rubric-based scoring.
  • Permutation Equivariance: Shared option encoders and an order-independent decision head make predictions follow the options when they are reordered, up to numerical precision.
  • Cross-Option Understanding: A lightweight Option Set Interactor models semantic relationships across candidates, including overlap, competition and references to other options.
  • Direct, Parallel Inference: Computes decision probabilities without autoregressive answer generation, with parallel option branches and shared-prefix reuse for attention caches and recurrent states.
  • Two-Stage Post-Training: Full-parameter SFT → REINFORCE-Analysis combines predominantly hard-label supervision with a larger share of soft-target contexts during RL to refine probability quality.
  • Efficient Outcome-Based RL: REINFORCE-Analysis samples 16 candidate actions with 10% uniform exploration and learns from binary outcome feedback. Analytic computation, importance correction and a leave-one-out baseline reduce reliance on sampled estimates.
  • Replay and Reference Regularization: A fixed 25% of SFT contexts are replayed with the same RL objective. A frozen SFT reference provides KL regularization to help retain learned behavior.

Training data: OpenJevData-140k.

3. Evaluation Results

  1. Accuracy (↑): the percentage of hard-label examples for which the highest-probability option matches the gold answer. Higher is better; unsupported or failed predictions count as incorrect.
  2. Brier (↓): the average sum of squared differences between predicted and gold option probabilities. Lower is better, with 0 indicating a perfect match. For hard labels, the gold distribution is one-hot; for soft labels, it is the provided target distribution.

We compare OpenJev-4B with Jev, Decider 4B, JevK5 4B, and Intern-Decision-4B on MMDM and JevBench Public.

Task type OpenJev-4B Jev1 Decider 4B3 JevK5 4B3 Intern-Decision-4B2
MMDM · Hard · Accuracy (%) ↑ / Brier ↓4
Answer preference
n = 185
90.81/0.1236 88.11/0.1794 85.95/0.2263 77.84/0.3276 77.84/0.3142
Code / database
n = 122
84.43/0.2128 83.61/0.2189 77.05/0.3293 72.95/0.3546 68.03/0.3364
Commonsense context
n = 355
93.80/0.0854 96.90/0.0409 95.21/0.0687 93.80/0.1073 87.04/0.2096
Constraint planning
n = 334
76.05/0.2944 75.75/0.2859 71.26/0.3685 64.37/0.4816 58.38/0.5222
Deductive reasoning
n = 619
90.47/0.1389 94.51/0.0781 86.27/0.1861 79.97/0.2809 68.66/0.3848
Emotion recognition
n = 9
100.00/0.0243 77.78/0.2353 100.00/0.0684 77.78/0.3633 77.78/0.3690
Evidence validation
n = 177
91.53/0.1221 90.96/0.1211 85.31/0.2103 88.14/0.1960 84.18/0.2301
Intent routing
n = 253
94.07/0.1045 93.28/0.1047 92.49/0.1223 87.75/0.1688 64.82/0.2352
Knowledge QA
n = 572
95.63/0.0693 98.95/0.0202 96.85/0.0527 94.76/0.0847 88.99/0.1722
Multi-hop reasoning
n = 211
92.42/0.0992 89.57/0.1274 73.93/0.3978 74.41/0.3046 64.45/0.4399
Option logic
n = 287
51.92/0.5694 41.11/0.6551 34.84/0.7715 35.19/0.7176 28.22/0.7979
Policy exceptions
n = 576
95.31/0.0688 90.97/0.1301 82.81/0.2620 72.74/0.3858 63.89/0.4630
Quantitative reasoning
n = 330
69.09/0.3752 77.88/0.2841 51.21/0.6098 46.36/0.6636 37.58/0.7533
Reading comprehension
n = 48
95.83/0.0523 100.00/0.0091 95.83/0.0640 91.67/0.1074 89.58/0.1868
Score rubrics
n = 135
78.52/0.2939 85.93/0.2055 72.59/0.3487 73.33/0.3190 69.63/0.3827
Temporal / numeric
n = 138
55.07/0.4695 64.49/0.4235 54.35/0.5197 53.62/0.5379 59.42/0.5422
Total
n = 4,351
85.57/0.1864 86.37/0.1733 78.88/0.2832 74.70/0.3330 66.95/0.4069
MMDM · Soft · Brier ↓4
Constraint planning
n = 1
0.0070 0.1412 0.1338 0.0251 0.0258
Evidence validation
n = 22
0.1066 0.0493 0.1517 0.0773 0.0914
Intent routing
n = 22
0.0681 0.1570 0.1366 0.1304 0.0838
Policy exceptions
n = 18
0.0423 0.0606 0.1128 0.0916 0.0674
Probabilistic reasoning
n = 505
0.0278 0.1164 0.1387 0.1092 0.1242
Score rubrics
n = 30
0.0864 0.0895 0.1653 0.1242 0.1746
Temporal / numeric
n = 12
0.2168 0.0822 0.1100 0.1283 0.1747
Total
n = 610
0.0391 0.1119 0.1391 0.1092 0.1232
JevBench · Public · Accuracy (%) ↑ / Brier ↓5
Original
n = 72
100.00/0.0312 98.61/0.0283 98.61/0.0574 94.44/0.1075 98.61/0.0340
Easy
n = 48
100.00/0.000011 100.00/0.000196 100.00/0.000045 100.00/0.002509 100.00/0.001119
Hard
n = 111
75.68/0.3469 72.07/0.3594 64.86/0.4686 77.48/0.3118 73.87/0.3439
Total
n = 231
88.31/0.1764 86.15/0.1816 82.68/0.2431 87.45/0.1839 87.01/0.1761
  1. Jev: official API, evaluated as Jev 1.13.0.
  2. Intern-Decision-4B: official inference code and temperature 1.99241824, BF16, SDPA, no input truncation, in the existing Torch 2.11.0 / Transformers 5.17.0 environment. Its option-count limit prevents scoring 57 MMDM hard examples; these count as incorrect in accuracy and are excluded from the common probability-metric cohort.
  3. External model versions: Decider 4B v2.1 and JevK5 4B v0.3. Their returned decision distributions are used for probability metrics. OpenJev-4B is our final release model; earlier OpenJev checkpoints are omitted.
  4. MMDM: the published test split contains 4,961 decisions: 4,351 hard and 610 soft. Accuracy uses all hard examples. Brier uses the same 4,294 common valid hard examples across models (196 of 253 for intent routing), and all 610 soft examples, matching the dataset card. All models' totals use the corresponding scoring counts. Soft rows show Brier only. Bold marks the best value for each metric; zero probabilities are not clipped.
  5. JevBench Public: local evaluations of the v1.3.0 public suite: Original (72), Easy (48), Hard (111). Accuracy and Brier follow its gold-label scoring protocol; all five models return valid predictions for all 231 examples. “Hard” names a difficulty cohort here, rather than a target type. These results are separate from the official private leaderboard. Easy Brier is shown to six decimal places to retain the small differences.
  6. Evaluation context: the published MMDM split combines the previously separate development and test sets, so these numbers are not a new blind evaluation. OpenJev training includes newly constructed scenarios informed by public decision-task patterns. The tables report performance on these fixed evaluation collections.

Machine-readable values and evaluation counts are included in evaluation.json.

4. Model Weights and Code

The complete inference and two-stage training implementation is available at ejhshen/OpenJev. This model repository contains the text backbone, decision head, tokenizer, model configuration and calibration settings. Load them together using the OpenJev runtime.

Run a decision

Install the GitHub repository in a compatible CUDA environment, then run:

from openjev.runtime.engine import DecisionEngine

engine = DecisionEngine("shenjunhao/OpenJev-4B", execution_mode="expanded")
result = engine.predict({
    "state": "A customer recognizes a purchase but reports being charged twice.",
    "questions": {
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "Duplicate payment": "The same purchase was charged more than once.",
                "Unrecognized payment": "The customer does not recognize the purchase."
            }
        }
    }
})
print(result["answers"]["route"]["choice"])
print(result["answers"]["route"]["probabilities"])

The model accepts a shared state and a questions object. For choice, each question supplies instructions and a mapping from candidate names to criteria; candidate names are defined by the request. It returns the selected candidate and its full probability distribution. The same interface supports noul for binary judgments and score for expected rubric-level scores. See the input/output format and HTTP example for details.

The runtime downloads the complete artifact from this repository or accepts a local directory. The release uses temperature 1.0, matching the evaluation results. Decision outputs are computed directly, without generating explanation text.

Train SFT and REINFORCE-A

Training data is available at shenjunhao/OpenJevData-140k. The repository includes data preparation, the optimized distributed training engine, sample replay, resumable caches, checkpointing and model export. To fine-tune OpenJev-4B, set up the official verl/SGLang environment described in the training guide, then load the complete released model:

hf download shenjunhao/OpenJev-4B --local-dir work/base-model
python -m openjev.cli validate-artifact work/base-model
python -m openjev.training.prepare_data \
  --dataset shenjunhao/OpenJevData-140k \
  --mandatory-replay training/mandatory-replay.json --output work/data
NPROC=8 bash training/run_two_stage.sh

The pipeline runs SFT, exports the SFT artifact, caches frozen reference probabilities, runs REINFORCE-A and exports the final decision model. Both stages use global batch 1,024 by default; their step counts are calculated from the prepared sample counts. Individual stage launchers and resume instructions are included in the training guide.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for shenjunhao/OpenJev-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(802)
this model