1. Introduction
OpenJev-4B is a decision model built from the text backbone of Qwen3.5-4B. Given a state, a question, and a request-defined set of natural-language options, it directly predicts a probability distribution over those options. It supports selection, binary judgments, and rubric-based scoring through the same probabilistic interface.
The model combines a language-model backbone with a lightweight Option Set Interactor and a shared decision head. Its output dimension follows the candidate set supplied with each request: option names and optional criteria are inputs, rather than classes fixed in the model parameters. Decision probabilities are computed without autoregressive answer generation.
OpenJev-4B is post-trained on OpenJevData-140k in two stages: supervised fine-tuning (SFT) followed by REINFORCE-Analysis (REINFORCE-A). SFT establishes decision behavior using predominantly hard labels. RL increases the share of soft-target contexts and refines probability quality through sampled outcome feedback.
2. Model Summary
OpenJev-4B brings together the following architecture and training features:
- Open-Vocabulary Decisions: Accepts request-defined options and criteria, returning explicit probabilities for selection, binary judgments and rubric-based scoring.
- Permutation Equivariance: Shared option encoders and an order-independent decision head make predictions follow the options when they are reordered, up to numerical precision.
- Cross-Option Understanding: A lightweight Option Set Interactor models semantic relationships across candidates, including overlap, competition and references to other options.
- Direct, Parallel Inference: Computes decision probabilities without autoregressive answer generation, with parallel option branches and shared-prefix reuse for attention caches and recurrent states.
- Two-Stage Post-Training: Full-parameter SFT → REINFORCE-Analysis combines predominantly hard-label supervision with a larger share of soft-target contexts during RL to refine probability quality.
- Efficient Outcome-Based RL: REINFORCE-Analysis samples 16 candidate actions with 10% uniform exploration and learns from binary outcome feedback. Analytic computation, importance correction and a leave-one-out baseline reduce reliance on sampled estimates.
- Replay and Reference Regularization: A fixed 25% of SFT contexts are replayed with the same RL objective. A frozen SFT reference provides KL regularization to help retain learned behavior.
Training data: OpenJevData-140k.
3. Evaluation Results
- Accuracy (↑): the percentage of hard-label examples for which the highest-probability option matches the gold answer. Higher is better; unsupported or failed predictions count as incorrect.
- Brier (↓): the average sum of squared differences between predicted and gold option probabilities. Lower is better, with 0 indicating a perfect match. For hard labels, the gold distribution is one-hot; for soft labels, it is the provided target distribution.
We compare OpenJev-4B with Jev, Decider 4B, JevK5 4B, and Intern-Decision-4B on MMDM and JevBench Public.
| Task type | OpenJev-4B | Jev1 | Decider 4B3 | JevK5 4B3 | Intern- |
|---|---|---|---|---|---|
| MMDM · Hard · Accuracy (%) ↑ / Brier ↓4 | |||||
| Answer preference n = 185 |
90.81/0.1236 | 88.11/0.1794 | 85.95/0.2263 | 77.84/0.3276 | 77.84/0.3142 |
| Code / database n = 122 |
84.43/0.2128 | 83.61/0.2189 | 77.05/0.3293 | 72.95/0.3546 | 68.03/0.3364 |
| Commonsense context n = 355 |
93.80/0.0854 | 96.90/0.0409 | 95.21/0.0687 | 93.80/0.1073 | 87.04/0.2096 |
| Constraint planning n = 334 |
76.05/0.2944 | 75.75/0.2859 | 71.26/0.3685 | 64.37/0.4816 | 58.38/0.5222 |
| Deductive reasoning n = 619 |
90.47/0.1389 | 94.51/0.0781 | 86.27/0.1861 | 79.97/0.2809 | 68.66/0.3848 |
| Emotion recognition n = 9 |
100.00/0.0243 | 77.78/0.2353 | 100.00/0.0684 | 77.78/0.3633 | 77.78/0.3690 |
| Evidence validation n = 177 |
91.53/0.1221 | 90.96/0.1211 | 85.31/0.2103 | 88.14/0.1960 | 84.18/0.2301 |
| Intent routing n = 253 |
94.07/0.1045 | 93.28/0.1047 | 92.49/0.1223 | 87.75/0.1688 | 64.82/0.2352 |
| Knowledge QA n = 572 |
95.63/0.0693 | 98.95/0.0202 | 96.85/0.0527 | 94.76/0.0847 | 88.99/0.1722 |
| Multi-hop reasoning n = 211 |
92.42/0.0992 | 89.57/0.1274 | 73.93/0.3978 | 74.41/0.3046 | 64.45/0.4399 |
| Option logic n = 287 |
51.92/0.5694 | 41.11/0.6551 | 34.84/0.7715 | 35.19/0.7176 | 28.22/0.7979 |
| Policy exceptions n = 576 |
95.31/0.0688 | 90.97/0.1301 | 82.81/0.2620 | 72.74/0.3858 | 63.89/0.4630 |
| Quantitative reasoning n = 330 |
69.09/0.3752 | 77.88/0.2841 | 51.21/0.6098 | 46.36/0.6636 | 37.58/0.7533 |
| Reading comprehension n = 48 |
95.83/0.0523 | 100.00/0.0091 | 95.83/0.0640 | 91.67/0.1074 | 89.58/0.1868 |
| Score rubrics n = 135 |
78.52/0.2939 | 85.93/0.2055 | 72.59/0.3487 | 73.33/0.3190 | 69.63/0.3827 |
| Temporal / numeric n = 138 |
55.07/0.4695 | 64.49/0.4235 | 54.35/0.5197 | 53.62/0.5379 | 59.42/0.5422 |
| Total n = 4,351 |
85.57/0.1864 | 86.37/0.1733 | 78.88/0.2832 | 74.70/0.3330 | 66.95/0.4069 |
| MMDM · Soft · Brier ↓4 | |||||
| Constraint planning n = 1 |
0.0070 | 0.1412 | 0.1338 | 0.0251 | 0.0258 |
| Evidence validation n = 22 |
0.1066 | 0.0493 | 0.1517 | 0.0773 | 0.0914 |
| Intent routing n = 22 |
0.0681 | 0.1570 | 0.1366 | 0.1304 | 0.0838 |
| Policy exceptions n = 18 |
0.0423 | 0.0606 | 0.1128 | 0.0916 | 0.0674 |
| Probabilistic reasoning n = 505 |
0.0278 | 0.1164 | 0.1387 | 0.1092 | 0.1242 |
| Score rubrics n = 30 |
0.0864 | 0.0895 | 0.1653 | 0.1242 | 0.1746 |
| Temporal / numeric n = 12 |
0.2168 | 0.0822 | 0.1100 | 0.1283 | 0.1747 |
| Total n = 610 |
0.0391 | 0.1119 | 0.1391 | 0.1092 | 0.1232 |
| JevBench · Public · Accuracy (%) ↑ / Brier ↓5 | |||||
| Original n = 72 |
100.00/0.0312 | 98.61/0.0283 | 98.61/0.0574 | 94.44/0.1075 | 98.61/0.0340 |
| Easy n = 48 |
100.00/0.000011 | 100.00/0.000196 | 100.00/0.000045 | 100.00/0.002509 | 100.00/0.001119 |
| Hard n = 111 |
75.68/0.3469 | 72.07/0.3594 | 64.86/0.4686 | 77.48/0.3118 | 73.87/0.3439 |
| Total n = 231 |
88.31/0.1764 | 86.15/0.1816 | 82.68/0.2431 | 87.45/0.1839 | 87.01/0.1761 |
- Jev: official API, evaluated as Jev 1.13.0.
- Intern-Decision-4B: official inference code and temperature 1.99241824, BF16, SDPA, no input truncation, in the existing Torch 2.11.0 / Transformers 5.17.0 environment. Its option-count limit prevents scoring 57 MMDM hard examples; these count as incorrect in accuracy and are excluded from the common probability-metric cohort.
- External model versions: Decider 4B v2.1 and JevK5 4B v0.3. Their returned decision distributions are used for probability metrics. OpenJev-4B is our final release model; earlier OpenJev checkpoints are omitted.
- MMDM: the published test split contains 4,961 decisions: 4,351 hard and 610 soft. Accuracy uses all hard examples. Brier uses the same 4,294 common valid hard examples across models (196 of 253 for intent routing), and all 610 soft examples, matching the dataset card. All models' totals use the corresponding scoring counts. Soft rows show Brier only. Bold marks the best value for each metric; zero probabilities are not clipped.
- JevBench Public: local evaluations of the v1.3.0 public suite: Original (72), Easy (48), Hard (111). Accuracy and Brier follow its gold-label scoring protocol; all five models return valid predictions for all 231 examples. “Hard” names a difficulty cohort here, rather than a target type. These results are separate from the official private leaderboard. Easy Brier is shown to six decimal places to retain the small differences.
- Evaluation context: the published MMDM split combines the previously separate development and test sets, so these numbers are not a new blind evaluation. OpenJev training includes newly constructed scenarios informed by public decision-task patterns. The tables report performance on these fixed evaluation collections.
Machine-readable values and evaluation counts are included in evaluation.json.
4. Model Weights and Code
The complete inference and two-stage training implementation is available at ejhshen/OpenJev. This model repository contains the text backbone, decision head, tokenizer, model configuration and calibration settings. Load them together using the OpenJev runtime.
Run a decision
Install the GitHub repository in a compatible CUDA environment, then run:
from openjev.runtime.engine import DecisionEngine
engine = DecisionEngine("shenjunhao/OpenJev-4B", execution_mode="expanded")
result = engine.predict({
"state": "A customer recognizes a purchase but reports being charged twice.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"Duplicate payment": "The same purchase was charged more than once.",
"Unrecognized payment": "The customer does not recognize the purchase."
}
}
}
})
print(result["answers"]["route"]["choice"])
print(result["answers"]["route"]["probabilities"])
The model accepts a shared state and a questions object. For choice, each question supplies instructions and a mapping from candidate names to criteria; candidate names are defined by the request. It returns the selected candidate and its full probability distribution. The same interface supports noul for binary judgments and score for expected rubric-level scores. See the input/output format and HTTP example for details.
The runtime downloads the complete artifact from this repository or accepts a local directory. The release uses temperature 1.0, matching the evaluation results. Decision outputs are computed directly, without generating explanation text.
Train SFT and REINFORCE-A
Training data is available at shenjunhao/OpenJevData-140k. The repository includes data preparation, the optimized distributed training engine, sample replay, resumable caches, checkpointing and model export. To fine-tune OpenJev-4B, set up the official verl/SGLang environment described in the training guide, then load the complete released model:
hf download shenjunhao/OpenJev-4B --local-dir work/base-model
python -m openjev.cli validate-artifact work/base-model
python -m openjev.training.prepare_data \
--dataset shenjunhao/OpenJevData-140k \
--mandatory-replay training/mandatory-replay.json --output work/data
NPROC=8 bash training/run_two_stage.sh
The pipeline runs SFT, exports the SFT artifact, caches frozen reference probabilities, runs REINFORCE-A and exports the final decision model. Both stages use global batch 1,024 by default; their step counts are calculated from the prepared sample counts. Individual stage launchers and resume instructions are included in the training guide.
