OpenJev-4B

Jun Huang*, Xin Ren* · University of Electronic Science and Technology of China · *Equal contribution

OpenJev (project wev) is an open-source, local alternative to Jev's System One decision API: same request shape, your own GPU, no API key.

A local decision model: typed questions in, calibrated probabilities out, in one forward pass. wev-4b answers the POST /v1/systemone request shape (choice, yes/no and score questions over a free-form state), for general decisions and for browser-agent steps (which operation? which element?). It runs on your own GPU: no API key, no per-call cost, nothing generated.

Code, training and evaluation: alanhuangyoo/OpenJev. Paper: doi:10.5281/zenodo.22941164. Independent project; not affiliated with TypeSafe AI.

Quickstart

pip install "wev-ai[serve]"   # or: pip install "openjev-ai[serve]"
import wev
m = wev.load("alanhuangya/OpenJev-4B")
out = m.predict(
    state="Refund request: order #4411 arrived damaged, customer attached photos, first refund this year.",
    questions={
        "action": {"type": "choice", "instructions": "What should support do?",
                   "criteria": {"refund": "Refund the order.", "replace": "Ship a replacement.",
                                "escalate": "Send to a human agent."}},
        "fraud_risk": {"type": "noul", "instructions": "This request looks fraudulent.",
                       "criteria": {"true": "Likely fraud.", "false": "No sign of fraud."}},
    },
)
print(out["answers"])
wev serve --model alanhuangya/OpenJev-4B --port 8009   # drop-in POST /v1/systemone, e.g. for jev-ultrafast

Results

Test splits, held out from training; every other model was run on the same requests and scored the same way (per-question accuracy; scripts/compare.py). This release was read on test once, and its calibrated export once more (the calibration is fitted on development rows and changes no answer).

General typed decisions

model kev decision-v7 kev transfer-v4 typed-decisions
wev-4b 88.1 81.4 78.8
Kev-4B 88.2 82.1 65.1
Kev-8B 88.1 76.8 62.7
Laya (typed-decisions) 65.7 62.8 76.8
Laya 64.3 63.7 36.2

wev-4b trains on 80% of the typed-decisions train split, like the Laya (typed-decisions) specialist; Kev and Laya do not, so on that column they are generalists. kev decision-v7 is Kev's own training suite (wev-4b also trains on its train split); transfer-v4 is out-of-domain for every model here.

Browser steps (Mind2Web test split: websites unseen in training, jev-ultrafast request format; step success = operation and target element both right)

model step success operation
wev-4b 78.8 92.0
Kev-4B 21.2 35.7
Kev-8B 19.0 73.3
Laya (typed-decisions) 0.7 13.1
Laya 0.0 2.5

873 requests; 12 exceed the context wev-4b is evaluated with and count as wrong for it.

NNetNav test split (live-web steps, DONE judged by an LLM): step success 66.8, DONE recall 87.4, premature DONE 7.8.

End to end (153 held-out tasks on live websites, run by jev-ultrafast with wev-4b as its System One; success = the agent says DONE and an LLM judge reading the final page agrees): 39–41/153 tasks (two judge passes), vs 27/153 for the qwen3-max teacher behind the same agent and 47/153 for the GLM-5.3-Flash teacher. Live sites differ from run to run; treat gaps of a few tasks as noise.

Model

  • Backbone: Qwen/Qwen3.5-4B-Base without its vocabulary head, LoRA r=16 on every attention, Gated DeltaNet and MLP projection, merged into the weights of this export; 32 layers, bf16.
  • Readout: a pointer head scores each option's </opt> state against the question's <decide> state.
  • Each question sees the state and itself only (positions restart after the state), so the state is encoded once per request however many questions it asks, and answers never depend on question order.
  • Probabilities are calibrated: serving divides the logits by a temperature of 1.5599 fitted on development rows, which leaves every answer unchanged.
  • Median latency on one RTX 5090 (bf16, one request at a time): 40 ms for a general decision, 184 ms for a browser step.
  • Context: state up to 4096 tokens, each question up to 8192 tokens (trained with 2048); longer page states are shrunk before encoding.

Training

1 epoch, lr 3e-05, one-cycle schedule, soft-label cross-entropy where the source has soft labels; the adapter starts from jaredpalmer/kev-4b. Recipe and data builders: alanhuangyoo/OpenJev.

source license what it adds
Mind2Web CC BY 4.0 human browser steps: click, type, select
NNetNav-live Apache-2.0 live-web steps; DONE relabelled by an LLM judge
teacher episodes outputs of qwen3-max jev-ultrafast on live sites with qwen3-max as System One, success judge-verified
GLM teacher episodes outputs of GLM-5.3-Flash the same collection with GLM-5.3-Flash as System One, success judge-verified
kev decision-v7 per source ten public classification / QA sources plus rule records
typed-decisions Apache-2.0 agent / ops workflows, 5 questions per case (80% of train)
tasksource-jev mixed (per source task; some research-only) hundreds of classification tasks as decisions
jev-distill-corpus-v3 Apache-2.0 synthetic operational scenarios, soft labels
typed-decisions-synth MIT multi-question cases over 149 domains

Use terms. Some training data carries its own terms: several tasksource-jev source tasks are research-only, and the teacher episodes are qwen3-max and GLM-5.3-Flash outputs subject to their providers' terms. Treat this model as a research artifact and check those terms before any commercial use.

Limitations

  • English only. Decisions, not text: TYPE values come from a separate text model, as in jev-ultrafast.
  • Browser targets are scored among the candidates the agent lists (8–40 per step), not every element on the page.
  • DONE and BLOCKED are the hardest operations; gate DONE on its probability when early stops are costly.
  • Not compared with Jev itself (no API access).

Citation

Jun Huang and Xin Ren contributed equally (University of Electronic Science and Technology of China).

@misc{huang2026wev,
  title     = {wev: Distilling LLM Browser Agents into Open, Local System-One Decision Models},
  author    = {Huang, Jun and Ren, Xin},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22941164},
  url       = {https://doi.org/10.5281/zenodo.22941164}
}

License

Apache-2.0, like the base model. Architecture code adapted from kev (Apache-2.0); weights initialised from jaredpalmer/kev-4b (Apache-2.0).

Downloads last month
76
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alanhuangya/OpenJev-4B

Adapter
(3)
this model

Datasets used to train alanhuangya/OpenJev-4B

Space using alanhuangya/OpenJev-4B 1

Collection including alanhuangya/OpenJev-4B