diffcider-browser
A typed-decision model made with sysone: the encoder dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 (lora) under a 0-layer yesno head, trained on osunlp/Multimodal-Mind2Web at 1b4c6a8c.
The adapter makes the masked diffusion language model Qwen3-0.6B-diffusion-mdlm-v0.1 answer a browser agent's questions, in the way of laya-ultrafast: at each step, which operation comes next (CLICK, TYPE_TEXT, SELECT, or a control such as DONE) and, for an operation, which element of the page it acts on, in one forward pass, reading the model's own Yes against its No at a mask after each option. When the agent types, the same model writes the text with the adapter switched off (dec.generate). It trained on 1,500 steps of Multimodal-Mind2Web's train split, read without the screenshots, with every action offered on every step, so that the actions a step offers never give its answer away, and with the rare operations oversampled to equal shares, since 85% of the steps are clicks. On 300 steps from 9 websites it never saw, each offering the actions its candidate elements allow, it picks the right operation 85% of the time, where always clicking is right 78% of the time and choosing the rarest action offered 80%; the right element among up to 45 65% of the time (25% untrained); and the whole step, the typed text included, 47% of the time. Offered every action on every step, it leans toward typing and selecting (57% right), having trained on the operations in equal shares: offer it the actions the page's elements allow. With the adapter off, on 120 typing steps whose value the task's wording contains, the model writes exactly what the person typed 49% of the time (8 unmasking steps, 372 ms on a T4). It has never seen a DONE page: Mind2Web records only the steps people took. An earlier version, in this repo's history, trained on the steps as they come and always answered CLICK. sysone's mdlm module loads the checkpoint's a2d-qwen3 model type with its own classes, so no remote code runs.
The repo holds the head and a LoRA adapter, not the encoder's weights: Decider.load downloads dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 at commit c8d24a3f4a from the Hub and puts the adapter on it.
Use it
Install sysone from GitHub, in Python 3.11 or newer:
pip install git+https://github.com/sgaseretto/sysonelib
A state is whatever the decision is about, as JSON-like data. Each question has a type: a choice among named options, a score on an ordered scale, or a noul, a statement that is true or false. The options are named when asking, so they can be new ones.
from sysone.inference import Decider
dec = Decider.load("sgaseretto/diffcider-browser")
from sysone.datasets import CONTROLS, NEXT_ACTION, OPERATIONS, TARGET
goal = "Find one-way flights from New York to Chicago on April 12 for one adult."
state = {
"page": {"url": "https://united.com/", "title": "united",
"text": "Book Flight status Check-in My trips Roundtrip One-way From New York To Depart Travelers 1 Adult Find flights"},
"recent_actions": [{"action": "[radio] One-way", "kind": "click", "text": None, "page_changed": True},
{"action": "[textbox] From", "kind": "fill", "text": "New York", "page_changed": False},
{"action": "[textbox] To", "kind": "click", "text": None, "page_changed": False}],
}
fields = {"1": "[1] From (textbox) = 'New York'", "2": "[2] To (textbox)", "3": "[3] Depart (textbox)"}
buttons = {"4": "[4] Roundtrip (radio) checked=False", "5": "[5] One-way (radio) checked=True", "6": "[6] Find flights (button)"}
def target(op, elements): return {"type": "choice", "instructions": {"goal": goal, "operation": op, "rules": [NEXT_ACTION, TARGET]}, "criteria": elements}
questions = {
"operation": {"type": "choice", "instructions": {"goal": goal, "rules": NEXT_ACTION},
"criteria": {k: OPERATIONS[k] for k in ("CLICK", "TYPE_TEXT")} | CONTROLS},
"type_text_target": target("TYPE_TEXT", fields),
"click_target": target("CLICK", buttons),
}
answers = dec.predict(state, questions) # System One: the next operation, and an element for each, in one pass with the adapter
# on this made-up page it hesitates, CLICK 0.51 against TYPE_TEXT 0.49, and would type into To (0.87)
field = fields[answers["type_text_target"]["choice"]].split("] ", 1)[1].split(" = ")[0]
text = dec.generate(f"The user's goal: {goal}\nThe page: {state['page']['title']}\nThe field: {field}\n"
"What should be typed into this field?", max_new_tokens=16, steps=8,
system="You fill in web forms for a user. Answer with the exact text to type, nothing else.")
# System Two: the text for that field, written by the same model without the adapter
print(text) # Chicago
answers holds an answer per question, in the Jev schema: a choice names the likeliest option and gives every option's probability, a score gives its expected level on the scale (0 for the first) and every level's probability, and a noul gives the probability that its statement is true; each says how confident it is. This model's answers to the example:
{
"operation": {
"type": "choice",
"choice": "CLICK",
"probabilities": {
"CLICK": 0.5107,
"TYPE_TEXT": 0.4856,
"WAIT": 0.0009,
"SCROLL_DOWN": 0.0007,
"SCROLL_UP": 0.0007,
"DONE": 0.0007,
"BLOCKED": 0.0007,
},
"confidence": 0.6297,
"answer_confidence": 0.5107,
},
"type_text_target": {
"type": "choice",
"choice": "2",
"probabilities": {"1": 0.0555, "2": 0.8736, "3": 0.0709},
"confidence": 0.5758,
"answer_confidence": 0.8736,
},
"click_target": {
"type": "choice",
"choice": "5",
"probabilities": {"4": 0.0165, "5": 0.6087, "6": 0.3748},
"confidence": 0.3285,
"answer_confidence": 0.6087,
},
}
Results
On osunlp/Multimodal-Mind2Web's test_website split (600 decisions), its first 300 steps, from 9 websites none of the training steps came from, two decisions each, each step offering the actions its candidate elements allow (the last metric: every action offered); calibrated, on a Kaggle T4:
| metric | value |
|---|---|
| accuracy_operation | 0.8533 |
| accuracy_element | 0.6467 |
| ece_operation | 0.0542 |
| ece_element | 0.0826 |
| accuracy_step | 0.47 |
| accuracy_operation_every_action_offered | 0.5733 |
Calibration temperatures: choice 1.3, choice:11+ 1.27, choice:2 5, choice:6-10 1.36.
Training
| Data | osunlp/Multimodal-Mind2Web at 1b4c6a8c; 1255 train, 89 valid, 156 calib cases |
| Encoder | dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1, lora (LoRA r=16, α=32, on down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj) |
| Head | 0 layers, yesno scorer, readout anchor |
| Loss | soft_ce_rps=0.25, preset t4, seed 0 |
| Fit 1 | 1 epoch, cosine, lr 0.001 (encoder 0.0002), batch 4×4, fp16 on cuda: 150 steps in 89.4 min, peak memory 7.8 GB, final loss 1.129 |
| Machine | Intel(R) Xeon(R) CPU @ 2.00GHz, 4 cores, 31.3 GB; accelerator cuda (Tesla T4, Tesla T4); a kaggle run (job mdlm-browser-e4) |
Reproduce
- sysone:
sysonelibat commit123ad8bac6(main), with uncommitted changes in 30 files (diff SHA-256b81003e7b8e8) - Ran:
run_f.py - Python 3.13.15, sysone 0.3.0, torch 2.11.0+cu128, transformers 5.16.1, peft 0.20.0, accelerate 1.14.0, datasets 4.8.5, huggingface_hub 1.29.0, tokenizers 0.23.1, safetensors 0.8.0, numpy 2.1.3, fastcore 2.2.32, plum-dispatch 2.10.1;
environment.txtlists every package
Model tree for sgaseretto/diffcider-browser
Base model
Qwen/Qwen3-0.6B-BaseDataset used to train sgaseretto/diffcider-browser
Evaluation results
- accuracy_operation on Multimodal-Mind2Webself-reported0.853
- accuracy_element on Multimodal-Mind2Webself-reported0.647
- ece_operation on Multimodal-Mind2Webself-reported0.054
- ece_element on Multimodal-Mind2Webself-reported0.083
- accuracy_step on Multimodal-Mind2Webself-reported0.470
- accuracy_operation_every_action_offered on Multimodal-Mind2Webself-reported0.573