--- base_model: dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 datasets: - osunlp/Multimodal-Mind2Web library_name: sysone tags: - sysone - typed-decisions - text - lora model-index: - name: diffcider-browser results: - task: type: zero-shot-classification dataset: name: Multimodal-Mind2Web type: osunlp/Multimodal-Mind2Web split: test_website metrics: - type: accuracy_operation value: 0.8533 - type: accuracy_element value: 0.6467 - type: ece_operation value: 0.0542 - type: ece_element value: 0.0826 - type: accuracy_step value: 0.47 - type: accuracy_operation_every_action_offered value: 0.5733 --- # diffcider-browser A typed-decision model made with [sysone](https://github.com/sgaseretto/sysonelib): the encoder `dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1` (lora) under a 0-layer `yesno` head, trained on [osunlp/Multimodal-Mind2Web](https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web) at `1b4c6a8c`. The adapter makes the masked diffusion language model [Qwen3-0.6B-diffusion-mdlm-v0.1](https://huggingface.co/dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1) answer a browser agent's questions, in the way of [laya-ultrafast](https://github.com/cklxx/laya): at each step, which operation comes next (CLICK, TYPE_TEXT, SELECT, or a control such as DONE) and, for an operation, which element of the page it acts on, in one forward pass, reading the model's own *Yes* against its *No* at a mask after each option. When the agent types, the same model writes the text with the adapter switched off (`dec.generate`). It trained on 1,500 steps of Multimodal-Mind2Web's `train` split, read without the screenshots, with every action offered on every step, so that the actions a step offers never give its answer away, and with the rare operations oversampled to equal shares, since 85% of the steps are clicks. On 300 steps from 9 websites it never saw, each offering the actions its candidate elements allow, it picks the right operation 85% of the time, where always clicking is right 78% of the time and choosing the rarest action offered 80%; the right element among up to 45 65% of the time (25% untrained); and the whole step, the typed text included, 47% of the time. Offered every action on every step, it leans toward typing and selecting (57% right), having trained on the operations in equal shares: offer it the actions the page's elements allow. With the adapter off, on 120 typing steps whose value the task's wording contains, the model writes exactly what the person typed 49% of the time (8 unmasking steps, 372 ms on a T4). It has never seen a DONE page: Mind2Web records only the steps people took. An earlier version, in this repo's history, trained on the steps as they come and always answered CLICK. sysone's `mdlm` module loads the checkpoint's `a2d-qwen3` model type with its own classes, so no remote code runs. The repo holds the head and a LoRA adapter, not the encoder's weights: `Decider.load` downloads `dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1` at commit `c8d24a3f4a` from the Hub and puts the adapter on it. ## Use it Install sysone from GitHub, in Python 3.11 or newer: ```sh pip install git+https://github.com/sgaseretto/sysonelib ``` A state is whatever the decision is about, as JSON-like data. Each question has a type: a `choice` among named options, a `score` on an ordered scale, or a `noul`, a statement that is true or false. The options are named when asking, so they can be new ones. ```python from sysone.inference import Decider dec = Decider.load("sgaseretto/diffcider-browser") from sysone.datasets import CONTROLS, NEXT_ACTION, OPERATIONS, TARGET goal = "Find one-way flights from New York to Chicago on April 12 for one adult." state = { "page": {"url": "https://united.com/", "title": "united", "text": "Book Flight status Check-in My trips Roundtrip One-way From New York To Depart Travelers 1 Adult Find flights"}, "recent_actions": [{"action": "[radio] One-way", "kind": "click", "text": None, "page_changed": True}, {"action": "[textbox] From", "kind": "fill", "text": "New York", "page_changed": False}, {"action": "[textbox] To", "kind": "click", "text": None, "page_changed": False}], } fields = {"1": "[1] From (textbox) = 'New York'", "2": "[2] To (textbox)", "3": "[3] Depart (textbox)"} buttons = {"4": "[4] Roundtrip (radio) checked=False", "5": "[5] One-way (radio) checked=True", "6": "[6] Find flights (button)"} def target(op, elements): return {"type": "choice", "instructions": {"goal": goal, "operation": op, "rules": [NEXT_ACTION, TARGET]}, "criteria": elements} questions = { "operation": {"type": "choice", "instructions": {"goal": goal, "rules": NEXT_ACTION}, "criteria": {k: OPERATIONS[k] for k in ("CLICK", "TYPE_TEXT")} | CONTROLS}, "type_text_target": target("TYPE_TEXT", fields), "click_target": target("CLICK", buttons), } answers = dec.predict(state, questions) # System One: the next operation, and an element for each, in one pass with the adapter # on this made-up page it hesitates, CLICK 0.51 against TYPE_TEXT 0.49, and would type into To (0.87) field = fields[answers["type_text_target"]["choice"]].split("] ", 1)[1].split(" = ")[0] text = dec.generate(f"The user's goal: {goal}\nThe page: {state['page']['title']}\nThe field: {field}\n" "What should be typed into this field?", max_new_tokens=16, steps=8, system="You fill in web forms for a user. Answer with the exact text to type, nothing else.") # System Two: the text for that field, written by the same model without the adapter print(text) # Chicago ``` `answers` holds an answer per question, in the Jev schema: a choice names the likeliest option and gives every option's probability, a score gives its expected level on the scale (0 for the first) and every level's probability, and a noul gives the probability that its statement is true; each says how confident it is. This model's answers to the example: ```python { "operation": { "type": "choice", "choice": "CLICK", "probabilities": { "CLICK": 0.5107, "TYPE_TEXT": 0.4856, "WAIT": 0.0009, "SCROLL_DOWN": 0.0007, "SCROLL_UP": 0.0007, "DONE": 0.0007, "BLOCKED": 0.0007, }, "confidence": 0.6297, "answer_confidence": 0.5107, }, "type_text_target": { "type": "choice", "choice": "2", "probabilities": {"1": 0.0555, "2": 0.8736, "3": 0.0709}, "confidence": 0.5758, "answer_confidence": 0.8736, }, "click_target": { "type": "choice", "choice": "5", "probabilities": {"4": 0.0165, "5": 0.6087, "6": 0.3748}, "confidence": 0.3285, "answer_confidence": 0.6087, }, } ``` ## Results On [osunlp/Multimodal-Mind2Web](https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web)'s `test_website` split (600 decisions), its first 300 steps, from 9 websites none of the training steps came from, two decisions each, each step offering the actions its candidate elements allow (the last metric: every action offered); calibrated, on a Kaggle T4: | metric | value | |---|---| | accuracy_operation | 0.8533 | | accuracy_element | 0.6467 | | ece_operation | 0.0542 | | ece_element | 0.0826 | | accuracy_step | 0.47 | | accuracy_operation_every_action_offered | 0.5733 | Calibration temperatures: choice 1.3, choice:11+ 1.27, choice:2 5, choice:6-10 1.36. ## Training | | | |---|---| | Data | [osunlp/Multimodal-Mind2Web](https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web) at `1b4c6a8c`; 1255 train, 89 valid, 156 calib cases | | Encoder | `dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1`, lora (LoRA r=16, α=32, on `down_proj`, `gate_proj`, `k_proj`, `o_proj`, `q_proj`, `up_proj`, `v_proj`) | | Head | 0 layers, `yesno` scorer, readout `anchor` | | Loss | `soft_ce_rps=0.25`, preset `t4`, seed 0 | | Fit 1 | 1 epoch, cosine, lr 0.001 (encoder 0.0002), batch 4×4, fp16 on cuda: 150 steps in 89.4 min, peak memory 7.8 GB, final loss 1.129 | | Machine | Intel(R) Xeon(R) CPU @ 2.00GHz, 4 cores, 31.3 GB; accelerator cuda (Tesla T4, Tesla T4); a kaggle run (job `mdlm-browser-e4`) | ## Reproduce - sysone: `sysonelib` at commit `123ad8bac6` (main), with uncommitted changes in 30 files (diff SHA-256 `b81003e7b8e8`) - Ran: `run_f.py` - Python 3.13.15, sysone 0.3.0, torch 2.11.0+cu128, transformers 5.16.1, peft 0.20.0, accelerate 1.14.0, datasets 4.8.5, huggingface_hub 1.29.0, tokenizers 0.23.1, safetensors 0.8.0, numpy 2.1.3, fastcore 2.2.32, plum-dispatch 2.10.1; `environment.txt` lists every package