decosa-commit-detector-xlmr-large

Before a browser or computer-use agent clicks an element, this model says whether the click commits something that is hard to undo, and what kind: pay, delete, send, publish, submit, security (passwords, 2FA, keys, permissions) or none. The agent, or a guard around it, stops and asks a person before any commit.

It reads the element's role and label plus the page title, URL path, heading and nearby text, so a generic "OK", "Continue" or "Weiter" is judged by the page it sits on. It covers 34 languages (36 language codes with pt-BR and es-419), including all 24 official EU languages. It runs on a CPU in about 50-80 ms per click; no LLM call.

Why: agents usually gate clicks with English keyword lists. Ours had 12 regexes. On the test sets below they caught under 1% of commits, because most commits are phrased in other languages or with words the list doesn't have.

Results

Two test sets, never trained on. Binary = commit vs none. False-commit rate = share of harmless clicks the model flags, i.e. how often a person is asked for no reason. The baselines saw the same fields.

Independent set (720 cases, 36 language codes, written by a different author from the label definitions only; the most honest test):

Commit recall False-commit rate Binary accuracy Kind accuracy
12 English keyword regexes 0.2% 0.7% 39.9% —
Zero-shot Qwen3.8-27B (one-word answer, greedy) 98.1% 1.4% 98.3% 97.1%
This model (argmax) 98.4% 10.1% 95.0% 91.4%
This model at the shipped safety threshold 99.1% 13.2% 94.2% 90.6%

Held-out site kinds (3,206 cases from 8 kinds of website never seen in training, same generator as the training data):

Commit recall False-commit rate Binary accuracy Kind accuracy
12 English keyword regexes 0.6% 0.5% 37.9% —
Zero-shot Qwen3.8-27B 93.2% 6.2% 93.4% 90.6%
This model (argmax) 97.2% 5.4% 96.2% 94.0%
This model at the shipped safety threshold 97.9% 6.9% 96.1% 93.6%

So: it catches commits as well as a 27B LLM does, far better than a keyword list, but on independently written cases it asks a person needlessly about 1 time in 10 (Qwen: 1 in 70). AUROC commit vs none: .989 (independent), .987 (held-out). By kind on the independent set: delete, security and send 72/72 caught, submit 71/72, pay and publish 69/72. Weakest languages there: Bulgarian and English 85% binary accuracy (20 cases each); Croatian, Italian, Maltese, Romanian 90%.

Speed on CPU (fp32, 8 threads, one click at a time): p50 77 ms, p95 88 ms on a busy server; 47 ms p50 on an idle 24-thread box.

All numbers are in eval_summary.json.

Intended use

  • A pre-click guard for browser and computer-use agents: if commit is true, pause and ask a person, show the kind.
  • Use it together with your own rules (ours: model OR keyword list), not instead of a permission system.

Not for: deciding on its own that an action is safe to take without a person, or any use where a missed commit is unacceptable without another safeguard. A none answer means the model didn't recognise a commit, not that there isn't one.

Limitations

  • Synthetic training data. Real sites have longer and messier accessibility trees, icon-only buttons, cookie banners and dialogs. It has not yet been measured on real agent traces.
  • More needless confirmations than an LLM (10% vs 1.4% of harmless clicks on the independent set). Typical false commits: buttons that only look final ("Continue" mid-flow, "Delete" on a search-filter chip).
  • Text only: it doesn't see the screenshot. A button whose effect is only visible in an image is out of scope.
  • Irish and Maltese training rows are machine translations, so those languages are less certain.
  • The label set is a policy choice: "add to cart" is not a commit; cancelling a subscription counts as delete. Change the threshold or which kinds need approval to fit your policy.

Training

Base FacebookAI/xlm-roberta-large @ c23d21b0 (560M parameters), MIT
Data 16,566 synthetic UI moments (element, page title, URL path, heading, nearby text, label) for 42 kinds of website, written natively in each language by Qwen3.8-27B; Irish and Maltese rows machine-translated from English by EuroLLM-9B-Instruct. The label is the kind the generator was asked to write; rows where it reported writing another kind were dropped. All sites, companies and people are invented. The data and generator are not released.
Method Full fine-tune for sequence classification, 7 classes; AdamW (weight decay 0.01), lr 1.5e-5, cosine with 6% warm-up, batch 32, 3 epochs, bf16 autocast, label smoothing 0.05, max 160 tokens (element first); augmentation drops title / URL / context 12% each and changes the label's case 10%. 145 s on one H200.
Selection Three bases were trained (XLM-R large, mmBERT-base, mmBERT-small); XLM-R large was picked on the dev split, before any test number was read. The safety threshold (p(none) < 0.70 means commit) was set on dev for ≥ 98% commit recall.

Usage

pip install transformers torch huggingface_hub
python usage.py '{"role":"button","name":"Weiter","heading":"Bestellung prüfen","title":"Kasse","url":"/checkout/review","context":"Gesamt 49,90 EUR, zahlungspflichtig"}'
# {"commit": true, "kind": "pay", "probs": {...}}
from usage import CommitDetector
cd = CommitDetector()                         # or CommitDetector(threshold_pnone=0.5) for fewer confirmations
cd([{"role": "button", "name": "...", "heading": "...", "title": "...", "url": "...", "context": "..."}])

The input string must be built exactly as usage.text_of() does (role: name | heading: ... | title: ... | url: ... | context: ..., empty fields left out). decosa_commit.json holds the class order, the max length and the threshold.

Files: model.safetensors (fp32), config.json, tokenizer files, decosa_commit.json, usage.py, eval_summary.json, SHA256SUMS, LICENSE, NOTICE.

Licence

Apache-2.0 for the weights and code in this repository. The base model is MIT (Facebook, Inc. and its affiliates); its licence notice is in NOTICE.

Downloads last month
16
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for decosaai/decosa-commit-detector-xlmr-large

Finetuned
(1027)
this model