Instructions to use decosaai/decosa-commit-detector-xlmr-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use decosaai/decosa-commit-detector-xlmr-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="decosaai/decosa-commit-detector-xlmr-large")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("decosaai/decosa-commit-detector-xlmr-large") model = AutoModelForSequenceClassification.from_pretrained("decosaai/decosa-commit-detector-xlmr-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
decosa-commit-detector-xlmr-large
Before a browser or computer-use agent clicks an element, this model says whether the click commits something that
is hard to undo, and what kind: pay, delete, send, publish, submit, security (passwords, 2FA, keys,
permissions) or none. The agent, or a guard around it, stops and asks a person before any commit.
It reads the element's role and label plus the page title, URL path, heading and nearby text, so a generic "OK", "Continue" or "Weiter" is judged by the page it sits on. It covers 34 languages (36 language codes with pt-BR and es-419), including all 24 official EU languages. It runs on a CPU in about 50-80 ms per click; no LLM call.
Why: agents usually gate clicks with English keyword lists. Ours had 12 regexes. On the test sets below they caught under 1% of commits, because most commits are phrased in other languages or with words the list doesn't have.
Results
Two test sets, never trained on. Binary = commit vs none. False-commit rate = share of harmless clicks the model flags, i.e. how often a person is asked for no reason. The baselines saw the same fields.
Independent set (720 cases, 36 language codes, written by a different author from the label definitions only; the most honest test):
| Commit recall | False-commit rate | Binary accuracy | Kind accuracy | |
|---|---|---|---|---|
| 12 English keyword regexes | 0.2% | 0.7% | 39.9% | — |
| Zero-shot Qwen3.8-27B (one-word answer, greedy) | 98.1% | 1.4% | 98.3% | 97.1% |
| This model (argmax) | 98.4% | 10.1% | 95.0% | 91.4% |
| This model at the shipped safety threshold | 99.1% | 13.2% | 94.2% | 90.6% |
Held-out site kinds (3,206 cases from 8 kinds of website never seen in training, same generator as the training data):
| Commit recall | False-commit rate | Binary accuracy | Kind accuracy | |
|---|---|---|---|---|
| 12 English keyword regexes | 0.6% | 0.5% | 37.9% | — |
| Zero-shot Qwen3.8-27B | 93.2% | 6.2% | 93.4% | 90.6% |
| This model (argmax) | 97.2% | 5.4% | 96.2% | 94.0% |
| This model at the shipped safety threshold | 97.9% | 6.9% | 96.1% | 93.6% |
So: it catches commits as well as a 27B LLM does, far better than a keyword list, but on independently written cases it asks a person needlessly about 1 time in 10 (Qwen: 1 in 70). AUROC commit vs none: .989 (independent), .987 (held-out). By kind on the independent set: delete, security and send 72/72 caught, submit 71/72, pay and publish 69/72. Weakest languages there: Bulgarian and English 85% binary accuracy (20 cases each); Croatian, Italian, Maltese, Romanian 90%.
Speed on CPU (fp32, 8 threads, one click at a time): p50 77 ms, p95 88 ms on a busy server; 47 ms p50 on an idle 24-thread box.
All numbers are in eval_summary.json.
Intended use
- A pre-click guard for browser and computer-use agents: if
commitis true, pause and ask a person, show the kind. - Use it together with your own rules (ours:
model OR keyword list), not instead of a permission system.
Not for: deciding on its own that an action is safe to take without a person, or any use where a missed commit is
unacceptable without another safeguard. A none answer means the model didn't recognise a commit, not that there
isn't one.
Limitations
- Synthetic training data. Real sites have longer and messier accessibility trees, icon-only buttons, cookie banners and dialogs. It has not yet been measured on real agent traces.
- More needless confirmations than an LLM (10% vs 1.4% of harmless clicks on the independent set). Typical false commits: buttons that only look final ("Continue" mid-flow, "Delete" on a search-filter chip).
- Text only: it doesn't see the screenshot. A button whose effect is only visible in an image is out of scope.
- Irish and Maltese training rows are machine translations, so those languages are less certain.
- The label set is a policy choice: "add to cart" is not a commit; cancelling a subscription counts as
delete. Change the threshold or which kinds need approval to fit your policy.
Training
| Base | FacebookAI/xlm-roberta-large @ c23d21b0 (560M parameters), MIT |
| Data | 16,566 synthetic UI moments (element, page title, URL path, heading, nearby text, label) for 42 kinds of website, written natively in each language by Qwen3.8-27B; Irish and Maltese rows machine-translated from English by EuroLLM-9B-Instruct. The label is the kind the generator was asked to write; rows where it reported writing another kind were dropped. All sites, companies and people are invented. The data and generator are not released. |
| Method | Full fine-tune for sequence classification, 7 classes; AdamW (weight decay 0.01), lr 1.5e-5, cosine with 6% warm-up, batch 32, 3 epochs, bf16 autocast, label smoothing 0.05, max 160 tokens (element first); augmentation drops title / URL / context 12% each and changes the label's case 10%. 145 s on one H200. |
| Selection | Three bases were trained (XLM-R large, mmBERT-base, mmBERT-small); XLM-R large was picked on the dev split, before any test number was read. The safety threshold (p(none) < 0.70 means commit) was set on dev for ≥ 98% commit recall. |
Usage
pip install transformers torch huggingface_hub
python usage.py '{"role":"button","name":"Weiter","heading":"Bestellung prüfen","title":"Kasse","url":"/checkout/review","context":"Gesamt 49,90 EUR, zahlungspflichtig"}'
# {"commit": true, "kind": "pay", "probs": {...}}
from usage import CommitDetector
cd = CommitDetector() # or CommitDetector(threshold_pnone=0.5) for fewer confirmations
cd([{"role": "button", "name": "...", "heading": "...", "title": "...", "url": "...", "context": "..."}])
The input string must be built exactly as usage.text_of() does (role: name | heading: ... | title: ... | url: ... | context: ..., empty fields left out). decosa_commit.json holds the class order, the max length and the threshold.
Files: model.safetensors (fp32), config.json, tokenizer files, decosa_commit.json, usage.py,
eval_summary.json, SHA256SUMS, LICENSE, NOTICE.
Licence
Apache-2.0 for the weights and code in this repository. The base model is MIT (Facebook, Inc. and its affiliates); its licence notice is in NOTICE.
- Downloads last month
- 16
Model tree for decosaai/decosa-commit-detector-xlmr-large
Base model
FacebookAI/xlm-roberta-large