LibreRFDETRm-ui

Class-agnostic UI element detector for screenshots (buttons, links, inputs, calendar cells, icons), repackaged for LibreYOLO. It is UI-DETR-1, an RF-DETR-Medium fine-tune built as the perception stage of a computer-use agent. One class: object (any interactive element).

Usage

from libreyolo import LibreYOLO

model = LibreYOLO("LibreRFDETRm-ui.pt")  # auto-downloads from this repo
results = model.predict("screenshot.png", conf=0.35)

The authors use a confidence threshold of 0.35. Native input size is 576.

Source

Weights: racineai/UI-DETR-1 at revision 0f0dda5. Copyright (c) 2025 UI-DETR-1 Team (racineai, TW3 Partners). Licensed under the MIT License.

Base model: RF-DETR-Medium from roboflow/rf-detr, Apache License 2.0, with a DINOv2 backbone from facebookresearch/dinov2, Apache License 2.0.

Training data: 2,656 screenshots from six Roboflow Universe datasets, merged by the authors into one class. The authors declare the training data MIT in their release notes.

Modifications

Checkpoint metadata wrapping only. Learned parameters are the upstream model state (not the EMA copy), bit-identical to model.pth. Optimizer, scheduler, EMA, and training args were dropped. The wrapper adds model_family, size, task, nc=1, names={0: "object"}, and imgsz=576 so the LibreYOLO() factory routes without filename heuristics.

Checked against upstream rfdetr 1.10.1 on five UI screenshots at conf 0.3: identical box counts, every box matched at IoU > 0.9, median confidence difference below 1e-3.

License

MIT (UI-DETR-1 fine-tune), on top of Apache License 2.0 base weights. See LICENSE and NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LibreYOLO/LibreRFDETRm-ui

Finetuned
(1)
this model

Collection including LibreYOLO/LibreRFDETRm-ui