gua / README.md
chenyvathf's picture
gua v2: continue v1 on 12,000 balanced questions from macos-cua v3 (adds Logic Pro, OBS, Plasticity)
1357427 verified
|
Raw History Blame Contribute Delete
10.7 kB
---
base_model: Qwen/Qwen3.5-4B
library_name: mlx
license: apache-2.0
pipeline_tag: image-text-to-text
tags:
- mlx
- qwen3.5
- lora
- computer-use
- gui-agent
- macos
---
# gua
**gua** is a computer-use decision model for macOS. Given a screenshot, the task and the last few
actions, it chooses **what kind of action comes next** and **which on-screen element it
acts on**, and can also **ground an element from a description**. It answers as typed
multiple-choice decisions and returns a probability for every option, so an agent can act
on confident answers and defer on the rest.
It is [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) fine-tuned (supervised
fine-tuning with LoRA, rank 8 on every linear layer of the language model) on GUI
trajectories extracted from screen-recorded tutorials of seven kinds of software: macOS
system apps, video editing, Blender, Godot, Logic Pro, OBS Studio and Plasticity. This repository holds the
LoRA merged into the base weights (MLX, bf16) and, in `adapter/`, the LoRA itself.
> Research model. The training data was derived from public YouTube tutorials and is not
> released. See [Training data](#training-data) and [Limitations](#limitations).
## How it answers
The model does not generate free text. Each step is asked as up to three questions, all
sharing one prompt prefix (system prompt, screenshot, task, past actions, numbered element
list), and the probability of every option is read from the model's next-token
distribution:
| Question | Options | Use |
|---|---|---|
| `action` | click, double-click, right-click, drag, type, press, shortcut, scroll, hover, wait, finish | the kind of the next action |
| `next_target` | the numbered elements on screen | where the next action goes |
| `grounding` | the numbered elements on screen | which element matches a description you give |
Elements are **candidates found on the screenshot**: text lines from the macOS Vision
framework, and, during training and evaluation, icons from the
[OmniParser v2](https://huggingface.co/microsoft/OmniParser-v2.0) icon detector. The model
chooses among candidates; it cannot point at something no candidate covers. When running
on a live Mac, the accessibility tree is a better candidate source than detection.
The prompt format is fixed by training. Use `inference/cua_decide.py`, which reproduces it
exactly (checked to give identical option probabilities to the evaluation code).
## Usage
Apple Silicon with [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) 0.7:
```bash
pip install "mlx-vlm>=0.7,<0.8" pyobjc-framework-vision pyobjc-framework-quartz
python inference/cua_decide.py --model <user>/gua \
--image screen.png --task "Turn on Dark Mode" \
--past "Click the Apple menu" --past "Click System Settings" \
--describe "the Appearance item in the sidebar" \
--temperatures temperatures.json
```
Each question prints its choice and calibrated confidence, for example:
```json
{"question": "action", "choice": "click", "confidence": 0.94, "action": "click"}
{"question": "next_target", "choice": "15", "confidence": 0.60, "element": {"text": "...", "centre": [0.033, 0.136], "box": [...]}}
```
To use the LoRA instead of the merged weights, pass the base model and the adapter:
`--model mlx-community/Qwen3.5-4B-MLX-bf16 --adapter adapter`.
OmniParser's icon detector improves candidate recall but its weights are AGPL-3.0 and are
not included; pass `--icon-weights icon_detect/model.pt` if you download them yourself.
Without it, only text candidates are offered, which lowers recall on icon-only controls.
## Training
gua v2 continues the v1 adapter on a larger, more varied dataset.
| | |
|---|---|
| Method | Supervised fine-tuning with LoRA: rank 8, alpha 16, on every linear layer of the language model (`q/k/v/o_proj`, linear-attention `in_proj_*`/`out_proj`, MLP `gate/up/down_proj`); vision tower frozen |
| Objective | Log loss of the right option label (and end-of-turn token) only, never of the prompt |
| Stage 1 (v1) | 8,000 questions from *macos-cua v2* (macOS system, video editing, Blender, Godot) |
| Stage 2 (v2) | 12,000 further questions from *macos-cua v3*, shared evenly across its seven topics (each topic at most about 1,940; smaller topics give all they have) |
| Optimiser | Adam, learning rate 5e-5, batch 1, gradient clipping 1.0 |
| Hardware | One Apple Silicon Mac (MLX), gradient checkpointing, about 110 h in total, peak 182 GB |
The recipe was chosen on held-out dev splits, never on an eval split. The released weights
are the end of stage 2, which scored higher on the full v3 dev split (mean accuracy 0.453)
than the checkpoint with the best interim dev check (0.442).
## Training data
Trajectories were extracted automatically from screen-recorded software tutorials on
YouTube: videos were searched and screened (recording quality, platform), segmented and
annotated into step-by-step actions by vision-language models, each click was located on
the frame, and before/after screenshots were cut. The **training split** of *macos-cua v3*,
and the questions stage 2 drew from it:
| Topic | Videos | Steps | Questions used in stage 2 |
|---|---|---|---|
| Plasticity (CAD) | 162 | 16,656 | 1,937 |
| Godot 4 (game engine) | 79 | 6,271 | 1,937 |
| Blender (3D) | 40 | 3,833 | 1,937 |
| OBS Studio (recording) | 37 | 2,890 | 1,937 |
| Logic Pro (audio) | 43 | 2,020 | 1,937 |
| macOS video editing (Final Cut Pro, Premiere Pro, CapCut, iMovie) | 7 | 640 | 1,468 |
| macOS system and productivity apps | 6 | 420 | 844 |
Splits are by video (dev 44 videos, eval 99 videos, 38 applications appear only in eval);
every video of the earlier v2 release kept its split, so v2's eval is part of v3's eval and
was never trained on. Godot typing steps were left out of training: an audit found their
labels mostly wrong. Many tutorials were recorded on Windows or full screen; steps that
would be wrong on a Mac (Windows shortcuts, taskbar, native dialogs) were removed.
**Label quality.** Labels are model-generated, not human-verified. A stronger model audited
a random sample per topic:
| Topic | Audited steps | Element located correctly | Action correct |
|---|---|---|---|
| Logic Pro | 150 | 94.0% | 84.0% |
| OBS Studio | 150 | 92.0% | 81.3% |
| macOS video editing | 395 | 92.4% | 78.2% |
| macOS system and productivity | 200 | 91.5% | 75.5% |
| Blender | 180 | 89.4% | 71.7% |
| Plasticity | 570 | 79.6% | 71.4% |
| Godot | 270 | 87.0% | 54.8% |
The dataset itself is not released.
## Evaluation
**macos-cua v3 eval split** (9,972 steps from 99 videos never seen in training, 6,509 with
an element target), gua v2 against gua v1 with the same prompt and candidates, 95% CI by
bootstrap paired by trajectory:
| Question | gua v1 | gua v2 | Difference |
|---|---|---|---|
| `action` accuracy | 0.507 | 0.537 | +0.030 (+0.021, +0.040) |
| `grounding` accuracy | 0.516 | 0.524 | +0.008 (+0.002, +0.015) |
| `next_target` accuracy | 0.303 | 0.323 | +0.020 (+0.012, +0.029) |
| `action` and `next_target` both right | 0.182 | 0.216 | |
On the 1,767 steps from **applications never seen in training**: action 0.602, grounding
0.628, next target 0.357. By topic (gua v2):
| Topic | `action` | `grounding` | `next_target` |
|---|---|---|---|
| OBS Studio | 0.806 | 0.803 | 0.453 |
| macOS system and productivity | 0.720 | 0.643 | 0.465 |
| Logic Pro | 0.687 | 0.566 | 0.341 |
| macOS video editing | 0.592 | 0.644 | 0.388 |
| Blender | 0.539 | 0.503 | 0.253 |
| Godot | 0.519 | 0.657 | 0.275 |
| Plasticity | 0.462 | 0.381 | 0.308 |
**macos-cua v2 eval split** (3,442 steps, the eval gua v1 was published with), against the
untrained base:
| Question | Base | gua v1 | gua v2 |
|---|---|---|---|
| `action` accuracy | 0.481 | 0.525 | 0.542 |
| `grounding` accuracy | 0.443 | 0.603 | 0.605 |
| `next_target` accuracy | 0.155 | 0.261 | 0.289 |
A candidate covers the target in 70.2% of v3's element steps (83.4% in v2's), which bounds
the element questions; Plasticity's small, dense CAD icons are often missed. Excluding the
Godot typing steps, whose labels are mostly wrong, v3 action accuracy is 0.563. Because
those steps were left out of training, the model rarely answers `type`.
**Selective accuracy** (v3 eval). Acting only on the most confident half of the answers:
grounding 0.711, action 0.665, next target 0.444.
**Calibration.** `temperatures.json` holds one temperature per question, fitted on the v3
dev split by log loss and applied as `p^(1/T)` renormalised (`inference/cua_decide.py
--temperatures`). Expected calibration error on the v3 eval split, which the fit never saw:
| Question | Temperature | ECE before | ECE after |
|---|---|---|---|
| `action` | 2.15 | 0.294 | 0.060 |
| `grounding` | 1.25 | 0.259 | 0.173 |
| `next_target` | 1.20 | 0.253 | 0.141 |
Without scaling the model is overconfident on every question; after it, element answers
remain somewhat overconfident.
## Limitations
- **macOS screenshots only**, English UI and tasks.
- **Candidate-bound.** The model picks among candidates; if no candidate covers the target
(about 30% of element steps in v3's eval, most often in CAD), the element questions have
no right answer. On a live Mac, the accessibility tree is a better candidate source.
- **Noisy labels**, as audited above; action-type labels are the weakest, and Plasticity's
element labels the least reliable.
- **Weak on dense professional UIs**: Plasticity grounding is 0.38.
- **Not an agent by itself.** It decides the action kind and target; the text to type, the
keys of a shortcut and drag end points are not predicted. It rarely predicts `type`.
- **Merged weights round to bf16.** Against base plus adapter, option probabilities differ
by at most 0.035 in total variation on our check; use `adapter/` for exact reproduction.
- Confidence is only meaningful after temperature scaling, and only for the question types
and candidate sources it was fitted with.
## Versions
| Version | Training | Notes |
|---|---|---|
| v2 (this) | v1 continued on 12,000 balanced questions from macos-cua v3 | adds Logic Pro, OBS Studio, Plasticity; better on every question |
| v1 | 8,000 questions from macos-cua v2 | first release; tagged `v1` in this repository |
## Licence and provenance
The weights are a derivative of Qwen3.5-4B (Apache-2.0) and are released under Apache-2.0
(`LICENSE`); the modification is the LoRA fine-tuning described above. The training data
was derived from public YouTube videos for research; no video frames are distributed here.
`inference/cua_decide.py` is released under the same licence. OmniParser v2 (optional at
inference) is licensed separately by its authors.