--- base_model: Qwen/Qwen3.5-4B library_name: mlx license: apache-2.0 pipeline_tag: image-text-to-text tags: - mlx - qwen3.5 - lora - computer-use - gui-agent - macos --- # gua **gua** is a computer-use decision model for macOS. Given a screenshot, the task and the last few actions, it chooses **what kind of action comes next** and **which on-screen element it acts on**, and can also **ground an element from a description**. It answers as typed multiple-choice decisions and returns a probability for every option, so an agent can act on confident answers and defer on the rest. It is [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) fine-tuned (supervised fine-tuning with LoRA, rank 8 on every linear layer of the language model) on GUI trajectories extracted from screen-recorded tutorials of seven kinds of software: macOS system apps, video editing, Blender, Godot, Logic Pro, OBS Studio and Plasticity. This repository holds the LoRA merged into the base weights (MLX, bf16) and, in `adapter/`, the LoRA itself. > Research model. The training data was derived from public YouTube tutorials and is not > released. See [Training data](#training-data) and [Limitations](#limitations). ## How it answers The model does not generate free text. Each step is asked as up to three questions, all sharing one prompt prefix (system prompt, screenshot, task, past actions, numbered element list), and the probability of every option is read from the model's next-token distribution: | Question | Options | Use | |---|---|---| | `action` | click, double-click, right-click, drag, type, press, shortcut, scroll, hover, wait, finish | the kind of the next action | | `next_target` | the numbered elements on screen | where the next action goes | | `grounding` | the numbered elements on screen | which element matches a description you give | Elements are **candidates found on the screenshot**: text lines from the macOS Vision framework, and, during training and evaluation, icons from the [OmniParser v2](https://huggingface.co/microsoft/OmniParser-v2.0) icon detector. The model chooses among candidates; it cannot point at something no candidate covers. When running on a live Mac, the accessibility tree is a better candidate source than detection. The prompt format is fixed by training. Use `inference/cua_decide.py`, which reproduces it exactly (checked to give identical option probabilities to the evaluation code). ## Usage Apple Silicon with [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) 0.7: ```bash pip install "mlx-vlm>=0.7,<0.8" pyobjc-framework-vision pyobjc-framework-quartz python inference/cua_decide.py --model /gua \ --image screen.png --task "Turn on Dark Mode" \ --past "Click the Apple menu" --past "Click System Settings" \ --describe "the Appearance item in the sidebar" \ --temperatures temperatures.json ``` Each question prints its choice and calibrated confidence, for example: ```json {"question": "action", "choice": "click", "confidence": 0.94, "action": "click"} {"question": "next_target", "choice": "15", "confidence": 0.60, "element": {"text": "...", "centre": [0.033, 0.136], "box": [...]}} ``` To use the LoRA instead of the merged weights, pass the base model and the adapter: `--model mlx-community/Qwen3.5-4B-MLX-bf16 --adapter adapter`. OmniParser's icon detector improves candidate recall but its weights are AGPL-3.0 and are not included; pass `--icon-weights icon_detect/model.pt` if you download them yourself. Without it, only text candidates are offered, which lowers recall on icon-only controls. ## Training gua v2 continues the v1 adapter on a larger, more varied dataset. | | | |---|---| | Method | Supervised fine-tuning with LoRA: rank 8, alpha 16, on every linear layer of the language model (`q/k/v/o_proj`, linear-attention `in_proj_*`/`out_proj`, MLP `gate/up/down_proj`); vision tower frozen | | Objective | Log loss of the right option label (and end-of-turn token) only, never of the prompt | | Stage 1 (v1) | 8,000 questions from *macos-cua v2* (macOS system, video editing, Blender, Godot) | | Stage 2 (v2) | 12,000 further questions from *macos-cua v3*, shared evenly across its seven topics (each topic at most about 1,940; smaller topics give all they have) | | Optimiser | Adam, learning rate 5e-5, batch 1, gradient clipping 1.0 | | Hardware | One Apple Silicon Mac (MLX), gradient checkpointing, about 110 h in total, peak 182 GB | The recipe was chosen on held-out dev splits, never on an eval split. The released weights are the end of stage 2, which scored higher on the full v3 dev split (mean accuracy 0.453) than the checkpoint with the best interim dev check (0.442). ## Training data Trajectories were extracted automatically from screen-recorded software tutorials on YouTube: videos were searched and screened (recording quality, platform), segmented and annotated into step-by-step actions by vision-language models, each click was located on the frame, and before/after screenshots were cut. The **training split** of *macos-cua v3*, and the questions stage 2 drew from it: | Topic | Videos | Steps | Questions used in stage 2 | |---|---|---|---| | Plasticity (CAD) | 162 | 16,656 | 1,937 | | Godot 4 (game engine) | 79 | 6,271 | 1,937 | | Blender (3D) | 40 | 3,833 | 1,937 | | OBS Studio (recording) | 37 | 2,890 | 1,937 | | Logic Pro (audio) | 43 | 2,020 | 1,937 | | macOS video editing (Final Cut Pro, Premiere Pro, CapCut, iMovie) | 7 | 640 | 1,468 | | macOS system and productivity apps | 6 | 420 | 844 | Splits are by video (dev 44 videos, eval 99 videos, 38 applications appear only in eval); every video of the earlier v2 release kept its split, so v2's eval is part of v3's eval and was never trained on. Godot typing steps were left out of training: an audit found their labels mostly wrong. Many tutorials were recorded on Windows or full screen; steps that would be wrong on a Mac (Windows shortcuts, taskbar, native dialogs) were removed. **Label quality.** Labels are model-generated, not human-verified. A stronger model audited a random sample per topic: | Topic | Audited steps | Element located correctly | Action correct | |---|---|---|---| | Logic Pro | 150 | 94.0% | 84.0% | | OBS Studio | 150 | 92.0% | 81.3% | | macOS video editing | 395 | 92.4% | 78.2% | | macOS system and productivity | 200 | 91.5% | 75.5% | | Blender | 180 | 89.4% | 71.7% | | Plasticity | 570 | 79.6% | 71.4% | | Godot | 270 | 87.0% | 54.8% | The dataset itself is not released. ## Evaluation **macos-cua v3 eval split** (9,972 steps from 99 videos never seen in training, 6,509 with an element target), gua v2 against gua v1 with the same prompt and candidates, 95% CI by bootstrap paired by trajectory: | Question | gua v1 | gua v2 | Difference | |---|---|---|---| | `action` accuracy | 0.507 | 0.537 | +0.030 (+0.021, +0.040) | | `grounding` accuracy | 0.516 | 0.524 | +0.008 (+0.002, +0.015) | | `next_target` accuracy | 0.303 | 0.323 | +0.020 (+0.012, +0.029) | | `action` and `next_target` both right | 0.182 | 0.216 | | On the 1,767 steps from **applications never seen in training**: action 0.602, grounding 0.628, next target 0.357. By topic (gua v2): | Topic | `action` | `grounding` | `next_target` | |---|---|---|---| | OBS Studio | 0.806 | 0.803 | 0.453 | | macOS system and productivity | 0.720 | 0.643 | 0.465 | | Logic Pro | 0.687 | 0.566 | 0.341 | | macOS video editing | 0.592 | 0.644 | 0.388 | | Blender | 0.539 | 0.503 | 0.253 | | Godot | 0.519 | 0.657 | 0.275 | | Plasticity | 0.462 | 0.381 | 0.308 | **macos-cua v2 eval split** (3,442 steps, the eval gua v1 was published with), against the untrained base: | Question | Base | gua v1 | gua v2 | |---|---|---|---| | `action` accuracy | 0.481 | 0.525 | 0.542 | | `grounding` accuracy | 0.443 | 0.603 | 0.605 | | `next_target` accuracy | 0.155 | 0.261 | 0.289 | A candidate covers the target in 70.2% of v3's element steps (83.4% in v2's), which bounds the element questions; Plasticity's small, dense CAD icons are often missed. Excluding the Godot typing steps, whose labels are mostly wrong, v3 action accuracy is 0.563. Because those steps were left out of training, the model rarely answers `type`. **Selective accuracy** (v3 eval). Acting only on the most confident half of the answers: grounding 0.711, action 0.665, next target 0.444. **Calibration.** `temperatures.json` holds one temperature per question, fitted on the v3 dev split by log loss and applied as `p^(1/T)` renormalised (`inference/cua_decide.py --temperatures`). Expected calibration error on the v3 eval split, which the fit never saw: | Question | Temperature | ECE before | ECE after | |---|---|---|---| | `action` | 2.15 | 0.294 | 0.060 | | `grounding` | 1.25 | 0.259 | 0.173 | | `next_target` | 1.20 | 0.253 | 0.141 | Without scaling the model is overconfident on every question; after it, element answers remain somewhat overconfident. ## Limitations - **macOS screenshots only**, English UI and tasks. - **Candidate-bound.** The model picks among candidates; if no candidate covers the target (about 30% of element steps in v3's eval, most often in CAD), the element questions have no right answer. On a live Mac, the accessibility tree is a better candidate source. - **Noisy labels**, as audited above; action-type labels are the weakest, and Plasticity's element labels the least reliable. - **Weak on dense professional UIs**: Plasticity grounding is 0.38. - **Not an agent by itself.** It decides the action kind and target; the text to type, the keys of a shortcut and drag end points are not predicted. It rarely predicts `type`. - **Merged weights round to bf16.** Against base plus adapter, option probabilities differ by at most 0.035 in total variation on our check; use `adapter/` for exact reproduction. - Confidence is only meaningful after temperature scaling, and only for the question types and candidate sources it was fitted with. ## Versions | Version | Training | Notes | |---|---|---| | v2 (this) | v1 continued on 12,000 balanced questions from macos-cua v3 | adds Logic Pro, OBS Studio, Plasticity; better on every question | | v1 | 8,000 questions from macos-cua v2 | first release; tagged `v1` in this repository | ## Licence and provenance The weights are a derivative of Qwen3.5-4B (Apache-2.0) and are released under Apache-2.0 (`LICENSE`); the modification is the LoRA fine-tuning described above. The training data was derived from public YouTube videos for research; no video frames are distributed here. `inference/cua_decide.py` is released under the same licence. OmniParser v2 (optional at inference) is licensed separately by its authors.