Instructions to use LaplAI/gua with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use LaplAI/gua with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("LaplAI/gua") config = load_config("LaplAI/gua") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use LaplAI/gua with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "LaplAI/gua"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "LaplAI/gua" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use LaplAI/gua with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "LaplAI/gua"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default LaplAI/gua
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use LaplAI/gua with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "LaplAI/gua"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "LaplAI/gua" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
gua v2: continue v1 on 12,000 balanced questions from macos-cua v3 (adds Logic Pro, OBS, Plasticity)
1357427 verified |
Download README.md from LaplAI/gua: direct link, hf CLI and curl.
- Browser
- Download file 10.7 kB
-
https://huggingface.co/LaplAI/gua/resolve/main/README.md
- Command line
-
hf download hf://LaplAI/gua/README.md
-
curl -L -o README.md https://huggingface.co/LaplAI/gua/resolve/main/README.md
10.7 kB
| base_model: Qwen/Qwen3.5-4B | |
| library_name: mlx | |
| license: apache-2.0 | |
| pipeline_tag: image-text-to-text | |
| tags: | |
| - mlx | |
| - qwen3.5 | |
| - lora | |
| - computer-use | |
| - gui-agent | |
| - macos | |
| # gua | |
| **gua** is a computer-use decision model for macOS. Given a screenshot, the task and the last few | |
| actions, it chooses **what kind of action comes next** and **which on-screen element it | |
| acts on**, and can also **ground an element from a description**. It answers as typed | |
| multiple-choice decisions and returns a probability for every option, so an agent can act | |
| on confident answers and defer on the rest. | |
| It is [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) fine-tuned (supervised | |
| fine-tuning with LoRA, rank 8 on every linear layer of the language model) on GUI | |
| trajectories extracted from screen-recorded tutorials of seven kinds of software: macOS | |
| system apps, video editing, Blender, Godot, Logic Pro, OBS Studio and Plasticity. This repository holds the | |
| LoRA merged into the base weights (MLX, bf16) and, in `adapter/`, the LoRA itself. | |
| > Research model. The training data was derived from public YouTube tutorials and is not | |
| > released. See [Training data](#training-data) and [Limitations](#limitations). | |
| ## How it answers | |
| The model does not generate free text. Each step is asked as up to three questions, all | |
| sharing one prompt prefix (system prompt, screenshot, task, past actions, numbered element | |
| list), and the probability of every option is read from the model's next-token | |
| distribution: | |
| | Question | Options | Use | | |
| |---|---|---| | |
| | `action` | click, double-click, right-click, drag, type, press, shortcut, scroll, hover, wait, finish | the kind of the next action | | |
| | `next_target` | the numbered elements on screen | where the next action goes | | |
| | `grounding` | the numbered elements on screen | which element matches a description you give | | |
| Elements are **candidates found on the screenshot**: text lines from the macOS Vision | |
| framework, and, during training and evaluation, icons from the | |
| [OmniParser v2](https://huggingface.co/microsoft/OmniParser-v2.0) icon detector. The model | |
| chooses among candidates; it cannot point at something no candidate covers. When running | |
| on a live Mac, the accessibility tree is a better candidate source than detection. | |
| The prompt format is fixed by training. Use `inference/cua_decide.py`, which reproduces it | |
| exactly (checked to give identical option probabilities to the evaluation code). | |
| ## Usage | |
| Apple Silicon with [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) 0.7: | |
| ```bash | |
| pip install "mlx-vlm>=0.7,<0.8" pyobjc-framework-vision pyobjc-framework-quartz | |
| python inference/cua_decide.py --model <user>/gua \ | |
| --image screen.png --task "Turn on Dark Mode" \ | |
| --past "Click the Apple menu" --past "Click System Settings" \ | |
| --describe "the Appearance item in the sidebar" \ | |
| --temperatures temperatures.json | |
| ``` | |
| Each question prints its choice and calibrated confidence, for example: | |
| ```json | |
| {"question": "action", "choice": "click", "confidence": 0.94, "action": "click"} | |
| {"question": "next_target", "choice": "15", "confidence": 0.60, "element": {"text": "...", "centre": [0.033, 0.136], "box": [...]}} | |
| ``` | |
| To use the LoRA instead of the merged weights, pass the base model and the adapter: | |
| `--model mlx-community/Qwen3.5-4B-MLX-bf16 --adapter adapter`. | |
| OmniParser's icon detector improves candidate recall but its weights are AGPL-3.0 and are | |
| not included; pass `--icon-weights icon_detect/model.pt` if you download them yourself. | |
| Without it, only text candidates are offered, which lowers recall on icon-only controls. | |
| ## Training | |
| gua v2 continues the v1 adapter on a larger, more varied dataset. | |
| | | | | |
| |---|---| | |
| | Method | Supervised fine-tuning with LoRA: rank 8, alpha 16, on every linear layer of the language model (`q/k/v/o_proj`, linear-attention `in_proj_*`/`out_proj`, MLP `gate/up/down_proj`); vision tower frozen | | |
| | Objective | Log loss of the right option label (and end-of-turn token) only, never of the prompt | | |
| | Stage 1 (v1) | 8,000 questions from *macos-cua v2* (macOS system, video editing, Blender, Godot) | | |
| | Stage 2 (v2) | 12,000 further questions from *macos-cua v3*, shared evenly across its seven topics (each topic at most about 1,940; smaller topics give all they have) | | |
| | Optimiser | Adam, learning rate 5e-5, batch 1, gradient clipping 1.0 | | |
| | Hardware | One Apple Silicon Mac (MLX), gradient checkpointing, about 110 h in total, peak 182 GB | | |
| The recipe was chosen on held-out dev splits, never on an eval split. The released weights | |
| are the end of stage 2, which scored higher on the full v3 dev split (mean accuracy 0.453) | |
| than the checkpoint with the best interim dev check (0.442). | |
| ## Training data | |
| Trajectories were extracted automatically from screen-recorded software tutorials on | |
| YouTube: videos were searched and screened (recording quality, platform), segmented and | |
| annotated into step-by-step actions by vision-language models, each click was located on | |
| the frame, and before/after screenshots were cut. The **training split** of *macos-cua v3*, | |
| and the questions stage 2 drew from it: | |
| | Topic | Videos | Steps | Questions used in stage 2 | | |
| |---|---|---|---| | |
| | Plasticity (CAD) | 162 | 16,656 | 1,937 | | |
| | Godot 4 (game engine) | 79 | 6,271 | 1,937 | | |
| | Blender (3D) | 40 | 3,833 | 1,937 | | |
| | OBS Studio (recording) | 37 | 2,890 | 1,937 | | |
| | Logic Pro (audio) | 43 | 2,020 | 1,937 | | |
| | macOS video editing (Final Cut Pro, Premiere Pro, CapCut, iMovie) | 7 | 640 | 1,468 | | |
| | macOS system and productivity apps | 6 | 420 | 844 | | |
| Splits are by video (dev 44 videos, eval 99 videos, 38 applications appear only in eval); | |
| every video of the earlier v2 release kept its split, so v2's eval is part of v3's eval and | |
| was never trained on. Godot typing steps were left out of training: an audit found their | |
| labels mostly wrong. Many tutorials were recorded on Windows or full screen; steps that | |
| would be wrong on a Mac (Windows shortcuts, taskbar, native dialogs) were removed. | |
| **Label quality.** Labels are model-generated, not human-verified. A stronger model audited | |
| a random sample per topic: | |
| | Topic | Audited steps | Element located correctly | Action correct | | |
| |---|---|---|---| | |
| | Logic Pro | 150 | 94.0% | 84.0% | | |
| | OBS Studio | 150 | 92.0% | 81.3% | | |
| | macOS video editing | 395 | 92.4% | 78.2% | | |
| | macOS system and productivity | 200 | 91.5% | 75.5% | | |
| | Blender | 180 | 89.4% | 71.7% | | |
| | Plasticity | 570 | 79.6% | 71.4% | | |
| | Godot | 270 | 87.0% | 54.8% | | |
| The dataset itself is not released. | |
| ## Evaluation | |
| **macos-cua v3 eval split** (9,972 steps from 99 videos never seen in training, 6,509 with | |
| an element target), gua v2 against gua v1 with the same prompt and candidates, 95% CI by | |
| bootstrap paired by trajectory: | |
| | Question | gua v1 | gua v2 | Difference | | |
| |---|---|---|---| | |
| | `action` accuracy | 0.507 | 0.537 | +0.030 (+0.021, +0.040) | | |
| | `grounding` accuracy | 0.516 | 0.524 | +0.008 (+0.002, +0.015) | | |
| | `next_target` accuracy | 0.303 | 0.323 | +0.020 (+0.012, +0.029) | | |
| | `action` and `next_target` both right | 0.182 | 0.216 | | | |
| On the 1,767 steps from **applications never seen in training**: action 0.602, grounding | |
| 0.628, next target 0.357. By topic (gua v2): | |
| | Topic | `action` | `grounding` | `next_target` | | |
| |---|---|---|---| | |
| | OBS Studio | 0.806 | 0.803 | 0.453 | | |
| | macOS system and productivity | 0.720 | 0.643 | 0.465 | | |
| | Logic Pro | 0.687 | 0.566 | 0.341 | | |
| | macOS video editing | 0.592 | 0.644 | 0.388 | | |
| | Blender | 0.539 | 0.503 | 0.253 | | |
| | Godot | 0.519 | 0.657 | 0.275 | | |
| | Plasticity | 0.462 | 0.381 | 0.308 | | |
| **macos-cua v2 eval split** (3,442 steps, the eval gua v1 was published with), against the | |
| untrained base: | |
| | Question | Base | gua v1 | gua v2 | | |
| |---|---|---|---| | |
| | `action` accuracy | 0.481 | 0.525 | 0.542 | | |
| | `grounding` accuracy | 0.443 | 0.603 | 0.605 | | |
| | `next_target` accuracy | 0.155 | 0.261 | 0.289 | | |
| A candidate covers the target in 70.2% of v3's element steps (83.4% in v2's), which bounds | |
| the element questions; Plasticity's small, dense CAD icons are often missed. Excluding the | |
| Godot typing steps, whose labels are mostly wrong, v3 action accuracy is 0.563. Because | |
| those steps were left out of training, the model rarely answers `type`. | |
| **Selective accuracy** (v3 eval). Acting only on the most confident half of the answers: | |
| grounding 0.711, action 0.665, next target 0.444. | |
| **Calibration.** `temperatures.json` holds one temperature per question, fitted on the v3 | |
| dev split by log loss and applied as `p^(1/T)` renormalised (`inference/cua_decide.py | |
| --temperatures`). Expected calibration error on the v3 eval split, which the fit never saw: | |
| | Question | Temperature | ECE before | ECE after | | |
| |---|---|---|---| | |
| | `action` | 2.15 | 0.294 | 0.060 | | |
| | `grounding` | 1.25 | 0.259 | 0.173 | | |
| | `next_target` | 1.20 | 0.253 | 0.141 | | |
| Without scaling the model is overconfident on every question; after it, element answers | |
| remain somewhat overconfident. | |
| ## Limitations | |
| - **macOS screenshots only**, English UI and tasks. | |
| - **Candidate-bound.** The model picks among candidates; if no candidate covers the target | |
| (about 30% of element steps in v3's eval, most often in CAD), the element questions have | |
| no right answer. On a live Mac, the accessibility tree is a better candidate source. | |
| - **Noisy labels**, as audited above; action-type labels are the weakest, and Plasticity's | |
| element labels the least reliable. | |
| - **Weak on dense professional UIs**: Plasticity grounding is 0.38. | |
| - **Not an agent by itself.** It decides the action kind and target; the text to type, the | |
| keys of a shortcut and drag end points are not predicted. It rarely predicts `type`. | |
| - **Merged weights round to bf16.** Against base plus adapter, option probabilities differ | |
| by at most 0.035 in total variation on our check; use `adapter/` for exact reproduction. | |
| - Confidence is only meaningful after temperature scaling, and only for the question types | |
| and candidate sources it was fitted with. | |
| ## Versions | |
| | Version | Training | Notes | | |
| |---|---|---| | |
| | v2 (this) | v1 continued on 12,000 balanced questions from macos-cua v3 | adds Logic Pro, OBS Studio, Plasticity; better on every question | | |
| | v1 | 8,000 questions from macos-cua v2 | first release; tagged `v1` in this repository | | |
| ## Licence and provenance | |
| The weights are a derivative of Qwen3.5-4B (Apache-2.0) and are released under Apache-2.0 | |
| (`LICENSE`); the modification is the LoRA fine-tuning described above. The training data | |
| was derived from public YouTube videos for research; no video frames are distributed here. | |
| `inference/cua_decide.py` is released under the same licence. OmniParser v2 (optional at | |
| inference) is licensed separately by its authors. | |