Instructions to use LaplAI/gua with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use LaplAI/gua with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("LaplAI/gua") config = load_config("LaplAI/gua") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use LaplAI/gua with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "LaplAI/gua"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "LaplAI/gua" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use LaplAI/gua with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "LaplAI/gua"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default LaplAI/gua
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use LaplAI/gua with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "LaplAI/gua"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "LaplAI/gua" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
gua
gua is a computer-use decision model for macOS. Given a screenshot, the task and the last few actions, it chooses what kind of action comes next and which on-screen element it acts on, and can also ground an element from a description. It answers as typed multiple-choice decisions and returns a probability for every option, so an agent can act on confident answers and defer on the rest.
It is Qwen/Qwen3.5-4B fine-tuned (supervised
fine-tuning with LoRA, rank 8 on every linear layer of the language model) on GUI
trajectories extracted from screen-recorded tutorials of seven kinds of software: macOS
system apps, video editing, Blender, Godot, Logic Pro, OBS Studio and Plasticity. This repository holds the
LoRA merged into the base weights (MLX, bf16) and, in adapter/, the LoRA itself.
Research model. The training data was derived from public YouTube tutorials and is not released. See Training data and Limitations.
How it answers
The model does not generate free text. Each step is asked as up to three questions, all sharing one prompt prefix (system prompt, screenshot, task, past actions, numbered element list), and the probability of every option is read from the model's next-token distribution:
| Question | Options | Use |
|---|---|---|
action |
click, double-click, right-click, drag, type, press, shortcut, scroll, hover, wait, finish | the kind of the next action |
next_target |
the numbered elements on screen | where the next action goes |
grounding |
the numbered elements on screen | which element matches a description you give |
Elements are candidates found on the screenshot: text lines from the macOS Vision framework, and, during training and evaluation, icons from the OmniParser v2 icon detector. The model chooses among candidates; it cannot point at something no candidate covers. When running on a live Mac, the accessibility tree is a better candidate source than detection.
The prompt format is fixed by training. Use inference/cua_decide.py, which reproduces it
exactly (checked to give identical option probabilities to the evaluation code).
Usage
Apple Silicon with mlx-vlm 0.7:
pip install "mlx-vlm>=0.7,<0.8" pyobjc-framework-vision pyobjc-framework-quartz
python inference/cua_decide.py --model <user>/gua \
--image screen.png --task "Turn on Dark Mode" \
--past "Click the Apple menu" --past "Click System Settings" \
--describe "the Appearance item in the sidebar" \
--temperatures temperatures.json
Each question prints its choice and calibrated confidence, for example:
{"question": "action", "choice": "click", "confidence": 0.94, "action": "click"}
{"question": "next_target", "choice": "15", "confidence": 0.60, "element": {"text": "...", "centre": [0.033, 0.136], "box": [...]}}
To use the LoRA instead of the merged weights, pass the base model and the adapter:
--model mlx-community/Qwen3.5-4B-MLX-bf16 --adapter adapter.
OmniParser's icon detector improves candidate recall but its weights are AGPL-3.0 and are
not included; pass --icon-weights icon_detect/model.pt if you download them yourself.
Without it, only text candidates are offered, which lowers recall on icon-only controls.
Training
gua v2 continues the v1 adapter on a larger, more varied dataset.
| Method | Supervised fine-tuning with LoRA: rank 8, alpha 16, on every linear layer of the language model (q/k/v/o_proj, linear-attention in_proj_*/out_proj, MLP gate/up/down_proj); vision tower frozen |
| Objective | Log loss of the right option label (and end-of-turn token) only, never of the prompt |
| Stage 1 (v1) | 8,000 questions from macos-cua v2 (macOS system, video editing, Blender, Godot) |
| Stage 2 (v2) | 12,000 further questions from macos-cua v3, shared evenly across its seven topics (each topic at most about 1,940; smaller topics give all they have) |
| Optimiser | Adam, learning rate 5e-5, batch 1, gradient clipping 1.0 |
| Hardware | One Apple Silicon Mac (MLX), gradient checkpointing, about 110 h in total, peak 182 GB |
The recipe was chosen on held-out dev splits, never on an eval split. The released weights are the end of stage 2, which scored higher on the full v3 dev split (mean accuracy 0.453) than the checkpoint with the best interim dev check (0.442).
Training data
Trajectories were extracted automatically from screen-recorded software tutorials on YouTube: videos were searched and screened (recording quality, platform), segmented and annotated into step-by-step actions by vision-language models, each click was located on the frame, and before/after screenshots were cut. The training split of macos-cua v3, and the questions stage 2 drew from it:
| Topic | Videos | Steps | Questions used in stage 2 |
|---|---|---|---|
| Plasticity (CAD) | 162 | 16,656 | 1,937 |
| Godot 4 (game engine) | 79 | 6,271 | 1,937 |
| Blender (3D) | 40 | 3,833 | 1,937 |
| OBS Studio (recording) | 37 | 2,890 | 1,937 |
| Logic Pro (audio) | 43 | 2,020 | 1,937 |
| macOS video editing (Final Cut Pro, Premiere Pro, CapCut, iMovie) | 7 | 640 | 1,468 |
| macOS system and productivity apps | 6 | 420 | 844 |
Splits are by video (dev 44 videos, eval 99 videos, 38 applications appear only in eval); every video of the earlier v2 release kept its split, so v2's eval is part of v3's eval and was never trained on. Godot typing steps were left out of training: an audit found their labels mostly wrong. Many tutorials were recorded on Windows or full screen; steps that would be wrong on a Mac (Windows shortcuts, taskbar, native dialogs) were removed.
Label quality. Labels are model-generated, not human-verified. A stronger model audited a random sample per topic:
| Topic | Audited steps | Element located correctly | Action correct |
|---|---|---|---|
| Logic Pro | 150 | 94.0% | 84.0% |
| OBS Studio | 150 | 92.0% | 81.3% |
| macOS video editing | 395 | 92.4% | 78.2% |
| macOS system and productivity | 200 | 91.5% | 75.5% |
| Blender | 180 | 89.4% | 71.7% |
| Plasticity | 570 | 79.6% | 71.4% |
| Godot | 270 | 87.0% | 54.8% |
The dataset itself is not released.
Evaluation
macos-cua v3 eval split (9,972 steps from 99 videos never seen in training, 6,509 with an element target), gua v2 against gua v1 with the same prompt and candidates, 95% CI by bootstrap paired by trajectory:
| Question | gua v1 | gua v2 | Difference |
|---|---|---|---|
action accuracy |
0.507 | 0.537 | +0.030 (+0.021, +0.040) |
grounding accuracy |
0.516 | 0.524 | +0.008 (+0.002, +0.015) |
next_target accuracy |
0.303 | 0.323 | +0.020 (+0.012, +0.029) |
action and next_target both right |
0.182 | 0.216 |
On the 1,767 steps from applications never seen in training: action 0.602, grounding 0.628, next target 0.357. By topic (gua v2):
| Topic | action |
grounding |
next_target |
|---|---|---|---|
| OBS Studio | 0.806 | 0.803 | 0.453 |
| macOS system and productivity | 0.720 | 0.643 | 0.465 |
| Logic Pro | 0.687 | 0.566 | 0.341 |
| macOS video editing | 0.592 | 0.644 | 0.388 |
| Blender | 0.539 | 0.503 | 0.253 |
| Godot | 0.519 | 0.657 | 0.275 |
| Plasticity | 0.462 | 0.381 | 0.308 |
macos-cua v2 eval split (3,442 steps, the eval gua v1 was published with), against the untrained base:
| Question | Base | gua v1 | gua v2 |
|---|---|---|---|
action accuracy |
0.481 | 0.525 | 0.542 |
grounding accuracy |
0.443 | 0.603 | 0.605 |
next_target accuracy |
0.155 | 0.261 | 0.289 |
A candidate covers the target in 70.2% of v3's element steps (83.4% in v2's), which bounds
the element questions; Plasticity's small, dense CAD icons are often missed. Excluding the
Godot typing steps, whose labels are mostly wrong, v3 action accuracy is 0.563. Because
those steps were left out of training, the model rarely answers type.
Selective accuracy (v3 eval). Acting only on the most confident half of the answers: grounding 0.711, action 0.665, next target 0.444.
Calibration. temperatures.json holds one temperature per question, fitted on the v3
dev split by log loss and applied as p^(1/T) renormalised (inference/cua_decide.py --temperatures). Expected calibration error on the v3 eval split, which the fit never saw:
| Question | Temperature | ECE before | ECE after |
|---|---|---|---|
action |
2.15 | 0.294 | 0.060 |
grounding |
1.25 | 0.259 | 0.173 |
next_target |
1.20 | 0.253 | 0.141 |
Without scaling the model is overconfident on every question; after it, element answers remain somewhat overconfident.
Limitations
- macOS screenshots only, English UI and tasks.
- Candidate-bound. The model picks among candidates; if no candidate covers the target (about 30% of element steps in v3's eval, most often in CAD), the element questions have no right answer. On a live Mac, the accessibility tree is a better candidate source.
- Noisy labels, as audited above; action-type labels are the weakest, and Plasticity's element labels the least reliable.
- Weak on dense professional UIs: Plasticity grounding is 0.38.
- Not an agent by itself. It decides the action kind and target; the text to type, the
keys of a shortcut and drag end points are not predicted. It rarely predicts
type. - Merged weights round to bf16. Against base plus adapter, option probabilities differ
by at most 0.035 in total variation on our check; use
adapter/for exact reproduction. - Confidence is only meaningful after temperature scaling, and only for the question types and candidate sources it was fitted with.
Versions
| Version | Training | Notes |
|---|---|---|
| v2 (this) | v1 continued on 12,000 balanced questions from macos-cua v3 | adds Logic Pro, OBS Studio, Plasticity; better on every question |
| v1 | 8,000 questions from macos-cua v2 | first release; tagged v1 in this repository |
Licence and provenance
The weights are a derivative of Qwen3.5-4B (Apache-2.0) and are released under Apache-2.0
(LICENSE); the modification is the LoRA fine-tuning described above. The training data
was derived from public YouTube videos for research; no video frames are distributed here.
inference/cua_decide.py is released under the same licence. OmniParser v2 (optional at
inference) is licensed separately by its authors.
- Downloads last month
- 24
Quantized
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("LaplAI/gua") config = load_config("LaplAI/gua") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output)