Instructions to use deskmind/eyes-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use deskmind/eyes-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="deskmind/eyes-4b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("deskmind/eyes-4b") model = AutoModelForMultimodalLM.from_pretrained("deskmind/eyes-4b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use deskmind/eyes-4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "deskmind/eyes-4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deskmind/eyes-4b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/deskmind/eyes-4b
- SGLang
How to use deskmind/eyes-4b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "deskmind/eyes-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deskmind/eyes-4b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "deskmind/eyes-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deskmind/eyes-4b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use deskmind/eyes-4b with Docker Model Runner:
docker model run hf.co/deskmind/eyes-4b
DeskMind Eyes 4B
A 4B GUI grounding model: given a screenshot and a description of a UI element, it returns the point to click. DeskMind uses it on this Mac for apps that expose no accessibility tree.
一个 4B 的 GUI 定位模型:输入截图和对界面元素的描述,输出要点击的位置。 得心(DeskMind)在本机用它操作没有辅助功能结构的应用。
| Revision | Contents |
|---|---|
main |
bf16 weights (Transformers / vLLM), for reproducing the results below |
mlx-4bit |
MLX 4-bit build (group size 64) that the DeskMind app runs on Apple silicon |
Results · 结果
The benchmarks were run with the bf16 weights, all on one platform: vLLM 0.19, the Qwen3-VL computer_use tool prompt, native resolution and greedy decoding. Comparisons with other models are paired, item by item, on that platform.
用 bf16 权重评测,所有结果都在同一平台上得出:vLLM 0.19、Qwen3-VL computer_use tool prompt、原生分辨率、贪心解码。与其他模型的对比是在该平台上逐题配对进行的。
| Benchmark | GUI-Owl-1.5-4B (base) | Eyes 4B |
|---|---|---|
| ScreenSpot-Pro, no zoom, full 1581 | 64.8 | 67.7 |
| ScreenSpot-Pro, with zoom-in, full | 76.2 | 77.5 |
| ScreenSpot-v2, 1272 | 92.8 | 95.0 |
On ScreenSpot-Pro (no zoom), Eyes 4B is +2.9 over the base (paired: +81 / −34, z = 4.4). It is +1.6 over KV-Ground-4B measured on the same platform (+68 / −42, z = 2.5).
The
mlx-4bitbuild as DeskMind deploys it (4-bit, screenshots scaled to ≤ 2 MP,point_2dprompt) has no benchmark number. Quantization and the smaller input may cost accuracy.ScreenSpot-Pro(不放大)上比基座高 2.9(配对 +81 / −34,z = 4.4);比在同一平台实测的 KV-Ground-4B 高 1.6(+68 / −42,z = 2.5)。
DeskMind 实际部署的
mlx-4bit版本(4-bit、截图缩到 ≤ 2 MP、point_2d提示词)没有评测分数。量化和更小的输入都可能降低准确率。
Training · 训练
The base is GUI-Owl-1.5-4B-Instruct. On top of it, a LoRA (r = 32, on all language-model linear layers and lm_head, vision tower frozen) was trained with GRPO and DAPO dynamic sampling, rewarded 1 for a click inside the target box and 0 otherwise. The LoRA is merged into the weights published here.
- DAPO on ShowUI-desktop and OS-Atlas desktop, with the OS-Atlas instructions rewritten into functional ones (see Disclosures).
- Continued DAPO on a harder pool of GroundCUA records, 3.3K with a base pass rate in [1/8, 4/8], at 2.5 MP.
基座是 GUI-Owl-1.5-4B-Instruct。在其上训练了一个 LoRA(r = 32,作用于语言模型的全部线性层和 lm_head,视觉塔冻结),用 GRPO 和 DAPO 动态采样;点击落在目标框内奖励为 1,否则为 0。这里发布的权重已经合并了该 LoRA。
- 在 ShowUI-desktop 和 OS-Atlas desktop 上做 DAPO。OS-Atlas 的指令改写成了功能描述式(见"说明"一节)。
- 在更难的 GroundCUA 子集上继续 DAPO:3.3K 条基座通过率在 [1/8, 4/8] 之间的记录,训练分辨率 2.5 MP。
Licence · 许可
The weights are released under Apache-2.0.
| Component | Licence |
|---|---|
| Base model GUI-Owl-1.5-4B-Instruct (mPLUG) | MIT: its notice is kept below |
| Qwen3-VL (the base of GUI-Owl-1.5) | Apache-2.0 |
| OS-Atlas data (OS-Copilot) | Apache-2.0 |
| GroundCUA (ServiceNow) | MIT |
| ShowUI-desktop (showlab) | not stated: see Disclosures |
权重以 Apache-2.0 发布。各组成部分的许可见上表:基座模型的 MIT 声明保留在本节末尾;ShowUI-desktop 的许可未标明,见"说明"一节。
GUI-Owl-1.5 — Copyright (c) the mPLUG / X-PLUG authors. Released under the MIT License: permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files, to deal in the Software without restriction, subject to the condition that the above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND.
Disclosures · 说明
ShowUI-desktop's licence is unconfirmed. Its dataset card states none, and we have asked the authors (showlab/ShowUI#104). If it turns out not to allow this use, a variant trained without it will replace these weights.
Rewritten instructions. The OS-Atlas desktop instructions (accessibility names) were rewritten into functional instructions by a frontier model. That rewriting prompt used a handful of ScreenSpot-Pro test instructions, text only, as style examples: no screenshots, boxes or answers. The rewritten data was used for training, not for evaluation, but the prompt did see those test strings.
UI-Vision is not reported here. GroundCUA shares its source with UI-Vision.
ShowUI-desktop 的许可尚未确认。 数据卡上没有写许可,我们已经向作者询问(showlab/ShowUI#104)。如果它不允许这种用途,会用不含它的数据重训一版,替换这里的权重。
改写过的指令。 OS-Atlas desktop 的指令(辅助功能名称)由一个前沿大模型改写成了功能描述式。改写用的提示词里,拿了几条 ScreenSpot-Pro 测试集的指令(仅文字)当风格示例,没有用截图、标注框或答案。改写后的数据只用于训练、不用于评测,但提示词确实见过这几条测试文字。
这里不报告 UI-Vision 分数。 因为 GroundCUA 与 UI-Vision 同源。
Use · 用法
Prompt (coordinates are 0–1000, relative to the image):
提示词(坐标为 0–1000,相对于图片):
Locate the UI element in the screenshot that matches the description, and output its click point in JSON as
{"point_2d": [x, y]}, with coordinates in 0-1000 relative to the image.
Description: <instruction>
pip install mlx-vlm huggingface_hub
hf download deskmind/eyes-4b --revision mlx-4bit --local-dir eyes-4b-mlx
python -m mlx_vlm.generate --model eyes-4b-mlx --max-tokens 32 --image screenshot.png --prompt "<the prompt above>"
Integrity: eyes-4b-bf16.sha256 on main and eyes-4b-mlx4.sha256 on mlx-4bit list every file's SHA-256.
完整性校验:main 上的 eyes-4b-bf16.sha256 和 mlx-4bit 上的 eyes-4b-mlx4.sha256 列出了每个文件的 SHA-256。
- Downloads last month
- 1
Model tree for deskmind/eyes-4b
Base model
mPLUG/GUI-Owl-1.5-4B-Instruct