Instructions to use duvoai/duvo-eye-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use duvoai/duvo-eye-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="duvoai/duvo-eye-2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("duvoai/duvo-eye-2") model = AutoModelForMultimodalLM.from_pretrained("duvoai/duvo-eye-2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use duvoai/duvo-eye-2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "duvoai/duvo-eye-2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "duvoai/duvo-eye-2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/duvoai/duvo-eye-2
- SGLang
How to use duvoai/duvo-eye-2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "duvoai/duvo-eye-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "duvoai/duvo-eye-2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "duvoai/duvo-eye-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "duvoai/duvo-eye-2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use duvoai/duvo-eye-2 with Docker Model Runner:
docker model run hf.co/duvoai/duvo-eye-2
duvo-eye-2
Model summary
duvo-eye-2 is a vision-language model for GUI grounding, built on Holo-3.1-35B-A3B and developed by Duvo. Give it a screenshot and a description of an element, and it returns where to click, or tells you the element isn't on the screen. It is built to be the eyes of computer-use agents that run business processes across enterprise software.
| Specification | Value |
|---|---|
| Model ID | duvoai/duvo-eye-2 |
| Architecture | Qwen3.5 MoE, 35B total, 3B active parameters |
| Checkpoint format | BF16 safetensors |
| Output | {"x", "y"} in [0, 1000], or {"x": -1, "y": -1} if the element is not on screen |
Performance
duvo-eye-2 is the most accurate single-pass model on the ScreenSpot-Pro leaderboard: 75.2% from one forward pass, ahead of every single-pass entry, including 8B and 32B dense models, with 3B active parameters. With one zoom pass it reaches 79.7%, #5 overall among 97 entries.
On OSWorld-G it scores 75.2%, the best single-pass result published, and it knows when the element it is asked for does not exist: it declines 23 of the benchmark's 54 infeasible tasks instead of clicking something else.
Benchmark results
| Benchmark | duvo-eye-2 | duvo-eye-1.5 | Best published single-pass |
|---|---|---|---|
| ScreenSpot-Pro | 75.2 | 73.6 | 73.4 · Indeed-UI-8B |
| ScreenSpot-Pro + zoom | 79.7 | 78.6 | – |
| OSWorld-G (564 tasks) | 75.2 | 73.4 | 70.6 · UI-Venus-1.5-30B-A3B |
| UI-I2E-Bench | 87.5 | 84.8 | 87.3 · UI-Ins-32B¹ |
| SynthUI² | 91.7 | 89.4 | – |
¹ With reasoning. ² Duvo's private benchmark of enterprise back-office UIs.
ScreenSpot-Pro is measured with its official harness; the other benchmarks with Duvo's harness, on all samples, temperature 0.
Open-source evaluation traces
For transparency, we share every prediction behind these numbers in duvoai/duvo-eye-2-evals.
Usage
Serve with vLLM and send a screenshot with the grounding prompt. Thinking must be disabled: duvo-eye-2 answers with the coordinate directly.
vllm serve duvoai/duvo-eye-2 --max-model-len 20480 --mm-processor-kwargs '{"max_pixels": 10000000}'
import base64, json, re
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
screenshot = base64.b64encode(open("screenshot.png", "rb").read()).decode()
prompt = (
"Localize an element on the GUI image according to the provided target and output a click position.\n"
' * You must output a valid JSON following the format: {"x": int 0-1000, "y": int 0-1000}\n'
" Your target is:\nthe Post button in the invoice toolbar"
)
reply = client.chat.completions.create(
model="duvoai/duvo-eye-2",
messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{screenshot}"}},
{"type": "text", "text": prompt}]}],
temperature=0, max_tokens=64,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
point = json.loads(re.search(r"\{.*?\}", reply.choices[0].message.content).group(0))
# {"x": -1, "y": -1}: the element is not on the screen. Otherwise scale x and y from [0, 1000] to pixels.
For 4K screenshots, keep max_pixels high so small icons stay legible. The BF16 weights are about 66 GB: serve on one GPU with 141 GB or more, or two 80 GB GPUs.
Training
duvo-eye-2 continues duvo-eye-1.5 with supervised fine-tuning on Duvo's SynthUI corpus, taught to answer "not on screen" and to refine zoomed views, then two rounds of reinforcement learning (GRPO) on real desktop applications from ServiceNow/GroundCUA. The second round trains on high-resolution screens with a reward that pays for landing inside small targets and penalises refusing an element that is there.
License
The model weights are available under the Apache License 2.0. This model is built on Holo-3.1-35B-A3B, which H Company also releases under the Apache License 2.0.
- Downloads last month
- 29