duvo-eye-2

Duvo Evaluation traces

Model summary

duvo-eye-2 is a vision-language model for GUI grounding, built on Holo-3.1-35B-A3B and developed by Duvo. Give it a screenshot and a description of an element, and it returns where to click, or tells you the element isn't on the screen. It is built to be the eyes of computer-use agents that run business processes across enterprise software.

SpecificationValue
Model IDduvoai/duvo-eye-2
ArchitectureQwen3.5 MoE, 35B total, 3B active parameters
Checkpoint formatBF16 safetensors
Output{"x", "y"} in [0, 1000], or {"x": -1, "y": -1} if the element is not on screen

Performance

duvo-eye-2 is the most accurate single-pass model on the ScreenSpot-Pro leaderboard: 75.2% from one forward pass, ahead of every single-pass entry, including 8B and 32B dense models, with 3B active parameters. With one zoom pass it reaches 79.7%, #5 overall among 97 entries.

On OSWorld-G it scores 75.2%, the best single-pass result published, and it knows when the element it is asked for does not exist: it declines 23 of the benchmark's 54 infeasible tasks instead of clicking something else.

Benchmark results

Benchmarkduvo-eye-2duvo-eye-1.5Best published single-pass
ScreenSpot-Pro75.273.673.4 · Indeed-UI-8B
ScreenSpot-Pro + zoom79.778.6–
OSWorld-G (564 tasks)75.273.470.6 · UI-Venus-1.5-30B-A3B
UI-I2E-Bench87.584.887.3 · UI-Ins-32B¹
SynthUI²91.789.4–

¹ With reasoning. ² Duvo's private benchmark of enterprise back-office UIs.

ScreenSpot-Pro is measured with its official harness; the other benchmarks with Duvo's harness, on all samples, temperature 0.

Open-source evaluation traces

For transparency, we share every prediction behind these numbers in duvoai/duvo-eye-2-evals.

Usage

Serve with vLLM and send a screenshot with the grounding prompt. Thinking must be disabled: duvo-eye-2 answers with the coordinate directly.

vllm serve duvoai/duvo-eye-2 --max-model-len 20480 --mm-processor-kwargs '{"max_pixels": 10000000}'
import base64, json, re
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
screenshot = base64.b64encode(open("screenshot.png", "rb").read()).decode()
prompt = (
    "Localize an element on the GUI image according to the provided target and output a click position.\n"
    ' * You must output a valid JSON following the format: {"x": int 0-1000, "y": int 0-1000}\n'
    " Your target is:\nthe Post button in the invoice toolbar"
)
reply = client.chat.completions.create(
    model="duvoai/duvo-eye-2",
    messages=[{"role": "user", "content": [
        {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{screenshot}"}},
        {"type": "text", "text": prompt}]}],
    temperature=0, max_tokens=64,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
point = json.loads(re.search(r"\{.*?\}", reply.choices[0].message.content).group(0))
# {"x": -1, "y": -1}: the element is not on the screen. Otherwise scale x and y from [0, 1000] to pixels.

For 4K screenshots, keep max_pixels high so small icons stay legible. The BF16 weights are about 66 GB: serve on one GPU with 141 GB or more, or two 80 GB GPUs.

Training

duvo-eye-2 continues duvo-eye-1.5 with supervised fine-tuning on Duvo's SynthUI corpus, taught to answer "not on screen" and to refine zoomed views, then two rounds of reinforcement learning (GRPO) on real desktop applications from ServiceNow/GroundCUA. The second round trains on high-resolution screens with a reward that pays for landing inside small targets and penalises refusing an element that is there.

License

The model weights are available under the Apache License 2.0. This model is built on Holo-3.1-35B-A3B, which H Company also releases under the Apache License 2.0.

Downloads last month
29
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for duvoai/duvo-eye-2

Finetuned
(2)
this model