Image-Text-to-Text
Transformers
Safetensors
English
qwen2_5_vl
gui
gui-agent
critic-model
vision-language-model
test-time-scaling
best-of-n
qwen2.5-vl
conversational
text-generation-inference
Instructions to use SeerRay-Lab/ICM-r2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SeerRay-Lab/ICM-r2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="SeerRay-Lab/ICM-r2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("SeerRay-Lab/ICM-r2") model = AutoModelForMultimodalLM.from_pretrained("SeerRay-Lab/ICM-r2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SeerRay-Lab/ICM-r2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SeerRay-Lab/ICM-r2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/ICM-r2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/SeerRay-Lab/ICM-r2
- SGLang
How to use SeerRay-Lab/ICM-r2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SeerRay-Lab/ICM-r2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/ICM-r2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SeerRay-Lab/ICM-r2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/ICM-r2", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use SeerRay-Lab/ICM-r2 with Docker Model Runner:
docker model run hf.co/SeerRay-Lab/ICM-r2
File size: 9,468 Bytes
979d677 036ccd9 979d677 036ccd9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 | ---
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen2.5-VL-7B-Instruct
datasets:
- SeerRay-Lab/GAIA-Dataset-v1.0
language:
- en
tags:
- gui
- gui-agent
- critic-model
- vision-language-model
- test-time-scaling
- best-of-n
- qwen2.5-vl
- arxiv:2601.18197
---
# ICM-r2
**ICM-r2** is the second-round **Intuitive Critic Model** introduced in [GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models](https://arxiv.org/abs/2601.18197).
<p align="center">
📄 <a href="https://arxiv.org/abs/2601.18197">Paper</a> |
💻 <a href="https://github.com/SeerRay-Lab/GAIA">Code</a> |
🤗 <a href="https://huggingface.co/datasets/SeerRay-Lab/GAIA-Dataset-v1.0">Dataset</a>
</p>
ICM-r2 is a Qwen2.5-VL-7B-based GUI action critic. Given a global task instruction, previous action history, the current GUI screenshot, and a candidate action, the model predicts whether the candidate action is **`correct`** or **`wrong`**.
> ICM-r2 is a critic, not a standalone GUI agent. It evaluates actions proposed by another GUI agent before they are executed.
## Model details
| Property | Value |
|---|---|
| Base model | [Qwen2.5-VL-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct) |
| Architecture | `Qwen2_5_VLForConditionalGeneration` |
| Parameters | Approximately 8.3B |
| Task | Binary GUI action correctness evaluation |
| Input | Task instruction, action history, candidate action, screenshot |
| Output | `correct` or `wrong` |
| Training framework | [ms-swift](https://github.com/modelscope/ms-swift) |
| License | Apache 2.0 |
ICM-r2 is trained on the second round of the GAIA data flywheel. In addition to the initial real-agent action samples, round two incorporates challenging actions collected under guidance from the first-round ICM.
## Intended use
ICM-r2 is intended for:
- pre-execution validation of GUI-agent actions;
- Best-of-N candidate-action selection;
- GUI action correctness evaluation;
- research on GUI agents, critic models, and test-time scaling.
The action space used in GAIA includes `Click`, `Swipe`, `Type`, `Open`, `Home`, `Back`, `Enter`, and `Wait`.
For click actions, we recommend including the action coordinates in the text and marking the proposed click location with a red circle in the screenshot, consistent with the released training data.
## Input and output format
The recommended user input is:
```text
The goal of the task (instruction): {instruction}
Action (plan) history: {action_history}
Current action of the agent: {candidate_action}
Screenshot: <image>
```
The expected output is one of:
```text
correct
```
or:
```text
wrong
```
Use this prompt format consistently. The model is trained to judge a provided action, not to generate the next GUI action.
## Quick start
### Installation
```bash
pip install "transformers>=4.51.3" accelerate qwen-vl-utils pillow
```
FlashAttention is optional but recommended on compatible GPUs.
### Inference
```python
import torch
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
from qwen_vl_utils import process_vision_info
model_id = "SeerRay-Lab/ICM-r2"
image_path = "screenshot_with_candidate_action.png"
system_prompt = """You are an expert in evaluating the performance of a phone operating agent.
The agent is designed to help a user to complete a task or retrieve information from the phone.
Given the user's task instruction, current action and current screenshot, your goal is to decide whether the agent's current action is correct or not.
Each action in the sequence is preceded by a corresponding screenshot that captures the context in which the action occurs.
## Evaluation Criteria
Whether the agent's current action is correct and corresponding to the user's task instruction.
## IMPORTANT
1. An action always follows a corresponding screenshot (even if only the last few are provided).
2. If the current action is a tap on the screen, the point where the action is clicked is marked with a red circle on the screenshot.
3. Answer only `correct` or `wrong`.
## Input
The input includes global_task_instruction, action_history, current_action, and screenshot.
"""
user_prompt = """The goal of the task (instruction): Open Settings and enable Wi-Fi.
Action (plan) history: Step 1: Return to the home screen.
Current action of the agent: Tap at [420, 760] to open Settings.
Screenshot:"""
messages = [
{
"role": "system",
"content": [{"type": "text", "text": system_prompt}],
},
{
"role": "user",
"content": [
{"type": "text", "text": user_prompt},
{"type": "image", "image": image_path},
],
},
]
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
).eval()
processor = AutoProcessor.from_pretrained(
model_id,
max_pixels=3600 * 28 * 28,
)
text = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
generated_ids = model.generate(
**inputs,
do_sample=False,
max_new_tokens=16,
)
generated_ids = [
output_ids[len(input_ids):]
for input_ids, output_ids in zip(inputs.input_ids, generated_ids)
]
response = processor.batch_decode(
generated_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0].strip()
print(response) # correct or wrong
```
A multi-GPU evaluation implementation is available in [`src/infer_critic.py`](https://github.com/SeerRay-Lab/GAIA/blob/main/src/infer_critic.py).
## Best-of-N action selection
At test time, a GUI agent can generate multiple candidate actions. ICM-r2 evaluates each candidate before execution:
1. Generate `N` candidate actions with the actor.
2. Evaluate every candidate using the same screenshot and task context.
3. Retain candidates judged as `correct`.
4. Select the candidate with the highest correctness confidence.
5. If no candidate is judged correct, fall back to the actor's first candidate.
The experiments in the GAIA paper use `N = 8`. Example integration code is available in [`benchmark_screenspot_best_of_n_critic.py`](https://github.com/SeerRay-Lab/GAIA/blob/main/src/benchmark_screenspot_best_of_n_critic.py).
## Training data
The model was trained with real action trajectories collected from GUI agents on AndroidControl and GUI-Odyssey.
| Round | Source | Positive samples | Negative samples |
|---|---|---:|---:|
| Initial data | AndroidControl | 68.2K | 69.9K |
| Initial data | GUI-Odyssey | 65.4K | 66.8K |
| Round-two additions | AndroidControl | 15.1K | 14.0K |
| Round-two additions | GUI-Odyssey | 26.1K | 26.3K |
Negative examples are derived from real GUI-agent mistakes rather than only randomly generated click locations. The released data is available at [SeerRay-Lab/GAIA-Dataset-v1.0](https://huggingface.co/datasets/SeerRay-Lab/GAIA-Dataset-v1.0).
## Evaluation
ICM-r2 is evaluated as a test-time critic using Best-of-8 action selection. All values below are percentages reported in the paper.
### GUI-Odyssey with UI-TARS 1.5
| Method | Action Type | Grounding | Step Success Rate |
|---|---:|---:|---:|
| UI-TARS 1.5 | 71.1 | 44.6 | 32.9 |
| + ICM | 78.2 | 52.9 | 47.8 |
| + ICM-r2 | **80.2** | **53.5** | **50.2** |
### Critic accuracy on GUI-Odyssey
| Critic | Critic accuracy |
|---|---:|
| RCM | 70.82 |
| ICM | 83.19 |
| ICM-r2 | **83.56** |
The critic-accuracy comparison uses UI-TARS 1.5 as the base agent.
### ScreenSpot-v2 with Qwen2.5-VL-7B
| Method | Average grounding accuracy |
|---|---:|
| Qwen2.5-VL-7B | 65.0 |
| + ICM | 70.4 |
| + ICM-r2 | **71.1** |
These results measure an actor combined with critic-guided Best-of-N selection. They should not be interpreted as standalone GUI execution results.
## Limitations and risks
- ICM-r2 predicts action correctness but does not guarantee that an action is correct or safe.
- Performance may degrade on unseen applications, layouts, languages, action representations, resolutions, or very long trajectories.
- Results are sensitive to the prompt format and the representation of click locations.
- False positives may allow an incorrect action to be executed.
- The model should not be the sole approval mechanism for payments, account changes, deletion, permission changes, or other high-risk operations.
- Screenshots may contain private information. Users are responsible for protecting sensitive data.
We recommend explicit user confirmation before executing irreversible or security-sensitive actions.
## Citation
```bibtex
@inproceedings{wang2026gaia,
title={GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models},
author={Wang, Shaokang and Fu, Pei and Zhang, Ruoceng and Zhang, Shaojie and Xi, Xiuwen and Yang, Jiahui and Qin, Bin and Huang, Ying and Luo, Zhenbo and Luan, Jian},
booktitle={European Conference on Computer Vision},
year={2026},
organization={Springer}
}
```
## License
The model is released under the Apache 2.0 License. Users should also comply with the licenses and terms of the base model, datasets, and third-party dependencies.
|