--- license: apache-2.0 library_name: transformers pipeline_tag: image-text-to-text base_model: Qwen/Qwen2.5-VL-7B-Instruct datasets: - SeerRay-Lab/GAIA-Dataset-v1.0 language: - en tags: - gui - gui-agent - critic-model - vision-language-model - test-time-scaling - best-of-n - qwen2.5-vl - arxiv:2601.18197 --- # ICM-r2 **ICM-r2** is the second-round **Intuitive Critic Model** introduced in [GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models](https://arxiv.org/abs/2601.18197).

📄 Paper  |  💻 Code  |  🤗 Dataset

ICM-r2 is a Qwen2.5-VL-7B-based GUI action critic. Given a global task instruction, previous action history, the current GUI screenshot, and a candidate action, the model predicts whether the candidate action is **`correct`** or **`wrong`**. > ICM-r2 is a critic, not a standalone GUI agent. It evaluates actions proposed by another GUI agent before they are executed. ## Model details | Property | Value | |---|---| | Base model | [Qwen2.5-VL-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct) | | Architecture | `Qwen2_5_VLForConditionalGeneration` | | Parameters | Approximately 8.3B | | Task | Binary GUI action correctness evaluation | | Input | Task instruction, action history, candidate action, screenshot | | Output | `correct` or `wrong` | | Training framework | [ms-swift](https://github.com/modelscope/ms-swift) | | License | Apache 2.0 | ICM-r2 is trained on the second round of the GAIA data flywheel. In addition to the initial real-agent action samples, round two incorporates challenging actions collected under guidance from the first-round ICM. ## Intended use ICM-r2 is intended for: - pre-execution validation of GUI-agent actions; - Best-of-N candidate-action selection; - GUI action correctness evaluation; - research on GUI agents, critic models, and test-time scaling. The action space used in GAIA includes `Click`, `Swipe`, `Type`, `Open`, `Home`, `Back`, `Enter`, and `Wait`. For click actions, we recommend including the action coordinates in the text and marking the proposed click location with a red circle in the screenshot, consistent with the released training data. ## Input and output format The recommended user input is: ```text The goal of the task (instruction): {instruction} Action (plan) history: {action_history} Current action of the agent: {candidate_action} Screenshot: ``` The expected output is one of: ```text correct ``` or: ```text wrong ``` Use this prompt format consistently. The model is trained to judge a provided action, not to generate the next GUI action. ## Quick start ### Installation ```bash pip install "transformers>=4.51.3" accelerate qwen-vl-utils pillow ``` FlashAttention is optional but recommended on compatible GPUs. ### Inference ```python import torch from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration from qwen_vl_utils import process_vision_info model_id = "SeerRay-Lab/ICM-r2" image_path = "screenshot_with_candidate_action.png" system_prompt = """You are an expert in evaluating the performance of a phone operating agent. The agent is designed to help a user to complete a task or retrieve information from the phone. Given the user's task instruction, current action and current screenshot, your goal is to decide whether the agent's current action is correct or not. Each action in the sequence is preceded by a corresponding screenshot that captures the context in which the action occurs. ## Evaluation Criteria Whether the agent's current action is correct and corresponding to the user's task instruction. ## IMPORTANT 1. An action always follows a corresponding screenshot (even if only the last few are provided). 2. If the current action is a tap on the screen, the point where the action is clicked is marked with a red circle on the screenshot. 3. Answer only `correct` or `wrong`. ## Input The input includes global_task_instruction, action_history, current_action, and screenshot. """ user_prompt = """The goal of the task (instruction): Open Settings and enable Wi-Fi. Action (plan) history: Step 1: Return to the home screen. Current action of the agent: Tap at [420, 760] to open Settings. Screenshot:""" messages = [ { "role": "system", "content": [{"type": "text", "text": system_prompt}], }, { "role": "user", "content": [ {"type": "text", "text": user_prompt}, {"type": "image", "image": image_path}, ], }, ] model = Qwen2_5_VLForConditionalGeneration.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", ).eval() processor = AutoProcessor.from_pretrained( model_id, max_pixels=3600 * 28 * 28, ) text = processor.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, ) image_inputs, video_inputs = process_vision_info(messages) inputs = processor( text=[text], images=image_inputs, videos=video_inputs, padding=True, return_tensors="pt", ).to(model.device) with torch.inference_mode(): generated_ids = model.generate( **inputs, do_sample=False, max_new_tokens=16, ) generated_ids = [ output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, generated_ids) ] response = processor.batch_decode( generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False, )[0].strip() print(response) # correct or wrong ``` A multi-GPU evaluation implementation is available in [`src/infer_critic.py`](https://github.com/SeerRay-Lab/GAIA/blob/main/src/infer_critic.py). ## Best-of-N action selection At test time, a GUI agent can generate multiple candidate actions. ICM-r2 evaluates each candidate before execution: 1. Generate `N` candidate actions with the actor. 2. Evaluate every candidate using the same screenshot and task context. 3. Retain candidates judged as `correct`. 4. Select the candidate with the highest correctness confidence. 5. If no candidate is judged correct, fall back to the actor's first candidate. The experiments in the GAIA paper use `N = 8`. Example integration code is available in [`benchmark_screenspot_best_of_n_critic.py`](https://github.com/SeerRay-Lab/GAIA/blob/main/src/benchmark_screenspot_best_of_n_critic.py). ## Training data The model was trained with real action trajectories collected from GUI agents on AndroidControl and GUI-Odyssey. | Round | Source | Positive samples | Negative samples | |---|---|---:|---:| | Initial data | AndroidControl | 68.2K | 69.9K | | Initial data | GUI-Odyssey | 65.4K | 66.8K | | Round-two additions | AndroidControl | 15.1K | 14.0K | | Round-two additions | GUI-Odyssey | 26.1K | 26.3K | Negative examples are derived from real GUI-agent mistakes rather than only randomly generated click locations. The released data is available at [SeerRay-Lab/GAIA-Dataset-v1.0](https://huggingface.co/datasets/SeerRay-Lab/GAIA-Dataset-v1.0). ## Evaluation ICM-r2 is evaluated as a test-time critic using Best-of-8 action selection. All values below are percentages reported in the paper. ### GUI-Odyssey with UI-TARS 1.5 | Method | Action Type | Grounding | Step Success Rate | |---|---:|---:|---:| | UI-TARS 1.5 | 71.1 | 44.6 | 32.9 | | + ICM | 78.2 | 52.9 | 47.8 | | + ICM-r2 | **80.2** | **53.5** | **50.2** | ### Critic accuracy on GUI-Odyssey | Critic | Critic accuracy | |---|---:| | RCM | 70.82 | | ICM | 83.19 | | ICM-r2 | **83.56** | The critic-accuracy comparison uses UI-TARS 1.5 as the base agent. ### ScreenSpot-v2 with Qwen2.5-VL-7B | Method | Average grounding accuracy | |---|---:| | Qwen2.5-VL-7B | 65.0 | | + ICM | 70.4 | | + ICM-r2 | **71.1** | These results measure an actor combined with critic-guided Best-of-N selection. They should not be interpreted as standalone GUI execution results. ## Limitations and risks - ICM-r2 predicts action correctness but does not guarantee that an action is correct or safe. - Performance may degrade on unseen applications, layouts, languages, action representations, resolutions, or very long trajectories. - Results are sensitive to the prompt format and the representation of click locations. - False positives may allow an incorrect action to be executed. - The model should not be the sole approval mechanism for payments, account changes, deletion, permission changes, or other high-risk operations. - Screenshots may contain private information. Users are responsible for protecting sensitive data. We recommend explicit user confirmation before executing irreversible or security-sensitive actions. ## Citation ```bibtex @inproceedings{wang2026gaia, title={GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models}, author={Wang, Shaokang and Fu, Pei and Zhang, Ruoceng and Zhang, Shaojie and Xi, Xiuwen and Yang, Jiahui and Qin, Bin and Huang, Ying and Luo, Zhenbo and Luan, Jian}, booktitle={European Conference on Computer Vision}, year={2026}, organization={Springer} } ``` ## License The model is released under the Apache 2.0 License. Users should also comply with the licenses and terms of the base model, datasets, and third-party dependencies.