AREX AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
Paper Homepage

Introduction

AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.

AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.

  • Architecture: Dense Qwen3.8-compatible multimodal model
  • Parameters: 27B
  • Context length: 262,144 tokens

Key features

  • Long-horizon self-improvement: turns extra test-time rounds into useful solution refinement.
  • Feedback-driven reflection: reads scores, logs, errors, and timings to decide what to change next.
  • Cross-domain performance: training on coding and machine-learning tasks also improves the model's deep-research performance.
  • Long-horizon reasoning: sustains productive iteration as the task budget grows.

Evaluation

AREX-2 is evaluated on algorithmic programming, machine-learning engineering, deep research, and general agentic reasoning. Results follow the protocols reported in the AREX-2 paper.

Coding and machine-learning engineering

Model Params Frontier-CS MLE-Lite
Closed-weight models
GPT-5.6 Sol-76.472.7
Claude Opus 4.8-74.563.6
GPT-5.5-72.168.2
Gemini-3.1-Pro-68.9-
Qwen3.7-Max-61.9-
Open-weight models
Kimi-K32.8T-72.7
Naive-N0.5-Flash309B-73.7
DeepSeek-V4-Pro1.6T44.754.5
DeepSeek-V4-Flash284B39.151.5
Kimi-K2.7-Code1T54.7-
GLM-5.3-Flash320B50.4-
Kimi-K2.61T46.966.7
Frontis-MA1-35B35B-71.2
BigBang-V135B-59.1
Qwen3.6-35B-A3B35B23.439.4
AREX-227B70.781.8

Frontier-CS is the 188-task Agent Track. MLE-Lite reports Any Medal averaged over three seeds.

General agentic reasoning and deep research

Model Params BrowseComp HLE GAIA DeepSearchQA
Frontier models
GPT-5.6 Sol-90.458.0*--
GPT-5.6 Terra-87.5---
GPT-5.6 Luna-83.3---
Kimi-K32.8T91.256.0*-95.0
Claude Fable 5-88.064.5*-94.2
Claude Opus 4.8-84.357.9*-93.1
GPT-5.5-84.452.2*87.4-
Gemini-3.1-Pro-85.951.4*80.693.3
Large models (>40B)
GLM-5744B75.950.470.0-
Kimi-K2.61T83.254.0*80.692.5
GLM-5.3-Flash320B-55.3*78.8-
DeepSeek-V4-Flash284B73.245.157.590.6
DeepSeek-V4-Pro1.6T83.448.271.188.7
MiroThinker-1.7235B74.042.982.772.1
XYZ-Aquila-pro397B84.853.3-92.5
Iris-pro397B88.656.4-92.9
AREX-Base122B82.552.485.489.9
Small models (≤40B)
Tongyi-DeepResearch-30B30B43.432.970.9-
Qwen3.5-35B35B61.047.480.068.5
XYZ-Aquila-mini35B78.851.197.189.5
BigBang-V135B76.550.3--
Quest-35B35B64.637.280.8-
Apodex-1.0-mini35B71.546.8-82.2
Agents-A135B75.547.696.0-
MiroThinker-1.7-mini30B67.936.480.367.9
Iris-mini35B82.252.3-86.9
AREX-Turbo4B70.740.681.678.5
AREX-227B84.052.692.293.8

HLE values marked * are from the full HLE set; unmarked values use the text-only subset.

Inference

Use a recent Transformers release with Qwen3.8 support.

pip install -U torch transformers accelerate
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "BAAI/AREX-2"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
    model_id, dtype=torch.bfloat16, device_map="auto"
)

messages = [{
    "role": "user",
    "content": "Propose a solution and explain how you would improve it over several rounds.",
}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(
    outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True
))

Intended use

AREX-2 is intended for research on long-horizon agents, iterative problem solving, machine-learning engineering, algorithmic coding, and tool-augmented deep research.

License

AREX-2 is released under the Apache License 2.0. Follow the terms and notices for the Qwen base model and any downstream data or tools.

Citation

@misc{baai2026arex,
  title={AREX: Towards a Recursively Self-Improving Agent for Deep Research},
  author={Lu, Shuqi and Li, Chaofan and Luo, Kun and Zhang, Zhang and Wang, Hui
          and Xiao, Hongwang and Xiong, Lei and Wang, Jiahao and Wang, Sen
          and Jiang, Xiyan and Li, Wanli and Hu, Yuyang and Qian, Hongjin
          and Yan, Bingyu and Xia, Ziyi and Shao, Yingxia and Liu, Kang
          and Dou, Zhicheng and He, Di and Li, Chaozhuo and Ye, Qiwei
          and Wang, Zhongyuan and Liu, Zheng},
  year={2026},
  eprint={2607.21461},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2607.21461}
}
Downloads last month
-
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BAAI/AREX-2

Base model

Qwen/Qwen3.8-27B
Finetuned
(423)
this model

Collection including BAAI/AREX-2

Paper for BAAI/AREX-2