AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.
AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.
AREX-2 is evaluated on algorithmic programming, machine-learning engineering, deep research, and general agentic reasoning. Results follow the protocols reported in the AREX-2 paper.
| Model | Params | Frontier-CS | MLE-Lite |
|---|---|---|---|
| Closed-weight models | |||
| GPT-5.6 Sol | - | 76.4 | 72.7 |
| Claude Opus 4.8 | - | 74.5 | 63.6 |
| GPT-5.5 | - | 72.1 | 68.2 |
| Gemini-3.1-Pro | - | 68.9 | - |
| Qwen3.7-Max | - | 61.9 | - |
| Open-weight models | |||
| Kimi-K3 | 2.8T | - | 72.7 |
| Naive-N0.5-Flash | 309B | - | 73.7 |
| DeepSeek-V4-Pro | 1.6T | 44.7 | 54.5 |
| DeepSeek-V4-Flash | 284B | 39.1 | 51.5 |
| Kimi-K2.7-Code | 1T | 54.7 | - |
| GLM-5.3-Flash | 320B | 50.4 | - |
| Kimi-K2.6 | 1T | 46.9 | 66.7 |
| Frontis-MA1-35B | 35B | - | 71.2 |
| BigBang-V1 | 35B | - | 59.1 |
| Qwen3.6-35B-A3B | 35B | 23.4 | 39.4 |
| AREX-2 | 27B | 70.7 | 81.8 |
Frontier-CS is the 188-task Agent Track. MLE-Lite reports Any Medal averaged over three seeds.
| Model | Params | BrowseComp | HLE | GAIA | DeepSearchQA |
|---|---|---|---|---|---|
| Frontier models | |||||
| GPT-5.6 Sol | - | 90.4 | 58.0* | - | - |
| GPT-5.6 Terra | - | 87.5 | - | - | - |
| GPT-5.6 Luna | - | 83.3 | - | - | - |
| Kimi-K3 | 2.8T | 91.2 | 56.0* | - | 95.0 |
| Claude Fable 5 | - | 88.0 | 64.5* | - | 94.2 |
| Claude Opus 4.8 | - | 84.3 | 57.9* | - | 93.1 |
| GPT-5.5 | - | 84.4 | 52.2* | 87.4 | - |
| Gemini-3.1-Pro | - | 85.9 | 51.4* | 80.6 | 93.3 |
| Large models (>40B) | |||||
| GLM-5 | 744B | 75.9 | 50.4 | 70.0 | - |
| Kimi-K2.6 | 1T | 83.2 | 54.0* | 80.6 | 92.5 |
| GLM-5.3-Flash | 320B | - | 55.3* | 78.8 | - |
| DeepSeek-V4-Flash | 284B | 73.2 | 45.1 | 57.5 | 90.6 |
| DeepSeek-V4-Pro | 1.6T | 83.4 | 48.2 | 71.1 | 88.7 |
| MiroThinker-1.7 | 235B | 74.0 | 42.9 | 82.7 | 72.1 |
| XYZ-Aquila-pro | 397B | 84.8 | 53.3 | - | 92.5 |
| Iris-pro | 397B | 88.6 | 56.4 | - | 92.9 |
| AREX-Base | 122B | 82.5 | 52.4 | 85.4 | 89.9 |
| Small models (≤40B) | |||||
| Tongyi-DeepResearch-30B | 30B | 43.4 | 32.9 | 70.9 | - |
| Qwen3.5-35B | 35B | 61.0 | 47.4 | 80.0 | 68.5 |
| XYZ-Aquila-mini | 35B | 78.8 | 51.1 | 97.1 | 89.5 |
| BigBang-V1 | 35B | 76.5 | 50.3 | - | - |
| Quest-35B | 35B | 64.6 | 37.2 | 80.8 | - |
| Apodex-1.0-mini | 35B | 71.5 | 46.8 | - | 82.2 |
| Agents-A1 | 35B | 75.5 | 47.6 | 96.0 | - |
| MiroThinker-1.7-mini | 30B | 67.9 | 36.4 | 80.3 | 67.9 |
| Iris-mini | 35B | 82.2 | 52.3 | - | 86.9 |
| AREX-Turbo | 4B | 70.7 | 40.6 | 81.6 | 78.5 |
| AREX-2 | 27B | 84.0 | 52.6 | 92.2 | 93.8 |
HLE values marked * are from the full HLE set; unmarked values use the text-only subset.
Use a recent Transformers release with Qwen3.8 support.
pip install -U torch transformers accelerate
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
model_id = "BAAI/AREX-2"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
)
messages = [{
"role": "user",
"content": "Propose a solution and explain how you would improve it over several rounds.",
}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt",
).to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(
outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True
))
AREX-2 is intended for research on long-horizon agents, iterative problem solving, machine-learning engineering, algorithmic coding, and tool-augmented deep research.
AREX-2 is released under the Apache License 2.0. Follow the terms and notices for the Qwen base model and any downstream data or tools.
@article{2026arex2,
title = {AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks},
author = {Qian, Hongjin and Li, Chaofan and Luo, Kun and Wei, Wenqing and Chen, Jianlyu and Lu, Shuqi and Hu, Yuyang and Xiao, Hongwang and Wang, Hui and Li, Chaozhuo and Ye, Qiwei and Dou, Zhicheng and Lian, Defu and Liu, Zheng},
journal = {arXiv preprint arXiv:2609.38288},
year = {2026},
url = {https://arxiv.org/abs/2609.38288}
}
AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.
AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.
AREX-2 is evaluated on algorithmic programming, machine-learning engineering, deep research, and general agentic reasoning. Results follow the protocols reported in the AREX-2 paper.
| Model | Params | Frontier-CS | MLE-Lite |
|---|---|---|---|
| Closed-weight models | |||
| GPT-5.6 Sol | - | 76.4 | 72.7 |
| Claude Opus 4.8 | - | 74.5 | 63.6 |
| GPT-5.5 | - | 72.1 | 68.2 |
| Gemini-3.1-Pro | - | 68.9 | - |
| Qwen3.7-Max | - | 61.9 | - |
| Open-weight models | |||
| Kimi-K3 | 2.8T | - | 72.7 |
| Naive-N0.5-Flash | 309B | - | 73.7 |
| DeepSeek-V4-Pro | 1.6T | 44.7 | 54.5 |
| DeepSeek-V4-Flash | 284B | 39.1 | 51.5 |
| Kimi-K2.7-Code | 1T | 54.7 | - |
| GLM-5.3-Flash | 320B | 50.4 | - |
| Kimi-K2.6 | 1T | 46.9 | 66.7 |
| Frontis-MA1-35B | 35B | - | 71.2 |
| BigBang-V1 | 35B | - | 59.1 |
| Qwen3.6-35B-A3B | 35B | 23.4 | 39.4 |
| AREX-2 | 27B | 70.7 | 81.8 |
Frontier-CS is the 188-task Agent Track. MLE-Lite reports Any Medal averaged over three seeds.
| Model | Params | BrowseComp | HLE | GAIA | DeepSearchQA |
|---|---|---|---|---|---|
| Frontier models | |||||
| GPT-5.6 Sol | - | 90.4 | 58.0* | - | - |
| GPT-5.6 Terra | - | 87.5 | - | - | - |
| GPT-5.6 Luna | - | 83.3 | - | - | - |
| Kimi-K3 | 2.8T | 91.2 | 56.0* | - | 95.0 |
| Claude Fable 5 | - | 88.0 | 64.5* | - | 94.2 |
| Claude Opus 4.8 | - | 84.3 | 57.9* | - | 93.1 |
| GPT-5.5 | - | 84.4 | 52.2* | 87.4 | - |
| Gemini-3.1-Pro | - | 85.9 | 51.4* | 80.6 | 93.3 |
| Large models (>40B) | |||||
| GLM-5 | 744B | 75.9 | 50.4 | 70.0 | - |
| Kimi-K2.6 | 1T | 83.2 | 54.0* | 80.6 | 92.5 |
| GLM-5.3-Flash | 320B | - | 55.3* | 78.8 | - |
| DeepSeek-V4-Flash | 284B | 73.2 | 45.1 | 57.5 | 90.6 |
| DeepSeek-V4-Pro | 1.6T | 83.4 | 48.2 | 71.1 | 88.7 |
| MiroThinker-1.7 | 235B | 74.0 | 42.9 | 82.7 | 72.1 |
| XYZ-Aquila-pro | 397B | 84.8 | 53.3 | - | 92.5 |
| Iris-pro | 397B | 88.6 | 56.4 | - | 92.9 |
| AREX-Base | 122B | 82.5 | 52.4 | 85.4 | 89.9 |
| Small models (≤40B) | |||||
| Tongyi-DeepResearch-30B | 30B | 43.4 | 32.9 | 70.9 | - |
| Qwen3.5-35B | 35B | 61.0 | 47.4 | 80.0 | 68.5 |
| XYZ-Aquila-mini | 35B | 78.8 | 51.1 | 97.1 | 89.5 |
| BigBang-V1 | 35B | 76.5 | 50.3 | - | - |
| Quest-35B | 35B | 64.6 | 37.2 | 80.8 | - |
| Apodex-1.0-mini | 35B | 71.5 | 46.8 | - | 82.2 |
| Agents-A1 | 35B | 75.5 | 47.6 | 96.0 | - |
| MiroThinker-1.7-mini | 30B | 67.9 | 36.4 | 80.3 | 67.9 |
| Iris-mini | 35B | 82.2 | 52.3 | - | 86.9 |
| AREX-Turbo | 4B | 70.7 | 40.6 | 81.6 | 78.5 |
| AREX-2 | 27B | 84.0 | 52.6 | 92.2 | 93.8 |
HLE values marked * are from the full HLE set; unmarked values use the text-only subset.
Use a recent Transformers release with Qwen3.8 support.
pip install -U torch transformers accelerate
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor
model_id = "BAAI/AREX-2"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
)
messages = [{
"role": "user",
"content": "Propose a solution and explain how you would improve it over several rounds.",
}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt",
).to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(
outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True
))
AREX-2 is intended for research on long-horizon agents, iterative problem solving, machine-learning engineering, algorithmic coding, and tool-augmented deep research.
AREX-2 is released under the Apache License 2.0. Follow the terms and notices for the Qwen base model and any downstream data or tools.
@article{2026arex2,
title = {AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks},
author = {Qian, Hongjin and Li, Chaofan and Luo, Kun and Wei, Wenqing and Chen, Jianlyu and Lu, Shuqi and Hu, Yuyang and Xiao, Hongwang and Wang, Hui and Li, Chaozhuo and Ye, Qiwei and Dou, Zhicheng and Lian, Defu and Liu, Zheng},
journal = {arXiv preprint arXiv:2609.38288},
year = {2026},
url = {https://arxiv.org/abs/2609.38288}
}