BAAI/AREX-2

Model

Introduction

85

7 commits

updated Oct 1, 2026

See the code

README

AREX AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
Paper Homepage

Introduction

AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.

AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.

  • Architecture: Dense Qwen3.8-compatible multimodal model
  • Parameters: 27B
  • Context length: 262,144 tokens

Key features

  • Long-horizon self-improvement: turns extra test-time rounds into useful solution refinement.
  • Feedback-driven reflection: reads scores, logs, errors, and timings to decide what to change next.
  • Cross-domain performance: training on coding and machine-learning tasks also improves the model's deep-research performance.
  • Long-horizon reasoning: sustains productive iteration as the task budget grows.

Evaluation

AREX-2 is evaluated on algorithmic programming, machine-learning engineering, deep research, and general agentic reasoning. Results follow the protocols reported in the AREX-2 paper.

Coding and machine-learning engineering

ModelParamsFrontier-CSMLE-Lite
Closed-weight models
GPT-5.6 Sol-76.472.7
Claude Opus 4.8-74.563.6
GPT-5.5-72.168.2
Gemini-3.1-Pro-68.9-
Qwen3.7-Max-61.9-
Open-weight models
Kimi-K32.8T-72.7
Naive-N0.5-Flash309B-73.7
DeepSeek-V4-Pro1.6T44.754.5
DeepSeek-V4-Flash284B39.151.5
Kimi-K2.7-Code1T54.7-
GLM-5.3-Flash320B50.4-
Kimi-K2.61T46.966.7
Frontis-MA1-35B35B-71.2
BigBang-V135B-59.1
Qwen3.6-35B-A3B35B23.439.4
AREX-227B70.781.8

Frontier-CS is the 188-task Agent Track. MLE-Lite reports Any Medal averaged over three seeds.

General agentic reasoning and deep research

ModelParamsBrowseCompHLEGAIADeepSearchQA
Frontier models
GPT-5.6 Sol-90.458.0*--
GPT-5.6 Terra-87.5---
GPT-5.6 Luna-83.3---
Kimi-K32.8T91.256.0*-95.0
Claude Fable 5-88.064.5*-94.2
Claude Opus 4.8-84.357.9*-93.1
GPT-5.5-84.452.2*87.4-
Gemini-3.1-Pro-85.951.4*80.693.3
Large models (>40B)
GLM-5744B75.950.470.0-
Kimi-K2.61T83.254.0*80.692.5
GLM-5.3-Flash320B-55.3*78.8-
DeepSeek-V4-Flash284B73.245.157.590.6
DeepSeek-V4-Pro1.6T83.448.271.188.7
MiroThinker-1.7235B74.042.982.772.1
XYZ-Aquila-pro397B84.853.3-92.5
Iris-pro397B88.656.4-92.9
AREX-Base122B82.552.485.489.9
Small models (≤40B)
Tongyi-DeepResearch-30B30B43.432.970.9-
Qwen3.5-35B35B61.047.480.068.5
XYZ-Aquila-mini35B78.851.197.189.5
BigBang-V135B76.550.3--
Quest-35B35B64.637.280.8-
Apodex-1.0-mini35B71.546.8-82.2
Agents-A135B75.547.696.0-
MiroThinker-1.7-mini30B67.936.480.367.9
Iris-mini35B82.252.3-86.9
AREX-Turbo4B70.740.681.678.5
AREX-227B84.052.692.293.8

HLE values marked * are from the full HLE set; unmarked values use the text-only subset.

Inference

Use a recent Transformers release with Qwen3.8 support.

pip install -U torch transformers accelerate
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "BAAI/AREX-2"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
    model_id, dtype=torch.bfloat16, device_map="auto"
)

messages = [{
    "role": "user",
    "content": "Propose a solution and explain how you would improve it over several rounds.",
}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(
    outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True
))

Intended use

AREX-2 is intended for research on long-horizon agents, iterative problem solving, machine-learning engineering, algorithmic coding, and tool-augmented deep research.

License

AREX-2 is released under the Apache License 2.0. Follow the terms and notices for the Qwen base model and any downstream data or tools.

Citation

@article{2026arex2,
  title   = {AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks},
  author  = {Qian, Hongjin and Li, Chaofan and Luo, Kun and Wei, Wenqing and Chen, Jianlyu and Lu, Shuqi and Hu, Yuyang and Xiao, Hongwang and Wang, Hui and Li, Chaozhuo and Ye, Qiwei and Dou, Zhicheng and Lian, Defu and Liu, Zheng},
  journal = {arXiv preprint arXiv:2609.38288},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.38288}
}
agent
conversational
deep-research
endpoints_compatible
image-text-to-text
long-context
qwen3_5
reasoning
safetensors
self-improvement
tool-use
transformers

BAAI/AREX-2

Model

Introduction

85

7 commits

updated Oct 1, 2026

See the code

README

AREX AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
Paper Homepage

Introduction

AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.

AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.

  • Architecture: Dense Qwen3.8-compatible multimodal model
  • Parameters: 27B
  • Context length: 262,144 tokens

Key features

  • Long-horizon self-improvement: turns extra test-time rounds into useful solution refinement.
  • Feedback-driven reflection: reads scores, logs, errors, and timings to decide what to change next.
  • Cross-domain performance: training on coding and machine-learning tasks also improves the model's deep-research performance.
  • Long-horizon reasoning: sustains productive iteration as the task budget grows.

Evaluation

AREX-2 is evaluated on algorithmic programming, machine-learning engineering, deep research, and general agentic reasoning. Results follow the protocols reported in the AREX-2 paper.

Coding and machine-learning engineering

ModelParamsFrontier-CSMLE-Lite
Closed-weight models
GPT-5.6 Sol-76.472.7
Claude Opus 4.8-74.563.6
GPT-5.5-72.168.2
Gemini-3.1-Pro-68.9-
Qwen3.7-Max-61.9-
Open-weight models
Kimi-K32.8T-72.7
Naive-N0.5-Flash309B-73.7
DeepSeek-V4-Pro1.6T44.754.5
DeepSeek-V4-Flash284B39.151.5
Kimi-K2.7-Code1T54.7-
GLM-5.3-Flash320B50.4-
Kimi-K2.61T46.966.7
Frontis-MA1-35B35B-71.2
BigBang-V135B-59.1
Qwen3.6-35B-A3B35B23.439.4
AREX-227B70.781.8

Frontier-CS is the 188-task Agent Track. MLE-Lite reports Any Medal averaged over three seeds.

General agentic reasoning and deep research

ModelParamsBrowseCompHLEGAIADeepSearchQA
Frontier models
GPT-5.6 Sol-90.458.0*--
GPT-5.6 Terra-87.5---
GPT-5.6 Luna-83.3---
Kimi-K32.8T91.256.0*-95.0
Claude Fable 5-88.064.5*-94.2
Claude Opus 4.8-84.357.9*-93.1
GPT-5.5-84.452.2*87.4-
Gemini-3.1-Pro-85.951.4*80.693.3
Large models (>40B)
GLM-5744B75.950.470.0-
Kimi-K2.61T83.254.0*80.692.5
GLM-5.3-Flash320B-55.3*78.8-
DeepSeek-V4-Flash284B73.245.157.590.6
DeepSeek-V4-Pro1.6T83.448.271.188.7
MiroThinker-1.7235B74.042.982.772.1
XYZ-Aquila-pro397B84.853.3-92.5
Iris-pro397B88.656.4-92.9
AREX-Base122B82.552.485.489.9
Small models (≤40B)
Tongyi-DeepResearch-30B30B43.432.970.9-
Qwen3.5-35B35B61.047.480.068.5
XYZ-Aquila-mini35B78.851.197.189.5
BigBang-V135B76.550.3--
Quest-35B35B64.637.280.8-
Apodex-1.0-mini35B71.546.8-82.2
Agents-A135B75.547.696.0-
MiroThinker-1.7-mini30B67.936.480.367.9
Iris-mini35B82.252.3-86.9
AREX-Turbo4B70.740.681.678.5
AREX-227B84.052.692.293.8

HLE values marked * are from the full HLE set; unmarked values use the text-only subset.

Inference

Use a recent Transformers release with Qwen3.8 support.

pip install -U torch transformers accelerate
import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "BAAI/AREX-2"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
    model_id, dtype=torch.bfloat16, device_map="auto"
)

messages = [{
    "role": "user",
    "content": "Propose a solution and explain how you would improve it over several rounds.",
}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(**inputs, max_new_tokens=1024)
print(processor.decode(
    outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True
))

Intended use

AREX-2 is intended for research on long-horizon agents, iterative problem solving, machine-learning engineering, algorithmic coding, and tool-augmented deep research.

License

AREX-2 is released under the Apache License 2.0. Follow the terms and notices for the Qwen base model and any downstream data or tools.

Citation

@article{2026arex2,
  title   = {AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks},
  author  = {Qian, Hongjin and Li, Chaofan and Luo, Kun and Wei, Wenqing and Chen, Jianlyu and Lu, Shuqi and Hu, Yuyang and Xiao, Hongwang and Wang, Hui and Li, Chaozhuo and Ye, Qiwei and Dou, Zhicheng and Lian, Defu and Liu, Zheng},
  journal = {arXiv preprint arXiv:2609.38288},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.38288}
}
agent
conversational
deep-research
endpoints_compatible
image-text-to-text
long-context
qwen3_5
reasoning
safetensors
self-improvement
tool-use
transformers