The OpAgent framework is used to automate web browser operations, achieving State-of-the-Art (SOTA) performance on the WebArena benchmark.
Python
235
44 commits
updated Sep 15, 2026

OpAgent is a powerful agentic framework designed for autonomous web navigation and operation. It now includes three complementary parts: a full-featured Agentic Framework for state-of-the-art performance, a streamlined Single-Model Mode for ease of use and quick deployment, and a newly released Multi-Agent RL Training Framework for agent training research and experimentation.
π₯π₯π₯ [2026/09/14] We have released the OpAgent Web Benchmark β a 665-case benchmark on mock websites (READ 573 + OPERATION 92), together with browser_use evaluation results across 8 models. β‘οΈ Go to the OpAgent Web Benchmark directory for details β¬ οΈ
π₯π₯π₯ [2026/04/24] We have open-sourced our multi-agent RL training framework under opagent_training/, covering training code, environment preparation helpers, and analysis/evaluation utilities.
β‘οΈ Go to the Multi-Agent RL Training Guide For Details β¬
οΈ
π₯π₯π₯ [2026/03/17] We have released the demo on HuggingFace and ModelScope. We invite everyone to try it out and share your feedback!
π₯π₯π₯ [2026/03/17] We have released the INT4-quantized version of the OpAgent-32B model, enabling efficient deployment on consumer-grade hardware with 24GB of VRAM. β‘οΈ Go to the Single-Model Mode Usage Guide For Details β¬ οΈ
πππ [2026/02/14] We have released our technical report. Please refer to OpAgent Technical Report for details.
π₯π₯π₯ [2026/01/22] We are pleased to announce that Opagent achieves a remarkable 71.6% resolve rate on the Webarena leaderboard.
This repository provides the code and models for OpAgent, an operator agent for web navigation. We offer three complementary parts:
OpAgent: Single-Model Mode (opagent_single_model/ directory)
OpAgent: The Full Agentic Framework (opagent/ directory)
OpAgent: Multi-Agent RL Training Framework (opagent_training/ directory)
Agent-R1 training codebase, environment preparation helpers, and analysis/evaluation utilities.opagent_training/README.md.We employ an innovative Online Agentic Reinforcement Learning (RL) pipeline to significantly improve the capability of a single VLM. Our RL-enhanced model (RL-HybridReward-Zero) achieves a 38.1% success rate (@Pass5) on WebArena, outperforming other monolithic baselines and demonstrating a 10.7% absolute improvement over the original model.

Our full agentic framework, OpAgent, achieves a state-of-the-art (SOTA) 71.6% resolve rate on the WebArena benchmark (formerly OAgent on the WebArena leaderboard), securing the #1 position on the leaderboard on Jan. 2026.

Depending on which part you'd like to use, please follow the instructions below.
opagent_single_model/)This mode provides a ready-to-use, interactive web agent powered by a single model. It's the quickest way to see OpAgent in action.

For detailed installation and usage instructions, please refer to the README in the opagent_single_model directory:
β‘οΈ Go to Single-Model Mode Usage Guide β¬ οΈ
A quick preview of how to get started:
cd opagent_single_model
pip install -r requirements.txt
python main.py
opagent/)This mode utilizes a multi-agent architecture (Planner, Grounder, etc.) to achieve the highest performance.
The core logic is implemented in the ./opagent/ directory, with evaluation scripts located in ./demo/local_agent_eval.py. This setup is primarily designed for benchmark evaluation and research.
To run the evaluation:
# Detailed setup and execution instructions are work-in-progress.
# Please refer to the code in the 'opagent' and 'demo' directories for now.
# We welcome community contributions to improve the documentation!
(Learn more about the Agentic Framework's architecture below)β
opagent_training/)This module contains our newly open-sourced multi-agent RL training framework for web-style agents.
For detailed setup, environment preparation, and training workflow, please refer to the README in the opagent_training directory:
β‘οΈ Go to Multi-Agent RL Training Guide β¬ οΈ
This section details the architecture of our high-performance, multi-agent framework.
This document describes the structure of the demo WebAgent framework implemented in the ./opagent/local_agent_eval.py script. This framework aims to execute and evaluate automated tasks in real Web environments (such as the WebArena Shopping environment) via local/remote model calls.
This Agent adopts a modular Planner-Grounder-Reflector-Summary architecture. The entire system consists of a task scheduler, multi-threaded Workers, browser environment management, and core Agent logic.
The execution flow of the Agent is a closed-loop system, mainly containing the following steps:
is_task_done).tips).instruction) and action type (action_type).coords) or operation parameters.LocalWebAgent ClassThe main body of the Agent, responsible for maintaining task status, calling various model modules, and executing the main loop.
steps (history steps), marked_notes (collected info), last_screenshot.call_reflector: Calls the reasoning model to judge status.call_planner: Calls the reasoning model to generate plans.call_grounder: Calls the visual model (usually an SFT model) to get precise coordinates.call_summary: Generates the final answer.get_domain_specific_tips dynamically loads operation guides for different sites like Shopping/Admin/Map based on the current URL.LocalModelCaller ClassA unified model call interface encapsulating requests to different backend services:
BrowserActor & Distributed ExecutionThe framework defines four core Prompt templates guiding different Agent roles:
REFLECTION_PROMPT: Emphasizes "based on observed facts", responsible for verifying task success criteria, detecting infinite loops, and collecting structured data.PLANNER_PROMPT: Responsible for generating atomic operation instructions. Includes detailed action definitions (scroll, click, type, etc.) and core principles (priority search, table pagination checks, etc.).GROUNDER_PROMPT: Concise visual instructions requiring the model to output <tool_call> or coordinates.SUMMARY_PROMPT: Responsible for formatting the final answer, handling sorting, counting, and specific format requirements.select_option (when Playwright standard selection fails).Alongside inference and evaluation, this repository now includes a dedicated sub-project for multi-agent RL training under opagent_training/.
For detailed setup and usage guidance, see opagent_training/README.md.
We release a 665-case web agent benchmark on self-contained mock websites (READ 573 + OPERATION 92, 71 mock sites), together with browser_use evaluation results for 8 models under a unified protocol (READ = LLM semantic judge, OPERATION = deterministic frontend state assertions).
| Rank | Model | READ (573) | OP (92) | Overall (665) |
|---|---|---|---|---|
| 1 | Kimi-K3 | 79.8% | 82.6% | 80.2% |
| 2 | MiniMax-M3 | 79.1% | 80.4% | 79.2% |
| 3 | Qwen3.5-397B-A17B | 77.1% | 71.7% | 76.4% |
| 4 | Qwen3.5-27B | 76.6% | 73.9% | 76.2% |
| 5 | GLM-5.2 (text-only DOM) | 74.2% | 71.7% | 73.8% |
| 6 | Kimi-K2.5 | 71.2% | 71.7% | 71.3% |
| 7 | Qwen3-VL-235B | 69.5% | 65.2% | 68.9% |
| 8 | DeepSeek-V4-Pro | 67.2% | 63.0% | 66.6% |
For the dataset, per-model results and judging protocol, see opagent_web_benchmark/.
If you use OpAgent in your research or project, please cite it as follows:
@article{guo2026opagent,
title={OpAgent: Operator Agent for Web Navigation},
author={Guo, Yuyu and Yang, Wenjie and Yang, Siyuan and Liu, Ziyang and Chen, Cheng and Wei, Yuan and Hu, Yun and Huang, Yang and Hao, Guoliang and Yuan, Dongsheng and others},
journal={arXiv preprint arXiv:2602.13559},
year={2026}
}
412 followers Β· starred Feb 2026
Python
87.5%
Shell
10.8%
The OpAgent framework is used to automate web browser operations, achieving State-of-the-Art (SOTA) performance on the WebArena benchmark.
Python
235
44 commits
updated Sep 15, 2026

OpAgent is a powerful agentic framework designed for autonomous web navigation and operation. It now includes three complementary parts: a full-featured Agentic Framework for state-of-the-art performance, a streamlined Single-Model Mode for ease of use and quick deployment, and a newly released Multi-Agent RL Training Framework for agent training research and experimentation.
π₯π₯π₯ [2026/09/14] We have released the OpAgent Web Benchmark β a 665-case benchmark on mock websites (READ 573 + OPERATION 92), together with browser_use evaluation results across 8 models. β‘οΈ Go to the OpAgent Web Benchmark directory for details β¬ οΈ
π₯π₯π₯ [2026/04/24] We have open-sourced our multi-agent RL training framework under opagent_training/, covering training code, environment preparation helpers, and analysis/evaluation utilities.
β‘οΈ Go to the Multi-Agent RL Training Guide For Details β¬
οΈ
π₯π₯π₯ [2026/03/17] We have released the demo on HuggingFace and ModelScope. We invite everyone to try it out and share your feedback!
π₯π₯π₯ [2026/03/17] We have released the INT4-quantized version of the OpAgent-32B model, enabling efficient deployment on consumer-grade hardware with 24GB of VRAM. β‘οΈ Go to the Single-Model Mode Usage Guide For Details β¬ οΈ
πππ [2026/02/14] We have released our technical report. Please refer to OpAgent Technical Report for details.
π₯π₯π₯ [2026/01/22] We are pleased to announce that Opagent achieves a remarkable 71.6% resolve rate on the Webarena leaderboard.
This repository provides the code and models for OpAgent, an operator agent for web navigation. We offer three complementary parts:
OpAgent: Single-Model Mode (opagent_single_model/ directory)
OpAgent: The Full Agentic Framework (opagent/ directory)
OpAgent: Multi-Agent RL Training Framework (opagent_training/ directory)
Agent-R1 training codebase, environment preparation helpers, and analysis/evaluation utilities.opagent_training/README.md.We employ an innovative Online Agentic Reinforcement Learning (RL) pipeline to significantly improve the capability of a single VLM. Our RL-enhanced model (RL-HybridReward-Zero) achieves a 38.1% success rate (@Pass5) on WebArena, outperforming other monolithic baselines and demonstrating a 10.7% absolute improvement over the original model.

Our full agentic framework, OpAgent, achieves a state-of-the-art (SOTA) 71.6% resolve rate on the WebArena benchmark (formerly OAgent on the WebArena leaderboard), securing the #1 position on the leaderboard on Jan. 2026.

Depending on which part you'd like to use, please follow the instructions below.
opagent_single_model/)This mode provides a ready-to-use, interactive web agent powered by a single model. It's the quickest way to see OpAgent in action.

For detailed installation and usage instructions, please refer to the README in the opagent_single_model directory:
β‘οΈ Go to Single-Model Mode Usage Guide β¬ οΈ
A quick preview of how to get started:
cd opagent_single_model
pip install -r requirements.txt
python main.py
opagent/)This mode utilizes a multi-agent architecture (Planner, Grounder, etc.) to achieve the highest performance.
The core logic is implemented in the ./opagent/ directory, with evaluation scripts located in ./demo/local_agent_eval.py. This setup is primarily designed for benchmark evaluation and research.
To run the evaluation:
# Detailed setup and execution instructions are work-in-progress.
# Please refer to the code in the 'opagent' and 'demo' directories for now.
# We welcome community contributions to improve the documentation!
(Learn more about the Agentic Framework's architecture below)β
opagent_training/)This module contains our newly open-sourced multi-agent RL training framework for web-style agents.
For detailed setup, environment preparation, and training workflow, please refer to the README in the opagent_training directory:
β‘οΈ Go to Multi-Agent RL Training Guide β¬ οΈ
This section details the architecture of our high-performance, multi-agent framework.
This document describes the structure of the demo WebAgent framework implemented in the ./opagent/local_agent_eval.py script. This framework aims to execute and evaluate automated tasks in real Web environments (such as the WebArena Shopping environment) via local/remote model calls.
This Agent adopts a modular Planner-Grounder-Reflector-Summary architecture. The entire system consists of a task scheduler, multi-threaded Workers, browser environment management, and core Agent logic.
The execution flow of the Agent is a closed-loop system, mainly containing the following steps:
is_task_done).tips).instruction) and action type (action_type).coords) or operation parameters.LocalWebAgent ClassThe main body of the Agent, responsible for maintaining task status, calling various model modules, and executing the main loop.
steps (history steps), marked_notes (collected info), last_screenshot.call_reflector: Calls the reasoning model to judge status.call_planner: Calls the reasoning model to generate plans.call_grounder: Calls the visual model (usually an SFT model) to get precise coordinates.call_summary: Generates the final answer.get_domain_specific_tips dynamically loads operation guides for different sites like Shopping/Admin/Map based on the current URL.LocalModelCaller ClassA unified model call interface encapsulating requests to different backend services:
BrowserActor & Distributed ExecutionThe framework defines four core Prompt templates guiding different Agent roles:
REFLECTION_PROMPT: Emphasizes "based on observed facts", responsible for verifying task success criteria, detecting infinite loops, and collecting structured data.PLANNER_PROMPT: Responsible for generating atomic operation instructions. Includes detailed action definitions (scroll, click, type, etc.) and core principles (priority search, table pagination checks, etc.).GROUNDER_PROMPT: Concise visual instructions requiring the model to output <tool_call> or coordinates.SUMMARY_PROMPT: Responsible for formatting the final answer, handling sorting, counting, and specific format requirements.select_option (when Playwright standard selection fails).Alongside inference and evaluation, this repository now includes a dedicated sub-project for multi-agent RL training under opagent_training/.
For detailed setup and usage guidance, see opagent_training/README.md.
We release a 665-case web agent benchmark on self-contained mock websites (READ 573 + OPERATION 92, 71 mock sites), together with browser_use evaluation results for 8 models under a unified protocol (READ = LLM semantic judge, OPERATION = deterministic frontend state assertions).
| Rank | Model | READ (573) | OP (92) | Overall (665) |
|---|---|---|---|---|
| 1 | Kimi-K3 | 79.8% | 82.6% | 80.2% |
| 2 | MiniMax-M3 | 79.1% | 80.4% | 79.2% |
| 3 | Qwen3.5-397B-A17B | 77.1% | 71.7% | 76.4% |
| 4 | Qwen3.5-27B | 76.6% | 73.9% | 76.2% |
| 5 | GLM-5.2 (text-only DOM) | 74.2% | 71.7% | 73.8% |
| 6 | Kimi-K2.5 | 71.2% | 71.7% | 71.3% |
| 7 | Qwen3-VL-235B | 69.5% | 65.2% | 68.9% |
| 8 | DeepSeek-V4-Pro | 67.2% | 63.0% | 66.6% |
For the dataset, per-model results and judging protocol, see opagent_web_benchmark/.
If you use OpAgent in your research or project, please cite it as follows:
@article{guo2026opagent,
title={OpAgent: Operator Agent for Web Navigation},
author={Guo, Yuyu and Yang, Wenjie and Yang, Siyuan and Liu, Ziyang and Chen, Cheng and Wei, Yuan and Hu, Yun and Huang, Yang and Hao, Guoliang and Yuan, Dongsheng and others},
journal={arXiv preprint arXiv:2602.13559},
year={2026}
}
412 followers Β· starred Feb 2026
Python
87.5%
Shell
10.8%