A curated list of of awesome UI agents resources, encompassing Web, App, OS, and beyond (continually updated)
319
29 commits
updated Jun 17, 2026
This is a collection of research papers for UI Agent, which includes models, tools, and datasets. And the repository will be continuously updated to track the frontier of UI Agent or related fields.
Welcome to follow and star!
UI Agent aims to build a generalist agent that can interact with various user interfaces (UIs) in different environments, such as mobile apps, web pages, and PC applications. The agent can understand the UIs through vision-language models and interact with them to complete tasks. The agent can be applied to various scenarios, such as mobile device operation, web browsing, and game playing. The agent can be trained in a simulated environment or with real-world data. The agent can be evaluated in terms of task completion rate, efficiency, and generalization ability.
The research on UI Agent is still in its early stage, and there are many challenges to be addressed, such as the scalability of the agent, the robustness of the agent, and the interpretability of the agent. The research on UI Agent is interdisciplinary, involving computer vision, natural language processing, reinforcement learning, human-computer interaction, and software engineering. The research on UI Agent has the potential to revolutionize the way we interact with computers and improve the efficiency and usability of computer systems.
format:
- [title](paper link) [links]
- author1, author2, and author3...
- year
- publisher
- key
- code
- experiment environment
PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX Evaluation
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as Reasoning
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
Go-Browse: Training Web Agents with Structured Exploration
WALT: Web Agents that Learn Tools
Safe and Scalable Web Agent Learning via Recreated Websites
CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
Programming with Pixels: Towards Generalist Software Engineering Agents
E-commerce UI/UX Optimization via Generative AI, MDD, and Multi-Agent Reinforcement Learning
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
Enhance Mobile Agents Thinking Process Via Iterative Preference Learning
UI-Venus Technical Report: Building High-performance UI Agents with RFT
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery for Foundation Model Internet Agents
AppVLM: A Lightweight Vision Language Model for Online App Control
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
Lightweight Neural App Control
Enhancing Software Agents with Monte Carlo Tree Search and Hindsight Feedback
On the Effects of Data Scale on UI Control Agents
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
Cradle: Empowering Foundation Agents Towards General Computer Control
Lightweight Neural App Control
SeeAct GPT-4V(ision) is a Generalist Web Agent, if Grounded
MMAC-Copilot: Multi-modal Agent Collaboration Operating System Copilot
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents
Autowebglm: Bootstrap and reinforce a large language model-based web navigating agent
Dual-view visual contextualization for web navigation
Agent-e: From autonomous web navigation to foundational design principles in agentic systems
Tree search for language model agents
Agent S: an open agentic framework that uses computers like a human
Apple Intelligence Foundation Language Models
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
LATS: Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
ScreenAgent: A Vision Language Model-driven Computer Control Agent
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
UFO: A UI-Focused Agent for Windows OS Interaction
Octopus v2: On-device language model for super agent
Openagents: An open platform for language agents in the wild
LASER: LLM Agent with State-Space Exploration for Web Navigation
AppAgent: Multimodal Agents as Smartphone Users
CogAgent: A Visual Language Model for GUI Agents
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
You Only Look at Screens: Multimodal Chain-of-Action Agents
LASER: LLM Agent with State-Space Exploration for Web Navigation
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Augmenting Autotelic Agents with Large Language Models
Language Models can Solve Computer Tasks
Opera Browser Operator: AI-based Agentic Browsing
OmniParser: Screen Parsing tool for Pure Vision Based GUI Agent
-Make Websites Accessible for Agents - Li Zhang and Shihe Wang and Xianqing Jia and Zhihan Zheng and Yunhe Yan and Longxi Gao and Yuanchun Li and Mengwei Xu - Key: websites, Agents - 2024 - code
ToolGen: Unified Tool Retrieval and Calling via Generation
LEGENT: An Open Platform for Embodied Agentb Agents on Large Language Models
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Automation Task Evaluation
WebArena: A Realistic Web Environment for Building Autonomous Agents
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
AndroidEnv: A Reinforcement Learning Platform for Android
TWZRD Agent Intel - Trust scoring MCP for AI agents on Solana. Verify agent wallet identity before x402 micropayments. Free: {"mcpServers":{"twzrd-agent-intel":{"url":"https://intel.twzrd.xyz/mcp"}}}
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks
WebCanvas: Benchmarking Web Agents in Online Environments
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions
Multi-Turn Mind2Web: On the Multi-turn Instruction Following for Conversational Web Agents
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Android in the Wild: A Large-Scale Dataset for Android Device Control
Mind2Web: Towards a Generalist Agent for the Web
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Rico: A Mobile App Dataset for Building Data-Driven Design Applications
Our purpose is to make this repo even better. If you are interested in contributing, please refer to HERE for instructions in contribution.
This repository is released under the Apache 2.0 license.
A curated list of of awesome UI agents resources, encompassing Web, App, OS, and beyond (continually updated)
319
29 commits
updated Jun 17, 2026
This is a collection of research papers for UI Agent, which includes models, tools, and datasets. And the repository will be continuously updated to track the frontier of UI Agent or related fields.
Welcome to follow and star!
UI Agent aims to build a generalist agent that can interact with various user interfaces (UIs) in different environments, such as mobile apps, web pages, and PC applications. The agent can understand the UIs through vision-language models and interact with them to complete tasks. The agent can be applied to various scenarios, such as mobile device operation, web browsing, and game playing. The agent can be trained in a simulated environment or with real-world data. The agent can be evaluated in terms of task completion rate, efficiency, and generalization ability.
The research on UI Agent is still in its early stage, and there are many challenges to be addressed, such as the scalability of the agent, the robustness of the agent, and the interpretability of the agent. The research on UI Agent is interdisciplinary, involving computer vision, natural language processing, reinforcement learning, human-computer interaction, and software engineering. The research on UI Agent has the potential to revolutionize the way we interact with computers and improve the efficiency and usability of computer systems.
format:
- [title](paper link) [links]
- author1, author2, and author3...
- year
- publisher
- key
- code
- experiment environment
PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX Evaluation
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as Reasoning
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
Go-Browse: Training Web Agents with Structured Exploration
WALT: Web Agents that Learn Tools
Safe and Scalable Web Agent Learning via Recreated Websites
CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
Programming with Pixels: Towards Generalist Software Engineering Agents
E-commerce UI/UX Optimization via Generative AI, MDD, and Multi-Agent Reinforcement Learning
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
Enhance Mobile Agents Thinking Process Via Iterative Preference Learning
UI-Venus Technical Report: Building High-performance UI Agents with RFT
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery for Foundation Model Internet Agents
AppVLM: A Lightweight Vision Language Model for Online App Control
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
Lightweight Neural App Control
Enhancing Software Agents with Monte Carlo Tree Search and Hindsight Feedback
On the Effects of Data Scale on UI Control Agents
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
Cradle: Empowering Foundation Agents Towards General Computer Control
Lightweight Neural App Control
SeeAct GPT-4V(ision) is a Generalist Web Agent, if Grounded
MMAC-Copilot: Multi-modal Agent Collaboration Operating System Copilot
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents
Autowebglm: Bootstrap and reinforce a large language model-based web navigating agent
Dual-view visual contextualization for web navigation
Agent-e: From autonomous web navigation to foundational design principles in agentic systems
Tree search for language model agents
Agent S: an open agentic framework that uses computers like a human
Apple Intelligence Foundation Language Models
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
LATS: Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
ScreenAgent: A Vision Language Model-driven Computer Control Agent
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
UFO: A UI-Focused Agent for Windows OS Interaction
Octopus v2: On-device language model for super agent
Openagents: An open platform for language agents in the wild
LASER: LLM Agent with State-Space Exploration for Web Navigation
AppAgent: Multimodal Agents as Smartphone Users
CogAgent: A Visual Language Model for GUI Agents
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
You Only Look at Screens: Multimodal Chain-of-Action Agents
LASER: LLM Agent with State-Space Exploration for Web Navigation
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Augmenting Autotelic Agents with Large Language Models
Language Models can Solve Computer Tasks
Opera Browser Operator: AI-based Agentic Browsing
OmniParser: Screen Parsing tool for Pure Vision Based GUI Agent
-Make Websites Accessible for Agents - Li Zhang and Shihe Wang and Xianqing Jia and Zhihan Zheng and Yunhe Yan and Longxi Gao and Yuanchun Li and Mengwei Xu - Key: websites, Agents - 2024 - code
ToolGen: Unified Tool Retrieval and Calling via Generation
LEGENT: An Open Platform for Embodied Agentb Agents on Large Language Models
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Automation Task Evaluation
WebArena: A Realistic Web Environment for Building Autonomous Agents
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
AndroidEnv: A Reinforcement Learning Platform for Android
TWZRD Agent Intel - Trust scoring MCP for AI agents on Solana. Verify agent wallet identity before x402 micropayments. Free: {"mcpServers":{"twzrd-agent-intel":{"url":"https://intel.twzrd.xyz/mcp"}}}
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks
WebCanvas: Benchmarking Web Agents in Online Environments
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions
Multi-Turn Mind2Web: On the Multi-turn Instruction Following for Conversational Web Agents
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Android in the Wild: A Large-Scale Dataset for Android Device Control
Mind2Web: Towards a Generalist Agent for the Web
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Rico: A Mobile App Dataset for Building Data-Driven Design Applications
Our purpose is to make this repo even better. If you are interested in contributing, please refer to HERE for instructions in contribution.
This repository is released under the Apache 2.0 license.