A curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.
See the codeAn AI Agent for Computer Use is an autonomous program that can reason about tasks, plan sequences of actions, and act within the domain of a computer or mobile device in the form of clicks, keystrokes, other computer events, command-line operations and internal/external API calls. These agents combine perception, decision-making, and control capabilities to interact with digital interfaces and accomplish user-specified goals independently.
A curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.
GUI Agents: A Survey (Dec. 2024)
Large Language Model-Brained GUI Agents: A Survey (Nov. 2024)
GUI Agents with Foundation Models: A Comprehensive Survey (Nov. 2024)
Reinforcement Learning for Long-Horizon Interactive LLM Agents (Feb. 2025)
Large Action Models: From Inception to Implementation (Dec. 2024)
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation (Dec. 2024)
SpiritSight Agent: Advanced GUI Agent with One Look (Dec. 2024)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs (Dec. 2024)
Simulate Before Act: Model-Based Planning for Web Agents (Dec. 2024)
Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery For Foundation Model Internet Agents (Dec. 2024)
Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (Dec. 2024)
Digi-Q: Transforming VLMs to Device-Control Agents via Value-Based Offline RL (Dec. 2024)
Magentic-One (Nov. 2024)
Agent Workflow Memory (Sep. 2024)
The Impact of Element Ordering on LM Agent Performance (Sep. 2024)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents (Aug. 2024)
OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models (Aug. 2024)
Agent-e: From autonomous web navigation to foundational design principles in agentic systems (Jul. 2024)
Apple Intelligence Foundation Language Models (Jul. 2024)
Tree search for language model agents (Jul. 2024)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning (Jun. 2024)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration (Jun. 2024)
Octopus Series: On-device Language Models for Computer Control (Apr. 2024)
AutoWebGLM: Bootstrap and reinforce a large language model-based web navigating agent (Apr. 2024)
Cradle: Empowering Foundation Agents towards General Computer Control (Mar. 2024)
Android in the Zoo: Chain-of-Action-Thought for GUI Agents (Mar. 2024)
ScreenAgent: A Computer Control Agent Driven by Visual Language Large Model (Feb. 2024)
OS-Copilot: Towards Generalist Computer Agents with Self-Improvement (Feb. 2024)
UFO: A UI-Focused Agent for Windows OS Interaction (Feb. 2024)
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation (Feb. 2024)
Intention-inInteraction (IN3): Tell Me More! (Feb. 2024)
Dual-view visual contextualization for web navigation (Feb. 2024)
ScreenAI: A Vision-Language Model for UI and Infographics Understanding (Feb. 2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded (Jan. 2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception (Jan. 2024)
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models (Jan. 2024)
CogAgent: A Visual Language Model for GUI Agents (Dec. 2023)
AppAgent: Multimodal Agents as Smartphone Users (Dec. 2023)
LASER: LLM Agent with State-Space Exploration for Web Navigation (Sep. 2023)
AndroidEnv: A Reinforcement Learning Platform for Android (May 2021)
OmniParser for Pure Vision Based GUI Agent (Aug. 2024)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs (Apr. 2024)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents (Jan. 2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms (Oct. 2024)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents (Oct. 2024)
OS-ATLAS: Foundation Action Model for Generalist GUI Agents (Oct. 2024)
UI-Pro: A Hidden Recipe for Building Vision-Language Models for GUI Grounding (Dec. 2024)
Grounding Multimodal Large Language Model in GUI World (Dec. 2024)
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents (Feb. 2025)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis (Dec. 2024)
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials (Dec. 2024)
ICAL: Continual Learning of Multimodal Agents by Transforming Trajectories into Actionable Insights (Jun. 2024)
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale (Sep. 2024)
UiPad: UI Parsing and Accessibility Dataset (Sep. 2024)
Multi-Turn Mind2Web: On the Multi-turn Instruction Following (Feb. 2024)
CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation (Aug. 2024)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks (Jul. 2024)
Mind2Web: Towards a Generalist Agent for the Web (Jun. 2023)
Android in the Wild: A Large-Scale Dataset for Android Device Control (Jul. 2023)
WebShop: Towards Scalable Real-World Web Interaction (Jul. 2022)
Rico: A Mobile App Dataset for Building Data-Driven Design Applications (Oct. 2017)
A3: Android Agent Arena for Mobile GUI Agents (Jan. 2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (Apr. 2024)
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents (May. 2024)
Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows? (Jul. 2024)
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents (Jul. 2024)
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (Jun. 2024)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents (Jun. 2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks (Jan. 2024)
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale (Sep. 2024)
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction (May. 2023)
Attacking Vision-Language Computer Agents via Pop-ups (Nov. 2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage (Sep. 2024)
GuardAgent: Safeguard LLM Agent by a Guard Agent via Knowledge-Enabled Reasoning (Jun. 2024)
Open Source Computer Use by E2B
Self-Operating Computer (Nov. 2023)
We welcome and encourage contributions from the community! Here's how you can help:
To contribute:
For an example of how to format your contribution, please refer to this PR.
Thank you for helping spread knowledge about AI agents for computer use!
A curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.
See the codeAn AI Agent for Computer Use is an autonomous program that can reason about tasks, plan sequences of actions, and act within the domain of a computer or mobile device in the form of clicks, keystrokes, other computer events, command-line operations and internal/external API calls. These agents combine perception, decision-making, and control capabilities to interact with digital interfaces and accomplish user-specified goals independently.
A curated list of resources about AI agents for Computer Use, including research papers, projects, frameworks, and tools.
GUI Agents: A Survey (Dec. 2024)
Large Language Model-Brained GUI Agents: A Survey (Nov. 2024)
GUI Agents with Foundation Models: A Comprehensive Survey (Nov. 2024)
Reinforcement Learning for Long-Horizon Interactive LLM Agents (Feb. 2025)
Large Action Models: From Inception to Implementation (Dec. 2024)
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation (Dec. 2024)
SpiritSight Agent: Advanced GUI Agent with One Look (Dec. 2024)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs (Dec. 2024)
Simulate Before Act: Model-Based Planning for Web Agents (Dec. 2024)
Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery For Foundation Model Internet Agents (Dec. 2024)
Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (Dec. 2024)
Digi-Q: Transforming VLMs to Device-Control Agents via Value-Based Offline RL (Dec. 2024)
Magentic-One (Nov. 2024)
Agent Workflow Memory (Sep. 2024)
The Impact of Element Ordering on LM Agent Performance (Sep. 2024)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents (Aug. 2024)
OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models (Aug. 2024)
Agent-e: From autonomous web navigation to foundational design principles in agentic systems (Jul. 2024)
Apple Intelligence Foundation Language Models (Jul. 2024)
Tree search for language model agents (Jul. 2024)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning (Jun. 2024)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration (Jun. 2024)
Octopus Series: On-device Language Models for Computer Control (Apr. 2024)
AutoWebGLM: Bootstrap and reinforce a large language model-based web navigating agent (Apr. 2024)
Cradle: Empowering Foundation Agents towards General Computer Control (Mar. 2024)
Android in the Zoo: Chain-of-Action-Thought for GUI Agents (Mar. 2024)
ScreenAgent: A Computer Control Agent Driven by Visual Language Large Model (Feb. 2024)
OS-Copilot: Towards Generalist Computer Agents with Self-Improvement (Feb. 2024)
UFO: A UI-Focused Agent for Windows OS Interaction (Feb. 2024)
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation (Feb. 2024)
Intention-inInteraction (IN3): Tell Me More! (Feb. 2024)
Dual-view visual contextualization for web navigation (Feb. 2024)
ScreenAI: A Vision-Language Model for UI and Infographics Understanding (Feb. 2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded (Jan. 2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception (Jan. 2024)
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models (Jan. 2024)
CogAgent: A Visual Language Model for GUI Agents (Dec. 2023)
AppAgent: Multimodal Agents as Smartphone Users (Dec. 2023)
LASER: LLM Agent with State-Space Exploration for Web Navigation (Sep. 2023)
AndroidEnv: A Reinforcement Learning Platform for Android (May 2021)
OmniParser for Pure Vision Based GUI Agent (Aug. 2024)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs (Apr. 2024)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents (Jan. 2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms (Oct. 2024)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents (Oct. 2024)
OS-ATLAS: Foundation Action Model for Generalist GUI Agents (Oct. 2024)
UI-Pro: A Hidden Recipe for Building Vision-Language Models for GUI Grounding (Dec. 2024)
Grounding Multimodal Large Language Model in GUI World (Dec. 2024)
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents (Feb. 2025)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis (Dec. 2024)
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials (Dec. 2024)
ICAL: Continual Learning of Multimodal Agents by Transforming Trajectories into Actionable Insights (Jun. 2024)
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale (Sep. 2024)
UiPad: UI Parsing and Accessibility Dataset (Sep. 2024)
Multi-Turn Mind2Web: On the Multi-turn Instruction Following (Feb. 2024)
CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation (Aug. 2024)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks (Jul. 2024)
Mind2Web: Towards a Generalist Agent for the Web (Jun. 2023)
Android in the Wild: A Large-Scale Dataset for Android Device Control (Jul. 2023)
WebShop: Towards Scalable Real-World Web Interaction (Jul. 2022)
Rico: A Mobile App Dataset for Building Data-Driven Design Applications (Oct. 2017)
A3: Android Agent Arena for Mobile GUI Agents (Jan. 2025)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (Apr. 2024)
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents (May. 2024)
Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows? (Jul. 2024)
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents (Jul. 2024)
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (Jun. 2024)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents (Jun. 2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks (Jan. 2024)
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale (Sep. 2024)
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction (May. 2023)
Attacking Vision-Language Computer Agents via Pop-ups (Nov. 2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage (Sep. 2024)
GuardAgent: Safeguard LLM Agent by a Guard Agent via Knowledge-Enabled Reasoning (Jun. 2024)
Open Source Computer Use by E2B
Self-Operating Computer (Nov. 2023)
We welcome and encourage contributions from the community! Here's how you can help:
To contribute:
For an example of how to format your contribution, please refer to this PR.
Thank you for helping spread knowledge about AI agents for computer use!