Survey on Multimodal Embodied Agents: From Computer-Use to Robot-Use
See the codeYanzhe Chen ·
Ziyi Yang ·
Jifeng Zhu ·
Qiming Huang ·
Ruihe An ·
Peiyao Xu
Hesen Yang ·
Runda Liu ·
Chang Gong ·
Zhijun Cao ·
Zechen Bai ·
Wenzheng Zeng
Yiqi Lin ·
Guoqiang Liang ·
Kevin Yuchen Ma ·
Kevin Qinghong Lin ·
Mike Zheng Shou†
This repository accompanies our survey on multimodal embodied agents (MMEAs): systems that couple multimodal reasoning with physical action and learn from the resulting feedback. It asks one question: what changes when a multimodal agent moves from computer-use to robot-use?
PAPAV: five capabilities that recur in every task loop, whether the agent operates a computer (top) or a robot (bottom).
[!IMPORTANT] This area is growing quickly, and so is this list. If we missed a paper, or you have just published one of your own, please open an issue or send a pull request. We are always glad to add new work!
A chronological view of representative robotic systems, multimodal embodied agents, and multimodal agents.
2026.09 Check out our survey paper on Multimodal Embodied Agents!2026.09 First release: 310 papers across MMEAs, multimodal agents, robotic systems, and benchmarks.Prior surveys and reviews adjacent to our scope.
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents
Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
Foundation-model-driven agents that couple multimodal perception and reasoning to embodied sensing and physical action in a closed loop — VLAs, LLM/VLM planners and critics for robots, language-conditioned robot world models, and embodied memory.
Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA
CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
eMEM: A Hybrid Spatio-Temporal Memory System For Embodied Agents
G³VLA: Geometric inductive bias for Vision-Language-Action Models
Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
KITE: Decoupling Kinematics and Interaction for Zero-Shot Cross-Embodiment Manipulation
What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory
What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents
Perturbation-Based Uncertainty for Failure Detection in Vision-Language-Action Models
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models
World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Dynamic Execution Commitment of Vision-Language-Action Models
EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
Robot Planning and Situation Handling with Active Perception
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
Using large language models for embodied planning introduces systematic safety risks
Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World
SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
Kinodynamic Task and Motion Planning using VLM-guided and Interleaved Sampling
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
SAFE: Multitask Failure Detection for Vision-Language-Action Models
A Unified Framework for Real-Time Failure Handling in Robotics Using Vision-Language Models, Reactive Planner and Behavior Trees
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
FAST: Efficient Action Tokenization for Vision-Language-Action Models
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
π₀: A Vision-Language-Action Flow Model for General Robot Control
CaStL: Constraints as Specifications through LLM Translation for Long-Horizon Task and Motion Planning
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
RoboDreamer: Learning Compositional World Models for Robot Imagination
Explore until Confident: Efficient Exploration for Embodied Question Answering
Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
Plug in the Safety Chip: Enforcing Constraints for LLM-driven Robot Agents
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
Liv: Language-image representations and rewards for robotic control
LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
Learning Universal Policies via Text-Guided Video Generation
Code as Policies: Language Model Programs for Embodied Control
ProgPrompt: Generating Situated Robot Task Plans Using Large Language Models
Inner Monologue: Embodied Reasoning through Planning with Language Models
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
Agents built on multi-modal foundation models that perceive and act in digital environments — screens, browsers, documents, and APIs — without a physical body.
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
Confident and Wrong: Silent Semantic Failures in Coding Agents
Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
Code2world: A gui world model via renderable code generation
Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents
Guitester: Enabling gui agents for exploratory defect discovery
WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands
Active perception agent for omnimodal audio-video understanding
WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes
Scaling Synthetic Task Generation for Agents via Exploration
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
Backtrackagent: Enhancing gui agent with error detection and backtracking mechanism
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
Webevolver: Enhancing web agent self-improvement with co-evolving world model
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Aguvis: Unified pure vision agents for autonomous gui interaction
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
Pandora: Towards general world model with natural language actions and video states
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Videoagent: A memory-augmented multimodal agent for video understanding
Os-copilot: Towards generalist computer agents with self-improvement
Position: LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Timechat: A time-sensitive multimodal large language model for long video understanding
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Salmonn: Towards generic hearing abilities for large language models
ControlLLM: Augment Language Models with Tools by Searching on Graphs
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Avis: Autonomous visual information seeking with large language model agent
Kosmos-2: Grounding multimodal large language models to the world
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
Voyager: An Open-Ended Embodied Agent with Large Language Models
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Mm-react: Prompting chatgpt for multimodal reasoning and action
Reflexion: Language Agents with Verbal Reinforcement Learning
Mastering atari, go, chess and shogi by planning with a learned model
Robot learning, perception, and control in the physical world — vision-language-action models, manipulation, navigation, humanoids, and the data and simulators behind them.
ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies
ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation
Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies
Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents
TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison
ComFree-Sim: A GPU-Parallelized Analytical Contact Physics Engine for Scalable Contact-Rich Robotics Simulation and Control
Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
Demystifying Action Space Design for Robotic Manipulation Policies
APPLE: Toward General Active Perception via Reinforcement Learning
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
Can we detect failures without failure data? uncertainty-aware runtime failure detection for imitation learning policies
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
MIRAGE: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting
TD-MPC2: Scalable, Robust World Models for Continuous Control
RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation
Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization
Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks
ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM
Kimera: an open-source library for real-time metric-semantic localization and mapping
Composable action-conditioned predictors: Flexible off-policy learning for robot navigation
Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
Self-Supervised Visual Planning with Temporal Skip Connections.
Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics
Impedance Control: An Approach to Manipulation, Part I-Theory
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
EMBODIEDBENCH: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
WebArena: A Realistic Web Environment for Building Autonomous Agents
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
WorldModelBench: Judging Video Generation Models As World Models
SWE-BENCH: CAN LANGUAGE MODELS RESOLVE REAL-WORLD GITHUB ISSUES?
Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
RLBench: The Robot Learning Benchmark and Learning Environment
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
Evaluating Real-World Robot Manipulation Policies in Simulation
Contributions are very welcome — new papers, corrections, better categorization, or dead-link fixes.
main. Include the first public release
date and your section choice when adding a paper.Every correction, addition, and categorization improvement is appreciated. Contributors are recorded in the repository history and on GitHub's contributor graph.
We thank the researchers who make their papers, code, models, datasets, and project pages publicly available, as well as the maintainers of the related collections that help the community navigate this fast-moving field.
The official survey citation is not public yet. A verified BibTeX entry will be added after the manuscript is released. Until then, please link to this repository rather than using a provisional citation.
Released under CC0-1.0. The listed papers remain under their own licenses and copyright.
If you like the repository, please give us a star ⭐ — it is how we hear that it is useful.
Survey on Multimodal Embodied Agents: From Computer-Use to Robot-Use
See the codeYanzhe Chen ·
Ziyi Yang ·
Jifeng Zhu ·
Qiming Huang ·
Ruihe An ·
Peiyao Xu
Hesen Yang ·
Runda Liu ·
Chang Gong ·
Zhijun Cao ·
Zechen Bai ·
Wenzheng Zeng
Yiqi Lin ·
Guoqiang Liang ·
Kevin Yuchen Ma ·
Kevin Qinghong Lin ·
Mike Zheng Shou†
This repository accompanies our survey on multimodal embodied agents (MMEAs): systems that couple multimodal reasoning with physical action and learn from the resulting feedback. It asks one question: what changes when a multimodal agent moves from computer-use to robot-use?
PAPAV: five capabilities that recur in every task loop, whether the agent operates a computer (top) or a robot (bottom).
[!IMPORTANT] This area is growing quickly, and so is this list. If we missed a paper, or you have just published one of your own, please open an issue or send a pull request. We are always glad to add new work!
A chronological view of representative robotic systems, multimodal embodied agents, and multimodal agents.
2026.09 Check out our survey paper on Multimodal Embodied Agents!2026.09 First release: 310 papers across MMEAs, multimodal agents, robotic systems, and benchmarks.Prior surveys and reviews adjacent to our scope.
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents
Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
Foundation-model-driven agents that couple multimodal perception and reasoning to embodied sensing and physical action in a closed loop — VLAs, LLM/VLM planners and critics for robots, language-conditioned robot world models, and embodied memory.
Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA
CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
eMEM: A Hybrid Spatio-Temporal Memory System For Embodied Agents
G³VLA: Geometric inductive bias for Vision-Language-Action Models
Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
KITE: Decoupling Kinematics and Interaction for Zero-Shot Cross-Embodiment Manipulation
What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory
What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents
Perturbation-Based Uncertainty for Failure Detection in Vision-Language-Action Models
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models
World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Dynamic Execution Commitment of Vision-Language-Action Models
EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
Robot Planning and Situation Handling with Active Perception
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
Using large language models for embodied planning introduces systematic safety risks
Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World
SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
Kinodynamic Task and Motion Planning using VLM-guided and Interleaved Sampling
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
SAFE: Multitask Failure Detection for Vision-Language-Action Models
A Unified Framework for Real-Time Failure Handling in Robotics Using Vision-Language Models, Reactive Planner and Behavior Trees
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
FAST: Efficient Action Tokenization for Vision-Language-Action Models
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
π₀: A Vision-Language-Action Flow Model for General Robot Control
CaStL: Constraints as Specifications through LLM Translation for Long-Horizon Task and Motion Planning
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
RoboDreamer: Learning Compositional World Models for Robot Imagination
Explore until Confident: Efficient Exploration for Embodied Question Answering
Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
Plug in the Safety Chip: Enforcing Constraints for LLM-driven Robot Agents
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
Liv: Language-image representations and rewards for robotic control
LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
Learning Universal Policies via Text-Guided Video Generation
Code as Policies: Language Model Programs for Embodied Control
ProgPrompt: Generating Situated Robot Task Plans Using Large Language Models
Inner Monologue: Embodied Reasoning through Planning with Language Models
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
Agents built on multi-modal foundation models that perceive and act in digital environments — screens, browsers, documents, and APIs — without a physical body.
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
Confident and Wrong: Silent Semantic Failures in Coding Agents
Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
Code2world: A gui world model via renderable code generation
Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents
Guitester: Enabling gui agents for exploratory defect discovery
WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands
Active perception agent for omnimodal audio-video understanding
WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes
Scaling Synthetic Task Generation for Agents via Exploration
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
Backtrackagent: Enhancing gui agent with error detection and backtracking mechanism
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
Webevolver: Enhancing web agent self-improvement with co-evolving world model
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Aguvis: Unified pure vision agents for autonomous gui interaction
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
Pandora: Towards general world model with natural language actions and video states
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Videoagent: A memory-augmented multimodal agent for video understanding
Os-copilot: Towards generalist computer agents with self-improvement
Position: LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Timechat: A time-sensitive multimodal large language model for long video understanding
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Salmonn: Towards generic hearing abilities for large language models
ControlLLM: Augment Language Models with Tools by Searching on Graphs
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Avis: Autonomous visual information seeking with large language model agent
Kosmos-2: Grounding multimodal large language models to the world
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
Voyager: An Open-Ended Embodied Agent with Large Language Models
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Mm-react: Prompting chatgpt for multimodal reasoning and action
Reflexion: Language Agents with Verbal Reinforcement Learning
Mastering atari, go, chess and shogi by planning with a learned model
Robot learning, perception, and control in the physical world — vision-language-action models, manipulation, navigation, humanoids, and the data and simulators behind them.
ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies
ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation
Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies
Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents
TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison
ComFree-Sim: A GPU-Parallelized Analytical Contact Physics Engine for Scalable Contact-Rich Robotics Simulation and Control
Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
Demystifying Action Space Design for Robotic Manipulation Policies
APPLE: Toward General Active Perception via Reinforcement Learning
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
Can we detect failures without failure data? uncertainty-aware runtime failure detection for imitation learning policies
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
MIRAGE: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting
TD-MPC2: Scalable, Robust World Models for Continuous Control
RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation
Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization
Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks
ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM
Kimera: an open-source library for real-time metric-semantic localization and mapping
Composable action-conditioned predictors: Flexible off-policy learning for robot navigation
Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
Self-Supervised Visual Planning with Temporal Skip Connections.
Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics
Impedance Control: An Approach to Manipulation, Part I-Theory
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
EMBODIEDBENCH: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
WebArena: A Realistic Web Environment for Building Autonomous Agents
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
WorldModelBench: Judging Video Generation Models As World Models
SWE-BENCH: CAN LANGUAGE MODELS RESOLVE REAL-WORLD GITHUB ISSUES?
Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
RLBench: The Robot Learning Benchmark and Learning Environment
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
Evaluating Real-World Robot Manipulation Policies in Simulation
Contributions are very welcome — new papers, corrections, better categorization, or dead-link fixes.
main. Include the first public release
date and your section choice when adding a paper.Every correction, addition, and categorization improvement is appreciated. Contributors are recorded in the repository history and on GitHub's contributor graph.
We thank the researchers who make their papers, code, models, datasets, and project pages publicly available, as well as the maintainers of the related collections that help the community navigate this fast-moving field.
The official survey citation is not public yet. A verified BibTeX entry will be added after the manuscript is released. Until then, please link to this repository rather than using a provisional citation.
Released under CC0-1.0. The listed papers remain under their own licenses and copyright.
If you like the repository, please give us a star ⭐ — it is how we hear that it is useful.