showlab/Awesome-Multimodal-Embodied-Agent

Survey on Multimodal Embodied Agents: From Computer-Use to Robot-Use

See the code

README


This repository accompanies our survey on multimodal embodied agents (MMEAs): systems that couple multimodal reasoning with physical action and learn from the resulting feedback. It asks one question: what changes when a multimodal agent moves from computer-use to robot-use?

PAPAV: Perceive, Anticipate, Plan, Act, and Verify, for computer-use (top) and robot-use (bottom)

PAPAV: five capabilities that recur in every task loop, whether the agent operates a computer (top) or a robot (bottom).

  • 🔁 One lens, three settings. Perceive, Anticipate, Plan, Act, Verify are capabilities defined by contribution, not architecture, so multimodal agents, robotic systems, and MMEAs sit on the same coordinate.
  • 🤖 Physics changes the loop. Five physical constraints leave fewer choices fixed in advance and less room to reverse errors.
  • 🧪 Success hides gaps. Of 64 benchmarks, only 5 assess Anticipate and 3 assess Verify.

[!IMPORTANT] This area is growing quickly, and so is this list. If we missed a paper, or you have just published one of your own, please open an issue or send a pull request. We are always glad to add new work!

🌳 Taxonomy

Evolution of multimodal embodied agents across robotic systems, multimodal embodied agents, and multimodal agents

A chronological view of representative robotic systems, multimodal embodied agents, and multimodal agents.

📢 News

  • 2026.09 Check out our survey paper on Multimodal Embodied Agents!
  • 2026.09 First release: 310 papers across MMEAs, multimodal agents, robotic systems, and benchmarks.

📑 Contents


Prior surveys and reviews adjacent to our scope.

  • How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
    Survey Paper

  • Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
    Survey Paper Code

  • Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents
    Survey Paper

  • A Survey on Agentic Multimodal Large Language Models
    Survey Paper Code

  • Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
    Survey Paper

  • A Survey on (M)LLM-Based GUI Agents
    Survey Paper Code

  • Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
    Survey Paper Code

  • A Survey on Vision-Language-Action Models for Embodied AI
    Survey Paper Code

  • Large Multimodal Agents: A Survey
    Survey Paper Code

  • Agent AI: Surveying the Horizons of Multimodal Interaction
    Survey Paper

  • Integrated Task and Motion Planning
    Survey Paper


🤖 B. Multimodal Embodied Agents

Foundation-model-driven agents that couple multimodal perception and reasoning to embodied sensing and physical action in a closed loop — VLAs, LLM/VLM planners and critics for robots, language-conditioned robot world models, and embodied memory.

  • Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments
    MMEA Paper

  • Claude Plays Robotics
    MMEA Blog

  • CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
    MMEA Paper

  • ETA: A New Agentic Paradigm for Embodied Tasks
    MMEA Paper Code Project

  • RoboTTT: Context Scaling for Robot Policies
    MMEA Paper Project

  • Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA
    MMEA Paper Code Project

  • CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
    MMEA Paper Paper

  • eMEM: A Hybrid Spatio-Temporal Memory System For Embodied Agents
    MMEA Paper

  • G³VLA: Geometric inductive bias for Vision-Language-Action Models
    MMEA Paper

  • Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation
    MMEA Paper

  • Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
    MMEA Paper Code

  • KITE: Decoupling Kinematics and Interaction for Zero-Shot Cross-Embodiment Manipulation
    MMEA Paper

  • What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory
    MMEA Paper

  • What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents
    MMEA Paper

  • Perturbation-Based Uncertainty for Failure Detection in Vision-Language-Action Models
    MMEA Paper

  • Robot Critics that Sweat the Small Stuff
    MMEA Paper Code Project

  • Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
    MMEA Paper Code Project

  • VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models
    MMEA Paper

  • World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
    MMEA Paper Code

  • X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
    MMEA Paper Code Project

  • Dynamic Execution Commitment of Vision-Language-Action Models
    MMEA Paper

  • EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
    MMEA Paper Code

  • Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
    MMEA Paper

  • OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
    MMEA Paper

  • 3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
    MMEA Paper Code Project

  • Robot Planning and Situation Handling with Active Perception
    MMEA Paper

  • Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
    MMEA Paper

  • Using large language models for embodied planning introduces systematic safety risks
    MMEA Paper

  • Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System
    MMEA Paper

  • World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
    MMEA Paper Code

  • GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
    MMEA Paper

  • Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences
    MMEA Paper

  • RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
    MMEA Paper Code Project

  • AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
    MMEA Paper

  • LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
    MMEA Paper Code Project

  • Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World
    MMEA Paper

  • Recursive Belief Vision Language Action Models
    MMEA Paper

  • SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems
    MMEA Paper

  • Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
    MMEA Paper Code

  • VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
    MMEA Paper Code Project

  • World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
    MMEA Paper Code Project

  • PhyCritic: Multimodal Critic Models for Physical AI
    MMEA Paper

  • ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
    MMEA Paper Code Project

  • RoboReward: General-Purpose Vision-Language Reward Models for Robotics
    MMEA Paper

  • TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control
    MMEA Paper

  • Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
    MMEA Paper Code

  • EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
    MMEA Paper Code Project

  • Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
    MMEA Paper

  • Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
    MMEA Paper

  • AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
    MMEA Paper

  • MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning
    MMEA Paper

  • X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
    MMEA Paper Code Project

  • Ctrl-World: A Controllable Generative World Model for Robot Manipulation
    MMEA Paper Code Project

  • Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
    MMEA Paper

  • Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
    MMEA Paper

  • Kinodynamic Task and Motion Planning using VLM-guided and Interleaved Sampling
    MMEA Paper

  • Using VLM Reasoning to Constrain Task and Motion Planning
    MMEA Paper

  • Leave No Observation Behind: Real-time Correction for VLA Action Chunks
    MMEA Paper

  • A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
    MMEA Paper Code

  • Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
    MMEA Paper Code Project

  • DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
    MMEA Paper Code Project

  • Real-Time Execution of Action Chunking Flow Policies
    MMEA Paper

  • Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
    MMEA Paper

  • GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering
    MMEA Paper

  • Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
    MMEA Paper

  • RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
    MMEA Paper Code Project

  • SAFE: Multitask Failure Detection for Vision-Language-Action Models
    MMEA Paper Code Project

  • A Unified Framework for Real-Time Failure Handling in Robotics Using Vision-Language Models, Reactive Planner and Behavior Trees
    MMEA Paper

  • Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
    MMEA Paper Project

  • Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
    MMEA Paper Code Project

  • FAST: Efficient Action Tokenization for Vision-Language-Action Models
    MMEA Paper Code Project

  • UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
    MMEA Paper

  • Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
    MMEA Paper Code Project

  • Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
    MMEA Paper

  • RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
    MMEA Paper Code Project

  • AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
    MMEA Paper Code Project

  • π₀: A Vision-Language-Action Flow Model for General Robot Control
    MMEA Paper Code Project

  • CaStL: Constraints as Specifications through LLM Translation for Long-Horizon Task and Motion Planning
    MMEA Paper

  • ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
    MMEA Paper Code Project

  • Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
    MMEA Paper Code Project

  • Octo: An Open-Source Generalist Robot Policy
    MMEA Paper Code Project

  • Clio: Real-time Task-Driven Open-Set 3D Scene Graphs
    MMEA Paper Code

  • RoboDreamer: Learning Compositional World Models for Robot Imagination
    MMEA Paper Code Project

  • Explore until Confident: Efficient Exploration for Embodied Question Answering
    MMEA Paper Code Project

  • Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation
    MMEA Paper Code Project

  • Vision-Language Models for Robot Success Detection
    MMEA Paper

  • Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
    MMEA Paper Code

  • RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
    MMEA Paper Code Project

  • OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
    MMEA Paper Code Project

  • SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
    MMEA Paper Code Project

  • Open X-Embodiment: Robotic Learning Datasets and RT-X Models
    MMEA Paper Code Project

  • Learning Interactive Real-World Simulators
    MMEA Paper Project

  • Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
    MMEA Paper Code Project

  • ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
    MMEA Paper Code Project

  • Plug in the Safety Chip: Enforcing Constraints for LLM-driven Robot Agents
    MMEA Paper

  • DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
    MMEA Paper Project

  • RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
    MMEA Paper Project

  • VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
    MMEA Paper Code Project

  • Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
    MMEA Paper Project

  • SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
    MMEA Paper Project

  • REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
    MMEA Paper Code Project

  • Liv: Language-image representations and rewards for robotic control
    MMEA Paper Code Project

  • LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
    MMEA Paper Code

  • Vision-Language Models as Success Detectors
    MMEA Paper

  • LERF: Language Embedded Radiance Fields
    MMEA Paper Code Project

  • PaLM-E: An Embodied Multimodal Language Model
    MMEA Paper Project

  • Learning Universal Policies via Text-Guided Video Generation
    MMEA Paper Project

  • Openscene: 3d scene understanding with open vocabularies
    MMEA Paper Code Project

  • Visual language maps for robot navigation
    MMEA Paper Code Project

  • VIMA: Robot Manipulation with Multimodal Prompts
    MMEA Paper Code Project

  • Code as Policies: Language Model Programs for Embodied Control
    MMEA Paper Code Project

  • ProgPrompt: Generating Situated Robot Task Plans Using Large Language Models
    MMEA Paper Code Project

  • Inner Monologue: Embodied Reasoning through Planning with Language Models
    MMEA Paper Project

  • Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
    MMEA Paper Code Project

  • Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
    MMEA Paper Code Project

  • Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
    MMEA Paper Project


💻 C. Multimodal Agents

Agents built on multi-modal foundation models that perceive and act in digital environments — screens, browsers, documents, and APIs — without a physical body.

  • AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
    MMA Paper

  • Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
    MMA Paper

  • Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
    MMA Paper

  • From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
    MMA Paper

  • Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
    MMA Paper Code

  • Qwen-AgentWorld: Language World Models for General Agents
    MMA Paper Code

  • A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction
    MMA Paper

  • DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
    MMA Paper

  • Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes
    MMA Paper

  • VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
    MMA Paper Code

  • Confident and Wrong: Silent Semantic Failures in Coding Agents
    MMA Paper

  • Generative Visual Code Mobile World Models
    MMA Paper Code

  • Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
    MMA Paper

  • Code2world: A gui world model via renderable code generation
    MMA Paper Code

  • Mobiledreamer: Generative sketch world model for gui agent
    MMA Paper

  • Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents
    MMA Paper

  • Guitester: Enabling gui agents for exploratory defect discovery
    MMA Paper Code

  • WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
    MMA Paper

  • ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands
    MMA Paper Code Project

  • Active perception agent for omnimodal audio-video understanding
    MMA Paper

  • WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
    MMA Paper Code

  • GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes
    MMA Paper

  • Scaling Synthetic Task Generation for Agents via Exploration
    MMA Paper

  • Learning GUI Grounding with Spatial Reasoning from Visual Feedback
    MMA Paper

  • Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory
    MMA Paper Code

  • Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
    MMA Paper

  • Magentic-UI: Towards Human-in-the-loop Agentic Systems
    MMA Paper Code

  • Agent-SAMA: State-Aware Mobile Assistant
    MMA Paper

  • Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
    MMA Paper Code

  • Backtrackagent: Enhancing gui agent with error detection and backtracking mechanism
    MMA Paper

  • Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
    MMA Paper Code

  • Webevolver: Enhancing web agent self-improvement with co-evolving world model
    MMA Paper Code

  • ViMo: A Generative Visual GUI World Model for App Agents
    MMA Paper

  • AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
    MMA Paper Code

  • UI-TARS: Pioneering Automated GUI Interaction with Native Agents
    MMA Paper Code

  • Aguvis: Unified pure vision agents for autonomous gui interaction
    MMA Paper Code Project

  • Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
    MMA Paper Code

  • ShowUI: One Vision-Language-Action Model for GUI Visual Agent
    MMA Paper Code

  • Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
    MMA Paper Code

  • Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
    MMA Paper Code Project

  • OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
    MMA Paper Code Project

  • AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
    MMA Paper Code

  • The Impact of Element Ordering on LM Agent Performance
    MMA Paper Code

  • Tree Search for Language Model Agents
    MMA Paper Code Project

  • Pandora: Towards general world model with natural language actions and video states
    MMA Paper Code Project

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
    MMA Paper Code Project

  • Videoagent: A memory-augmented multimodal agent for video understanding
    MMA Paper Code Project

  • Genie: Generative Interactive Environments
    MMA Paper Project

  • Os-copilot: Towards generalist computer agents with self-improvement
    MMA Paper Code Project

  • Position: LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
    MMA Paper

  • SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
    MMA Paper Code

  • GPT-4V(ision) Is a Generalist Web Agent, If Grounded
    MMA Paper Code Project

  • CogAgent: A Visual Language Model for GUI Agents
    MMA Paper Code

  • Timechat: A time-sensitive multimodal large language model for long video understanding
    MMA Paper Code

  • LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
    MMA Paper Code Project

  • Salmonn: Towards generic hearing abilities for large language models
    MMA Paper Code

  • ControlLLM: Augment Language Models with Tools by Searching on Graphs
    MMA Paper Code

  • Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
    MMA Paper Code Project

  • Identifying the Risks of LM Agents with an LM-Emulated Sandbox
    MMA Paper Code

  • Learning to model the world with language
    MMA Paper Code Project

  • Avis: Autonomous visual information seeking with large language model agent
    MMA Paper

  • Kosmos-2: Grounding multimodal large language models to the world
    MMA Paper Code

  • Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
    MMA Paper Code Project

  • Voyager: An Open-Ended Embodied Agent with Large Language Models
    MMA Paper Code Project

  • Self-refine: Iterative refinement with self-feedback
    MMA Paper Code Project

  • Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
    MMA Paper Code

  • Mm-react: Prompting chatgpt for multimodal reasoning and action
    MMA Paper Code Project

  • Reflexion: Language Agents with Verbal Reinforcement Learning
    MMA Paper Code

  • ReAct: Synergizing Reasoning and Acting in Language Models
    MMA Paper Code Project

  • Mastering atari, go, chess and shogi by planning with a learned model
    MMA Paper


🦾 D. Robotic Systems

Robot learning, perception, and control in the physical world — vision-language-action models, manipulation, navigation, humanoids, and the data and simulators behind them.

  • ActionMap: Robot Policy Learning via Voxel Action Heatmap
    RS Paper

  • ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies
    RS Paper Code Project

  • ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation
    RS Paper

  • Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies
    RS Paper

  • Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents
    RS Paper Project

  • TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
    RS Paper

  • ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
    RS Paper

  • Going Beyond World Models and VLAs
    RS Blog

  • Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison
    RS Paper

  • ComFree-Sim: A GPU-Parallelized Analytical Contact Physics Engine for Scalable Contact-Rich Robotics Simulation and Control
    RS Paper

  • Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
    RS Paper Code Project

  • Demystifying Action Space Design for Robotic Manipulation Policies
    RS Paper

  • Mixture of Horizons in Action Chunking
    RS Paper Code

  • APPLE: Toward General Active Perception via Reinforcement Learning
    RS Paper Code Project

  • Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
    RS Paper Code Project

  • Can we detect failures without failure data? uncertainty-aware runtime failure detection for imitation learning policies
    RS Paper

  • Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
    RS Paper

  • 3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
    RS Paper Code Project

  • Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
    RS Paper Code Project

  • IRASim: A Fine-Grained World Model for Robot Manipulation
    RS Paper Code Project

  • DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset
    RS Paper Code Project

  • MIRAGE: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting
    RS Paper

  • TD-MPC2: Scalable, Robust World Models for Continuous Control
    RS Paper Code Project

  • RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
    RS Paper Project

  • Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
    RS Paper Code Project

  • Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
    RS Paper Code Project

  • Mastering diverse control tasks through world models
    RS Paper Code Project

  • See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation
    RS Paper Code Project

  • RT-1: Robotics Transformer for Real-World Control at Scale
    RS Paper Code Project

  • DayDreamer: World Models for Physical Robot Learning
    RS Paper Code Project

  • Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization
    RS Paper Code

  • Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks
    RS Paper

  • The PANDA Framework for Hierarchical Planning
    RS Paper

  • Multimodal sensor fusion with differentiable filters
    RS Paper

  • ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM
    RS Paper Code

  • Dream to control: Learning behaviors by latent imagination
    RS Paper Code Project

  • Kimera: an open-source library for real-time metric-semantic localization and mapping
    RS Paper Code

  • Learning Latent Dynamics for Planning from Pixels
    RS Paper Code Project

  • Composable action-conditioned predictors: Flexible off-policy learning for robot navigation
    RS Paper Code

  • Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
    RS Paper Code

  • Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
    RS Paper Code

  • Self-Supervised Visual Planning with Temporal Skip Connections.
    RS Paper

  • Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics
    RS Paper Code Project

  • Deep visual foresight for planning robot motion
    RS Paper

  • Hierarchical task and motion planning in the now
    RS Paper

  • Impedance Control: An Approach to Manipulation, Part I-Theory
    RS Paper


🧪 E. Benchmarks

Multimodal Embodied Agents

Sim

  • ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
    Bench Paper Code Project

  • ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
    Bench Paper Code Project

  • RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
    Bench Paper Code Project

  • IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
    Bench Paper Code

  • EMBODIEDBENCH: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
    Bench Paper Code Project

  • PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
    Bench Paper Code Project

  • EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
    Bench Paper Code Project

  • RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
    Bench Paper Code Project

Real

  • PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
    Bench Paper Project

Hybrid

  • CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
    Bench Paper Code Project

Multimodal Agents

Understanding

  • Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
    Bench Paper Code Project

  • Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
    Bench Paper Code Project

  • MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
    Bench Paper Code Project

Interaction

  • OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks
    Bench Paper Code Project

  • GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
    Bench Paper Code Project

  • OmniGAIA: Towards Native Omni-Modal AI Agents
    Bench Paper Code

  • AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
    Bench Paper Code Project

  • MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
    Bench Paper Code Project

  • iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
    Bench Paper Project

  • CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
    Bench Paper Code Project

  • τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
    Bench Paper Code

  • AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
    Bench Paper Code Project

  • MMInA: Benchmarking Multihop Multimodal Internet Agents
    Bench Paper Code Project

  • AgentStudio: A Toolkit for Building General Virtual Agents
    Bench Paper Code Project

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
    Bench Paper Code Project

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
    Bench Paper Code

  • GAIA: A Benchmark for General AI Assistants
    Bench Paper Hugging Face Project

  • WebArena: A Realistic Web Environment for Building Autonomous Agents
    Bench Paper Code Project

Generation

  • GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
    Bench Paper Code Project

  • PBench: A Physical AI Benchmark for World Models
    Bench Paper Project

  • WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
    Bench Paper Code

  • WorldModelBench: Judging Video Generation Models As World Models
    Bench Paper Code Project

  • SWE-BENCH: CAN LANGUAGE MODELS RESOLVE REAL-WORLD GITHUB ISSUES?
    Bench Paper Code Project

Robotic Systems

Sim

  • Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
    Bench Paper Code

  • MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
    Bench Paper

  • SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
    Bench Paper Code Project

  • RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
    Bench Paper Code Project

  • RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
    Bench Paper Code Project

  • LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
    Bench Paper Code Project

  • Meta-World+: An Improved, Standardized, RL Benchmark
    Bench Paper Code Project

  • VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
    Bench Paper Code Project

  • GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
    Bench Paper Code Project

  • BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
    Bench Paper Code Project

  • Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
    Bench Paper Code Project

  • LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
    Bench Paper Code Project

  • ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
    Bench Paper Code Project

  • CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
    Bench Paper Code Project

  • TEACh: Task-driven Embodied Agents that Chat
    Bench Paper Code

  • robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
    Bench Paper Code Project

  • ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
    Bench Paper Code Project

  • Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
    Bench Paper Code Project

  • RLBench: The Robot Learning Benchmark and Learning Environment
    Bench Paper Code Project

  • Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
    Bench Paper Code Project

Real

  • PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology
    Bench Paper Code

  • VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
    Bench Paper Code Project

  • RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
    Bench Paper Project

  • RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
    Bench Paper Code Project

  • AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
    Bench Paper Code Project

  • BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
    Bench Paper Code Project

  • FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
    Bench Paper Code Project

  • FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation
    Bench Paper Code Project

Hybrid

  • RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
    Bench Paper Code Project

  • Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance
    Bench Paper Code Project

  • RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
    Bench Paper Code Project

  • RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
    Bench Paper Code Project

  • RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
    Bench Paper Code Project

  • RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
    Bench Paper Code Project

  • Evaluating Real-World Robot Manipulation Policies in Simulation
    Bench Paper Code Project

  • HomeRobot: Open-Vocabulary Mobile Manipulation
    Bench Paper Code Project

🤝 Contributing

Contributions are very welcome — new papers, corrections, better categorization, or dead-link fixes.

  1. Read CONTRIBUTING.md for the entry format and placement rules.
  2. Edit this README and submit a pull request against main. Include the first public release date and your section choice when adding a paper.

Every correction, addition, and categorization improvement is appreciated. Contributors are recorded in the repository history and on GitHub's contributor graph.

🙏 Acknowledgements

We thank the researchers who make their papers, code, models, datasets, and project pages publicly available, as well as the maintainers of the related collections that help the community navigate this fast-moving field.

📝 Citation

The official survey citation is not public yet. A verified BibTeX entry will be added after the manuscript is released. Until then, please link to this repository rather than using a provisional citation.

⚖️ License

Released under CC0-1.0. The listed papers remain under their own licenses and copyright.

📌 Citation

If you like the repository, please give us a star ⭐ — it is how we hear that it is useful.

Contributors

Ziyi510

34 commits

ChenAnno

27 commits

anruihe

13 commits

showlab/Awesome-Multimodal-Embodied-Agent

Survey on Multimodal Embodied Agents: From Computer-Use to Robot-Use

See the code

README


This repository accompanies our survey on multimodal embodied agents (MMEAs): systems that couple multimodal reasoning with physical action and learn from the resulting feedback. It asks one question: what changes when a multimodal agent moves from computer-use to robot-use?

PAPAV: Perceive, Anticipate, Plan, Act, and Verify, for computer-use (top) and robot-use (bottom)

PAPAV: five capabilities that recur in every task loop, whether the agent operates a computer (top) or a robot (bottom).

  • 🔁 One lens, three settings. Perceive, Anticipate, Plan, Act, Verify are capabilities defined by contribution, not architecture, so multimodal agents, robotic systems, and MMEAs sit on the same coordinate.
  • 🤖 Physics changes the loop. Five physical constraints leave fewer choices fixed in advance and less room to reverse errors.
  • 🧪 Success hides gaps. Of 64 benchmarks, only 5 assess Anticipate and 3 assess Verify.

[!IMPORTANT] This area is growing quickly, and so is this list. If we missed a paper, or you have just published one of your own, please open an issue or send a pull request. We are always glad to add new work!

🌳 Taxonomy

Evolution of multimodal embodied agents across robotic systems, multimodal embodied agents, and multimodal agents

A chronological view of representative robotic systems, multimodal embodied agents, and multimodal agents.

📢 News

  • 2026.09 Check out our survey paper on Multimodal Embodied Agents!
  • 2026.09 First release: 310 papers across MMEAs, multimodal agents, robotic systems, and benchmarks.

📑 Contents


Prior surveys and reviews adjacent to our scope.

  • How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
    Survey Paper

  • Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
    Survey Paper Code

  • Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents
    Survey Paper

  • A Survey on Agentic Multimodal Large Language Models
    Survey Paper Code

  • Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
    Survey Paper

  • A Survey on (M)LLM-Based GUI Agents
    Survey Paper Code

  • Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
    Survey Paper Code

  • A Survey on Vision-Language-Action Models for Embodied AI
    Survey Paper Code

  • Large Multimodal Agents: A Survey
    Survey Paper Code

  • Agent AI: Surveying the Horizons of Multimodal Interaction
    Survey Paper

  • Integrated Task and Motion Planning
    Survey Paper


🤖 B. Multimodal Embodied Agents

Foundation-model-driven agents that couple multimodal perception and reasoning to embodied sensing and physical action in a closed loop — VLAs, LLM/VLM planners and critics for robots, language-conditioned robot world models, and embodied memory.

  • Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments
    MMEA Paper

  • Claude Plays Robotics
    MMEA Blog

  • CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
    MMEA Paper

  • ETA: A New Agentic Paradigm for Embodied Tasks
    MMEA Paper Code Project

  • RoboTTT: Context Scaling for Robot Policies
    MMEA Paper Project

  • Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA
    MMEA Paper Code Project

  • CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
    MMEA Paper Paper

  • eMEM: A Hybrid Spatio-Temporal Memory System For Embodied Agents
    MMEA Paper

  • G³VLA: Geometric inductive bias for Vision-Language-Action Models
    MMEA Paper

  • Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation
    MMEA Paper

  • Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
    MMEA Paper Code

  • KITE: Decoupling Kinematics and Interaction for Zero-Shot Cross-Embodiment Manipulation
    MMEA Paper

  • What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory
    MMEA Paper

  • What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents
    MMEA Paper

  • Perturbation-Based Uncertainty for Failure Detection in Vision-Language-Action Models
    MMEA Paper

  • Robot Critics that Sweat the Small Stuff
    MMEA Paper Code Project

  • Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
    MMEA Paper Code Project

  • VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models
    MMEA Paper

  • World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
    MMEA Paper Code

  • X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
    MMEA Paper Code Project

  • Dynamic Execution Commitment of Vision-Language-Action Models
    MMEA Paper

  • EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
    MMEA Paper Code

  • Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
    MMEA Paper

  • OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
    MMEA Paper

  • 3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
    MMEA Paper Code Project

  • Robot Planning and Situation Handling with Active Perception
    MMEA Paper

  • Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
    MMEA Paper

  • Using large language models for embodied planning introduces systematic safety risks
    MMEA Paper

  • Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System
    MMEA Paper

  • World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
    MMEA Paper Code

  • GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
    MMEA Paper

  • Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences
    MMEA Paper

  • RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
    MMEA Paper Code Project

  • AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
    MMEA Paper

  • LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
    MMEA Paper Code Project

  • Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World
    MMEA Paper

  • Recursive Belief Vision Language Action Models
    MMEA Paper

  • SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems
    MMEA Paper

  • Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
    MMEA Paper Code

  • VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
    MMEA Paper Code Project

  • World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
    MMEA Paper Code Project

  • PhyCritic: Multimodal Critic Models for Physical AI
    MMEA Paper

  • ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
    MMEA Paper Code Project

  • RoboReward: General-Purpose Vision-Language Reward Models for Robotics
    MMEA Paper

  • TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control
    MMEA Paper

  • Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
    MMEA Paper Code

  • EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
    MMEA Paper Code Project

  • Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
    MMEA Paper

  • Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
    MMEA Paper

  • AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
    MMEA Paper

  • MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning
    MMEA Paper

  • X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
    MMEA Paper Code Project

  • Ctrl-World: A Controllable Generative World Model for Robot Manipulation
    MMEA Paper Code Project

  • Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
    MMEA Paper

  • Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
    MMEA Paper

  • Kinodynamic Task and Motion Planning using VLM-guided and Interleaved Sampling
    MMEA Paper

  • Using VLM Reasoning to Constrain Task and Motion Planning
    MMEA Paper

  • Leave No Observation Behind: Real-time Correction for VLA Action Chunks
    MMEA Paper

  • A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
    MMEA Paper Code

  • Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
    MMEA Paper Code Project

  • DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
    MMEA Paper Code Project

  • Real-Time Execution of Action Chunking Flow Policies
    MMEA Paper

  • Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
    MMEA Paper

  • GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering
    MMEA Paper

  • Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
    MMEA Paper

  • RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
    MMEA Paper Code Project

  • SAFE: Multitask Failure Detection for Vision-Language-Action Models
    MMEA Paper Code Project

  • A Unified Framework for Real-Time Failure Handling in Robotics Using Vision-Language Models, Reactive Planner and Behavior Trees
    MMEA Paper

  • Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
    MMEA Paper Project

  • Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
    MMEA Paper Code Project

  • FAST: Efficient Action Tokenization for Vision-Language-Action Models
    MMEA Paper Code Project

  • UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
    MMEA Paper

  • Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
    MMEA Paper Code Project

  • Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
    MMEA Paper

  • RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
    MMEA Paper Code Project

  • AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
    MMEA Paper Code Project

  • π₀: A Vision-Language-Action Flow Model for General Robot Control
    MMEA Paper Code Project

  • CaStL: Constraints as Specifications through LLM Translation for Long-Horizon Task and Motion Planning
    MMEA Paper

  • ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
    MMEA Paper Code Project

  • Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
    MMEA Paper Code Project

  • Octo: An Open-Source Generalist Robot Policy
    MMEA Paper Code Project

  • Clio: Real-time Task-Driven Open-Set 3D Scene Graphs
    MMEA Paper Code

  • RoboDreamer: Learning Compositional World Models for Robot Imagination
    MMEA Paper Code Project

  • Explore until Confident: Efficient Exploration for Embodied Question Answering
    MMEA Paper Code Project

  • Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation
    MMEA Paper Code Project

  • Vision-Language Models for Robot Success Detection
    MMEA Paper

  • Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
    MMEA Paper Code

  • RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
    MMEA Paper Code Project

  • OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
    MMEA Paper Code Project

  • SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
    MMEA Paper Code Project

  • Open X-Embodiment: Robotic Learning Datasets and RT-X Models
    MMEA Paper Code Project

  • Learning Interactive Real-World Simulators
    MMEA Paper Project

  • Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
    MMEA Paper Code Project

  • ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
    MMEA Paper Code Project

  • Plug in the Safety Chip: Enforcing Constraints for LLM-driven Robot Agents
    MMEA Paper

  • DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
    MMEA Paper Project

  • RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
    MMEA Paper Project

  • VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
    MMEA Paper Code Project

  • Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
    MMEA Paper Project

  • SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
    MMEA Paper Project

  • REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
    MMEA Paper Code Project

  • Liv: Language-image representations and rewards for robotic control
    MMEA Paper Code Project

  • LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
    MMEA Paper Code

  • Vision-Language Models as Success Detectors
    MMEA Paper

  • LERF: Language Embedded Radiance Fields
    MMEA Paper Code Project

  • PaLM-E: An Embodied Multimodal Language Model
    MMEA Paper Project

  • Learning Universal Policies via Text-Guided Video Generation
    MMEA Paper Project

  • Openscene: 3d scene understanding with open vocabularies
    MMEA Paper Code Project

  • Visual language maps for robot navigation
    MMEA Paper Code Project

  • VIMA: Robot Manipulation with Multimodal Prompts
    MMEA Paper Code Project

  • Code as Policies: Language Model Programs for Embodied Control
    MMEA Paper Code Project

  • ProgPrompt: Generating Situated Robot Task Plans Using Large Language Models
    MMEA Paper Code Project

  • Inner Monologue: Embodied Reasoning through Planning with Language Models
    MMEA Paper Project

  • Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
    MMEA Paper Code Project

  • Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
    MMEA Paper Code Project

  • Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
    MMEA Paper Project


💻 C. Multimodal Agents

Agents built on multi-modal foundation models that perceive and act in digital environments — screens, browsers, documents, and APIs — without a physical body.

  • AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
    MMA Paper

  • Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
    MMA Paper

  • Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
    MMA Paper

  • From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
    MMA Paper

  • Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human
    MMA Paper Code

  • Qwen-AgentWorld: Language World Models for General Agents
    MMA Paper Code

  • A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction
    MMA Paper

  • DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
    MMA Paper

  • Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes
    MMA Paper

  • VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
    MMA Paper Code

  • Confident and Wrong: Silent Semantic Failures in Coding Agents
    MMA Paper

  • Generative Visual Code Mobile World Models
    MMA Paper Code

  • Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
    MMA Paper

  • Code2world: A gui world model via renderable code generation
    MMA Paper Code

  • Mobiledreamer: Generative sketch world model for gui agent
    MMA Paper

  • Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents
    MMA Paper

  • Guitester: Enabling gui agents for exploratory defect discovery
    MMA Paper Code

  • WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
    MMA Paper

  • ShowUI-π: Flow-based Generative Models as GUI Dexterous Hands
    MMA Paper Code Project

  • Active perception agent for omnimodal audio-video understanding
    MMA Paper

  • WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
    MMA Paper Code

  • GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes
    MMA Paper

  • Scaling Synthetic Task Generation for Agents via Exploration
    MMA Paper

  • Learning GUI Grounding with Spatial Reasoning from Visual Feedback
    MMA Paper

  • Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory
    MMA Paper Code

  • Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
    MMA Paper

  • Magentic-UI: Towards Human-in-the-loop Agentic Systems
    MMA Paper Code

  • Agent-SAMA: State-Aware Mobile Assistant
    MMA Paper

  • Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
    MMA Paper Code

  • Backtrackagent: Enhancing gui agent with error detection and backtracking mechanism
    MMA Paper

  • Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
    MMA Paper Code

  • Webevolver: Enhancing web agent self-improvement with co-evolving world model
    MMA Paper Code

  • ViMo: A Generative Visual GUI World Model for App Agents
    MMA Paper

  • AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
    MMA Paper Code

  • UI-TARS: Pioneering Automated GUI Interaction with Native Agents
    MMA Paper Code

  • Aguvis: Unified pure vision agents for autonomous gui interaction
    MMA Paper Code Project

  • Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
    MMA Paper Code

  • ShowUI: One Vision-Language-Action Model for GUI Visual Agent
    MMA Paper Code

  • Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
    MMA Paper Code

  • Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
    MMA Paper Code Project

  • OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
    MMA Paper Code Project

  • AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
    MMA Paper Code

  • The Impact of Element Ordering on LM Agent Performance
    MMA Paper Code

  • Tree Search for Language Model Agents
    MMA Paper Code Project

  • Pandora: Towards general world model with natural language actions and video states
    MMA Paper Code Project

  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
    MMA Paper Code Project

  • Videoagent: A memory-augmented multimodal agent for video understanding
    MMA Paper Code Project

  • Genie: Generative Interactive Environments
    MMA Paper Project

  • Os-copilot: Towards generalist computer agents with self-improvement
    MMA Paper Code Project

  • Position: LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
    MMA Paper

  • SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
    MMA Paper Code

  • GPT-4V(ision) Is a Generalist Web Agent, If Grounded
    MMA Paper Code Project

  • CogAgent: A Visual Language Model for GUI Agents
    MMA Paper Code

  • Timechat: A time-sensitive multimodal large language model for long video understanding
    MMA Paper Code

  • LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
    MMA Paper Code Project

  • Salmonn: Towards generic hearing abilities for large language models
    MMA Paper Code

  • ControlLLM: Augment Language Models with Tools by Searching on Graphs
    MMA Paper Code

  • Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
    MMA Paper Code Project

  • Identifying the Risks of LM Agents with an LM-Emulated Sandbox
    MMA Paper Code

  • Learning to model the world with language
    MMA Paper Code Project

  • Avis: Autonomous visual information seeking with large language model agent
    MMA Paper

  • Kosmos-2: Grounding multimodal large language models to the world
    MMA Paper Code

  • Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
    MMA Paper Code Project

  • Voyager: An Open-Ended Embodied Agent with Large Language Models
    MMA Paper Code Project

  • Self-refine: Iterative refinement with self-feedback
    MMA Paper Code Project

  • Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
    MMA Paper Code

  • Mm-react: Prompting chatgpt for multimodal reasoning and action
    MMA Paper Code Project

  • Reflexion: Language Agents with Verbal Reinforcement Learning
    MMA Paper Code

  • ReAct: Synergizing Reasoning and Acting in Language Models
    MMA Paper Code Project

  • Mastering atari, go, chess and shogi by planning with a learned model
    MMA Paper


🦾 D. Robotic Systems

Robot learning, perception, and control in the physical world — vision-language-action models, manipulation, navigation, humanoids, and the data and simulators behind them.

  • ActionMap: Robot Policy Learning via Voxel Action Heatmap
    RS Paper

  • ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies
    RS Paper Code Project

  • ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation
    RS Paper

  • Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies
    RS Paper

  • Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents
    RS Paper Project

  • TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
    RS Paper

  • ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
    RS Paper

  • Going Beyond World Models and VLAs
    RS Blog

  • Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison
    RS Paper

  • ComFree-Sim: A GPU-Parallelized Analytical Contact Physics Engine for Scalable Contact-Rich Robotics Simulation and Control
    RS Paper

  • Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
    RS Paper Code Project

  • Demystifying Action Space Design for Robotic Manipulation Policies
    RS Paper

  • Mixture of Horizons in Action Chunking
    RS Paper Code

  • APPLE: Toward General Active Perception via Reinforcement Learning
    RS Paper Code Project

  • Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
    RS Paper Code Project

  • Can we detect failures without failure data? uncertainty-aware runtime failure detection for imitation learning policies
    RS Paper

  • Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
    RS Paper

  • 3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
    RS Paper Code Project

  • Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
    RS Paper Code Project

  • IRASim: A Fine-Grained World Model for Robot Manipulation
    RS Paper Code Project

  • DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset
    RS Paper Code Project

  • MIRAGE: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting
    RS Paper

  • TD-MPC2: Scalable, Robust World Models for Continuous Control
    RS Paper Code Project

  • RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
    RS Paper Project

  • Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
    RS Paper Code Project

  • Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
    RS Paper Code Project

  • Mastering diverse control tasks through world models
    RS Paper Code Project

  • See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation
    RS Paper Code Project

  • RT-1: Robotics Transformer for Real-World Control at Scale
    RS Paper Code Project

  • DayDreamer: World Models for Physical Robot Learning
    RS Paper Code Project

  • Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization
    RS Paper Code

  • Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks
    RS Paper

  • The PANDA Framework for Hierarchical Planning
    RS Paper

  • Multimodal sensor fusion with differentiable filters
    RS Paper

  • ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM
    RS Paper Code

  • Dream to control: Learning behaviors by latent imagination
    RS Paper Code Project

  • Kimera: an open-source library for real-time metric-semantic localization and mapping
    RS Paper Code

  • Learning Latent Dynamics for Planning from Pixels
    RS Paper Code Project

  • Composable action-conditioned predictors: Flexible off-policy learning for robot navigation
    RS Paper Code

  • Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
    RS Paper Code

  • Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
    RS Paper Code

  • Self-Supervised Visual Planning with Temporal Skip Connections.
    RS Paper

  • Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics
    RS Paper Code Project

  • Deep visual foresight for planning robot motion
    RS Paper

  • Hierarchical task and motion planning in the now
    RS Paper

  • Impedance Control: An Approach to Manipulation, Part I-Theory
    RS Paper


🧪 E. Benchmarks

Multimodal Embodied Agents

Sim

  • ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
    Bench Paper Code Project

  • ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
    Bench Paper Code Project

  • RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
    Bench Paper Code Project

  • IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
    Bench Paper Code

  • EMBODIEDBENCH: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
    Bench Paper Code Project

  • PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
    Bench Paper Code Project

  • EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
    Bench Paper Code Project

  • RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
    Bench Paper Code Project

Real

  • PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
    Bench Paper Project

Hybrid

  • CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
    Bench Paper Code Project

Multimodal Agents

Understanding

  • Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
    Bench Paper Code Project

  • Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
    Bench Paper Code Project

  • MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
    Bench Paper Code Project

Interaction

  • OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks
    Bench Paper Code Project

  • GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
    Bench Paper Code Project

  • OmniGAIA: Towards Native Omni-Modal AI Agents
    Bench Paper Code

  • AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
    Bench Paper Code Project

  • MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
    Bench Paper Code Project

  • iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
    Bench Paper Project

  • CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
    Bench Paper Code Project

  • τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
    Bench Paper Code

  • AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
    Bench Paper Code Project

  • MMInA: Benchmarking Multihop Multimodal Internet Agents
    Bench Paper Code Project

  • AgentStudio: A Toolkit for Building General Virtual Agents
    Bench Paper Code Project

  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
    Bench Paper Code Project

  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
    Bench Paper Code

  • GAIA: A Benchmark for General AI Assistants
    Bench Paper Hugging Face Project

  • WebArena: A Realistic Web Environment for Building Autonomous Agents
    Bench Paper Code Project

Generation

  • GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
    Bench Paper Code Project

  • PBench: A Physical AI Benchmark for World Models
    Bench Paper Project

  • WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
    Bench Paper Code

  • WorldModelBench: Judging Video Generation Models As World Models
    Bench Paper Code Project

  • SWE-BENCH: CAN LANGUAGE MODELS RESOLVE REAL-WORLD GITHUB ISSUES?
    Bench Paper Code Project

Robotic Systems

Sim

  • Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
    Bench Paper Code

  • MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
    Bench Paper

  • SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
    Bench Paper Code Project

  • RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
    Bench Paper Code Project

  • RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
    Bench Paper Code Project

  • LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
    Bench Paper Code Project

  • Meta-World+: An Improved, Standardized, RL Benchmark
    Bench Paper Code Project

  • VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
    Bench Paper Code Project

  • GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
    Bench Paper Code Project

  • BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
    Bench Paper Code Project

  • Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
    Bench Paper Code Project

  • LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
    Bench Paper Code Project

  • ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
    Bench Paper Code Project

  • CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
    Bench Paper Code Project

  • TEACh: Task-driven Embodied Agents that Chat
    Bench Paper Code

  • robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
    Bench Paper Code Project

  • ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
    Bench Paper Code Project

  • Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
    Bench Paper Code Project

  • RLBench: The Robot Learning Benchmark and Learning Environment
    Bench Paper Code Project

  • Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
    Bench Paper Code Project

Real

  • PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology
    Bench Paper Code

  • VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
    Bench Paper Code Project

  • RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
    Bench Paper Project

  • RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
    Bench Paper Code Project

  • AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
    Bench Paper Code Project

  • BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
    Bench Paper Code Project

  • FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
    Bench Paper Code Project

  • FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation
    Bench Paper Code Project

Hybrid

  • RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
    Bench Paper Code Project

  • Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance
    Bench Paper Code Project

  • RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
    Bench Paper Code Project

  • RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
    Bench Paper Code Project

  • RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
    Bench Paper Code Project

  • RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
    Bench Paper Code Project

  • Evaluating Real-World Robot Manipulation Policies in Simulation
    Bench Paper Code Project

  • HomeRobot: Open-Vocabulary Mobile Manipulation
    Bench Paper Code Project

🤝 Contributing

Contributions are very welcome — new papers, corrections, better categorization, or dead-link fixes.

  1. Read CONTRIBUTING.md for the entry format and placement rules.
  2. Edit this README and submit a pull request against main. Include the first public release date and your section choice when adding a paper.

Every correction, addition, and categorization improvement is appreciated. Contributors are recorded in the repository history and on GitHub's contributor graph.

🙏 Acknowledgements

We thank the researchers who make their papers, code, models, datasets, and project pages publicly available, as well as the maintainers of the related collections that help the community navigate this fast-moving field.

📝 Citation

The official survey citation is not public yet. A verified BibTeX entry will be added after the manuscript is released. Until then, please link to this repository rather than using a provisional citation.

⚖️ License

Released under CC0-1.0. The listed papers remain under their own licenses and copyright.

📌 Citation

If you like the repository, please give us a star ⭐ — it is how we hear that it is useful.

Contributors

Ziyi510

34 commits

ChenAnno

27 commits

anruihe

13 commits