dongqianyu99/Awesome-Agentic-Embodided-Systems

A curated research map of agentic embodied systems, covering foundation models, embodied harnesses, in-context adaptation, self-improvement, and evaluation.

4

1 commits

updated Aug 29, 2026

See the code

README

🤖 Awesome Agentic Embodied Systems

Awesome License: BSD-3-Clause PRs Welcome

A curated research map of the models, runtimes, and learning loops behind goal-directed embodied agents.

Last literature sweep: 2026-08-30


News and Updates

  • 2026-08 — Initial release. Published the initial system-oriented research map.
  • Ongoing — Contributions welcome. Missing papers, corrected links, and better placements are welcome through focused issues or pull requests; see CONTRIBUTING.md.

Overview

Motivation: Why Agentic Embodied Systems, Why Now?

A useful way to read recent language and vision-language AI is:

scaling -> foundation models -> agents / harnesses

Empirical scaling laws linked model performance to data, parameters, and compute. Broad pretraining produced adaptable foundation models. As multimodal reasoning, code generation, and tool use improved, research expanded toward reliable multi-step goal completion. ReAct interleaved reasoning with action and observation, and SWE-agent showed that agent interfaces materially shape behavior. A harness supplies runtime interfaces, persistent state, and outcome-driven control around a model. All three stages continue to co-evolve.

Embodied AI can draw on semantic and multimodal priors that are costly to learn from robot trajectories alone. RT-2 transferred Internet-scale vision-language training to robotic control, Open X-Embodiment found positive transfer across embodiments, and π₀.₅ combined heterogeneous robot, semantic, and web data for manipulation in unseen homes. Together, these works provide evidence for reusable cross-task transfer and control within the studied settings.

Physical interaction makes state uncertain, action constrained, and recovery costly. System interfaces and outcome verification therefore become central to long-horizon performance: SayCan grounds plans in executable affordances, and LIBERO-PRO measures degradation under task and environment perturbations.

Reusable models and skills now make harness design a visible research target. Thea, ENPIRE, and ASPIRE study closed-loop state management, recovery, and improvement in embodied systems. The central question is how model and runtime capabilities should scale together in grounded interaction.

Scope and System Map

This repository maps agentic embodied research across enabling models, system design, and evaluation. It prioritizes verified primary links, concise synthesis, and one primary section per work.

Papers are the default evidence unit. Technically substantive non-paper sources are included when they materially shape the field; source types and company reports are labeled explicitly.

Agentic embodied systems organize embodied models and policies through a runtime that supports goal-directed interaction, feedback, and improvement.

The scope covers physical and simulated environments, within and across episodes.

The collection moves from scaling evidence and reusable models through the agentic system stack to evaluation. Historical robotics and adjacent software-agent precedents appear within the corresponding module. Entries follow earliest public release, use one primary home, and cross-reference other relationships.

Tutorials, Surveys, and Starter Resources

  • Embodied AI Survey: "A Survey of Embodied AI: From Simulators to Research Tasks". arXiv
  • Large Language Models for Robotics: "Large Language Models for Robotics: A Survey". arXiv
  • Foundation Models in Robotics: "Foundation Models in Robotics: Applications, Challenges, and the Future". arXiv Code
  • LLMs for Robotics: Opportunities and Challenges: "Large Language Models for Robotics: Opportunities, Challenges, and Perspectives". arXiv
  • Robotics with Foundation Models: "A Survey on Robotics with Foundation Models: toward Embodied AI". arXiv
  • VLA Survey: "A Survey on Vision-Language-Action Models for Embodied AI". arXiv
  • Building Effective Agents: "Building effective agents". Website
  • Harness Engineering with Codex: "Harness engineering: leveraging Codex in an agent-first world". Website
  • Harness Engineering for Physical AI: "Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer". arXiv
  • Scaling Laws, Carefully: "Scaling Laws, Carefully". Website
  • Embodied Collective Intelligence: "When Multi-Robot Systems Meet Agentic AI: Towards Embodied Collective Intelligence". arXiv
  • Harness Engineering for Self-Improvement: "Harness Engineering for Self-Improvement". Website

Scaling Laws and Scaling Evidence

These works connect scaling research in language, vision-language, and embodied AI.

Language and Vision-Language Models

  • Kaplan Scaling Laws: "Scaling Laws for Neural Language Models". arXiv
  • Chinchilla: "Training Compute-Optimal Large Language Models". arXiv
  • OpenCLIP Scaling: "Reproducible scaling laws for contrastive language-image learning". arXiv Code
  • Zero-Shot Data Limit: "No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance". arXiv

Embodied AI Scaling Evidence

Embodied scaling evidence remains less mature than language-model scaling laws. Proprietary company reports are labeled to reflect their limited reproducibility.

  • Manipulation Data Scaling: "Data Scaling Laws in Imitation Learning for Robotic Manipulation". arXiv Website
  • Agents and World Models: "Scaling Laws for Pre-training Agents and World Models". arXiv
  • GEN-0: "GEN-0 / Embodied Foundation Models That Scale with Physical Interaction". Website
  • EgoScale: "EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data". arXiv Website
  • Precision Scaling: "The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation". arXiv
  • Dyna-2: "Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models". Website

Foundation Models, Policies, and Data

Foundation models, generalist policies, and shared datasets supply reusable capabilities for agentic runtimes.

Generalist Embodied Models and Policies

This section selects models and policies that materially advance reusable embodied capability or its integration into agentic systems. Checkpoint updates are grouped by model family.

  • CLIPort: "CLIPort: What and Where Pathways for Robotic Manipulation". arXiv Website
  • BC-Z: "BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning". arXiv
  • Gato: "A Generalist Agent". arXiv
  • VIMA: "VIMA: General Robot Manipulation with Multimodal Prompts". arXiv Website
  • RT-1: "RT-1: Robotics Transformer for Real-World Control at Scale". arXiv
  • PaLM-E: "PaLM-E: An Embodied Multimodal Language Model". arXiv
  • RT-2: "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control". arXiv
  • Open X-Embodiment / RT-X: "Open X-Embodiment: Robotic Learning Datasets and RT-X Models". arXiv Website
  • RT-H: "RT-H: Action Hierarchies Using Language". arXiv Website
  • Octo: "Octo: An Open-Source Generalist Robot Policy". arXiv Website
  • OpenVLA: "OpenVLA: An Open-Source Vision-Language-Action Model". arXiv Website
  • RDT-1B: "RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation". arXiv
  • π₀: "$π_0$: A Vision-Language-Action Flow Model for General Robot Control". arXiv Website
  • Magma: "Magma: A Foundation Model for Multimodal AI Agents". arXiv Website Code
  • Helix family: "Helix: A Vision-Language-Action Model for Generalist Humanoid Control". Helix Helix 02
  • GR00T N family: "GR00T N1: An Open Foundation Model for Generalist Humanoid Robots". arXiv N1.6 N1.7 Code
  • Gemini Robotics: "Gemini Robotics: Bringing AI into the Physical World". arXiv Website
  • π₀.₅: "$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization". arXiv
  • SmolVLA: "SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics". arXiv
  • TRI LBM 1.0: "A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation". arXiv Website
  • GR-3: "GR-3 Technical Report". arXiv Website
  • Gemini Robotics 1.5: "Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer". arXiv
  • X-VLA: "X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model". arXiv Website Code
  • LingBot-VLA: "A Pragmatic VLA Foundation Model". arXiv
  • DreamZero: "World Action Models are Zero-shot Policies". arXiv Website Code
  • GEN-1: "GEN-1: Scaling Embodied Foundation Models to Mastery". Website
  • π₀.₇: "$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities". arXiv Website
  • MolmoAct 2: "MolmoAct2: Action Reasoning Models for Real-world Deployment". arXiv Website Code
  • Embodied-R1.5: "Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models". arXiv Website Code
  • Vesta: "Vesta: A Generalist Embodied Reasoning Model". arXiv Website
  • LingBot-VLA 2.0: "From Foundation to Application: Improving VLA Models in Practice". arXiv Website Code
  • Xiaomi-Robotics-1: "Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories". arXiv Website Code
  • Gemini Robotics 2: "Gemini Robotics 2 brings whole body intelligence to robots". Website

Promptable and In-Context Robot Policies

Policies that adapt at inference time through demonstrations, interaction history, or persistent context.

  • ICRT: "In-Context Imitation Learning via Next-Token Prediction". arXiv Website
  • Instant Policy: "Instant Policy: In-Context Imitation Learning via Graph Diffusion". arXiv Website
  • LocoFormer: "LocoFormer: Generalist Locomotion via Long-context Adaptation". arXiv Website
  • BPP: "Behavior Prompting Policy: Demonstrations as Prompts for Manipulation". arXiv Website
  • RoboTTT: "RoboTTT: Context Scaling for Robot Policies". arXiv Website
  • Skild S1: "Introducing S1: In-Context Learning for Robotics". Website
  • GEN-1.5: "GEN-1.5: Embodied Foundation Models are One-Shot Learners". Website

Data and Open Infrastructure

  • RoboNet: "RoboNet: Large-Scale Multi-Robot Learning". arXiv Website
  • BridgeData V2: "BridgeData V2: A Dataset for Robot Learning at Scale". arXiv Website
  • DROID: "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset". arXiv Website
  • LeRobot: "LeRobot". Website Code
  • OpenVLA-OFT: "Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success". arXiv Website
  • AgiBot World / GO-1: "AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems". arXiv Website Code
  • DreamGen: "DreamGen: Unlocking Generalization in Robot Learning through Video World Models". arXiv Website Code

Agentic Embodied System Stack

The primary classification axis groups each work by its main system responsibility, with embodied systems foregrounded and relevant precedents placed alongside them.

Architectures and Harnesses

Architectures, middleware, and runtimes that span multiple stages of the interaction loop.

  • Subsumption Architecture: "A Robust Layered Control System for a Mobile Robot". Paper
  • ATLANTIS: "Integrating Planning and Reacting in a Heterogeneous Asynchronous Architecture for Controlling Real-World Mobile Robots". Paper
  • 3T: "Experiences with an Architecture for Intelligent, Reactive Agents". Paper
  • Remote Agent: "Remote Agent: To Boldly Go Where No AI System Has Gone Before". Paper
  • CRAM: "CRAM—A Cognitive Robot Abstract Machine for Everyday Manipulation in Human Environments". Paper Website
  • OK-Robot: "OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics". arXiv Website
  • ROS-LLM: "ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning". arXiv Code
  • Being-0: "Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills". arXiv Website
  • RoboNeuron: "RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI". arXiv Code
  • VoLo / RoboVoLo: "VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation". arXiv Website Code
  • Harness VLA: "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents". arXiv
  • OpenETA: "ETA: A New Agentic Paradigm for Embodied Tasks". arXiv Website Code
  • Thea: "Towards the Harness of Embodied Agents". arXiv Website Code

State, Context, and Memory

Methods for constructing world state and retaining knowledge across steps or episodes.

Embodied Systems and Robotics Foundations

  • KnowRob: "KnowRob—Knowledge Processing for Autonomous Personal Robots". Website
  • RoboEarth: "RoboEarth—A World Wide Web for Robots". Website
  • RoboBrain: "RoboBrain: Large-Scale Knowledge Engine for Robots". arXiv
  • Statler: "Statler: State-Maintaining Language Models for Embodied Reasoning". arXiv
  • HELPER: "Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language Models". arXiv Website
  • HELPER-X: "HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models". arXiv
  • KARMA: "KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems". arXiv
  • ViReSkill: "ViReSkill: Vision-Grounded Replanning with Skill Memory for LLM-Based Planning in Lifelong Robot Learning". arXiv
  • MEM: "MEM: Multi-Scale Embodied Memory for Vision Language Action Models". arXiv Website

Adjacent Memory Architectures

  • Reflexion: "Reflexion: Language Agents with Verbal Reinforcement Learning". OpenReview
  • Generative Agents: "Generative Agents: Interactive Simulacra of Human Behavior". arXiv
  • ExpeL: "ExpeL: LLM Agents Are Experiential Learners". arXiv Paper
  • MemGPT: "MemGPT: Towards LLMs as Operating Systems". arXiv

Planning and Orchestration

Goal decomposition and high-level coordination of plans, actions, and skills.

  • STRIPS: "STRIPS: A New Approach to the Application of Theorem Proving to Problem Solving". Paper
  • Generalized Robot Plans: "Learning and Executing Generalized Robot Plans". Paper
  • PRS: "Reactive Reasoning and Planning". Paper
  • PDDL: "PDDL—The Planning Domain Definition Language". Paper
  • HTAMP: "Hierarchical Task and Motion Planning in the Now". Paper
  • ROSPlan: "ROSPlan: Planning in the Robot Operating System". Paper Code
  • Logic-Geometric Programming: "Logic-Geometric Programming: An Optimization-Based Approach to Combined Task and Motion Planning". Paper
  • PDDLStream: "PDDLStream: Integrating Symbolic Planners and Blackbox Samplers via Optimistic Adaptive Planning". arXiv
  • Zero-Shot Planner: "Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents". arXiv
  • SayCan: "Do As I Can, Not As I Say: Grounding Language in Robotic Affordances". arXiv Website
  • ProgPrompt: "ProgPrompt: Generating Situated Robot Task Plans using Large Language Models". arXiv
  • LLM-Planner: "LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models". arXiv
  • ASC: "ASC: Adaptive Skill Coordination for Robotic Mobile Manipulation". arXiv
  • LLM+P: "LLM+P: Empowering Large Language Models with Optimal Planning Proficiency". arXiv Code
  • SayPlan: "SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning". arXiv Website
  • Reflective Planning: "Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation". arXiv Website
  • Hi Robot: "Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models". arXiv Website
  • VLA-Reasoner: "VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search". arXiv
  • Goal2Skill: "Goal2Skill: Long-Horizon Manipulation with Adaptive Planning and Reflection". arXiv
  • τ₀-VLA: "$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation". arXiv

Tools, Programs, and Action Grounding

Representations and interfaces that turn selected capabilities into executable actions.

Embodied Systems and Robotics Foundations

  • RAPs: "An Investigation into Reactive Planning in Complex Domains". Paper
  • Teleo-Reactive Programs: "Teleo-Reactive Programs for Agent Control". arXiv
  • Options: "Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning". Paper
  • Walk the Talk / MARCO: "Walk the Talk: Connecting Language, Knowledge, and Action in Route Instructions". Paper
  • : "Understanding Natural Language Commands for Robotic Navigation and Mobile Manipulation". Paper
  • Navigation from Observation: "Learning to Interpret Natural Language Navigation Instructions from Observations". Paper
  • Grounded Semantic Parsing: "Weakly Supervised Learning of Semantic Parsers for Mapping Instructions to Actions". Paper
  • Behavior Trees: "Behavior Trees in Robotics and AI: An Introduction". arXiv
  • Code as Policies: "Code as Policies: Language Model Programs for Embodied Control". arXiv Website
  • ChatGPT for Robotics: "ChatGPT for Robotics: Design Principles and Model Abilities". Paper Code
  • VoxPoser: "VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models". arXiv Website
  • ReKep: "ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation". arXiv Website
  • VLAs-as-Tools: "Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models". arXiv
  • Project Fetch: Phase Two: "Project Fetch: Phase two". Website

Adjacent Agent Interfaces

  • MRKL Systems: "MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning". arXiv
  • Toolformer: "Toolformer: Language Models Can Teach Themselves to Use Tools". arXiv
  • MM-REACT: "MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action". arXiv
  • HuggingGPT: "HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face". Paper
  • CodeAct: "Executable Code Actions Elicit Better LLM Agents". Paper
  • SWE-agent: "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering". arXiv
  • Model Context Protocol: "Model Context Protocol". Website

Monitoring, Verification, and Recovery

Mechanisms for tracking progress, verifying outcomes, and choosing recovery or termination.

Embodied Systems and Robotics Foundations

  • IPEM: "Integrating Planning, Execution and Monitoring". Paper
  • Inner Monologue: "Inner Monologue: Embodied Reasoning through Planning with Language Models". arXiv Website
  • REFLECT: "REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction". arXiv
  • DoReMi: "DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment". arXiv Website
  • BrainBody-LLM: "Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback". arXiv
  • VASO: "VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents". arXiv

Adjacent Verification and Feedback

  • Training Verifiers: "Training Verifiers to Solve Math Word Problems". arXiv
  • ReAct: "ReAct: Synergizing Reasoning and Acting in Language Models". OpenReview
  • Let's Verify Step by Step: "Let's Verify Step by Step". arXiv
  • LLM Critics: "LLM Critics Help Catch LLM Bugs". Paper

Skill Learning and Lifelong Adaptation

Methods for acquiring and reusing skills across tasks, environments, and embodiments.

  • Lifelong Robot Learning: "Lifelong Robot Learning". Paper
  • Portable Options: "Building Portable Options: Skill Transfer in Reinforcement Learning". Paper
  • Robot Learning from Demonstration Survey: "A Survey of Robot Learning from Demonstration". Paper
  • Neural Task Programming: "Neural Task Programming: Learning to Generalize Across Hierarchical Tasks". arXiv
  • Voyager: "Voyager: An Open-Ended Embodied Agent with Large Language Models". arXiv Website
  • RoboCat: "RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation". arXiv
  • LRLL: "Lifelong Robot Library Learning: Bootstrapping Composable and Generalizable Skills for Embodied Control with Language Models". arXiv Website
  • ASPIRE: "ASPIRE: Agentic /Skills Discovery for Robotics". arXiv Website Code

Autonomous System and Policy Improvement

Feedback-driven loops that improve policies and their supporting data, tools, or runtimes.

Embodied Data, Policy, and Harness Loops

  • GenSim: "GenSim: Generating Robotic Simulation Tasks via Large Language Models". arXiv
  • Eureka: "Eureka: Human-Level Reward Design via Coding Large Language Models". arXiv Website
  • RoboGen: "RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation". arXiv
  • AutoRT: "AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents". arXiv
  • DrEureka: "DrEureka: Language Model Guided Sim-To-Real Transfer". arXiv Website
  • π*₀.₆ / RECAP: "$π^{*}_{0.6}$: a VLA That Learns From Experience". arXiv Website
  • RoboClaw: "RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks". arXiv
  • HARBOR: "HARBOR: A Harness Framework for Agentic Robot Reinforcement Learning". arXiv
  • RHO: "RHO: Your Coding Agent is Secretly a Roboticist". arXiv Website Code
  • RATs: "Playful Agentic Robot Learning". arXiv Website Code
  • ENPIRE: "ENPIRE: Agentic Robot Policy Self-Improvement in the Real World". arXiv Website
  • GaP: "GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks". arXiv Website Code
  • SHAPER: "Self-Evolving Embodied Agents via Skill-Harness Evolution". arXiv
  • Zetta ζ: "Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence". arXiv Website
  • Agentic Push-T: "Revisiting the “Push-T” Robot Manipulation Task with Agentic Robotics". arXiv

Adjacent Autonomous-Agent Precedents

  • Self-Refine: "Self-Refine: Iterative Refinement with Self-Feedback". Paper
  • Coscientist: "Autonomous Chemical Research with Large Language Models". Paper
  • ChemCrow: "ChemCrow: Augmenting large-language models with chemistry tools". arXiv Paper
  • DSPy: "DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines". arXiv
  • A-Lab: "An autonomous laboratory for the accelerated synthesis of inorganic materials". Paper

Multi-Agent, Multi-Robot, and Fleet Coordination

Embodied and Multi-Robot Systems

  • ALLIANCE: "ALLIANCE: An Architecture for Fault Tolerant Multirobot Cooperation". Paper
  • CoELA: "Building Cooperative Embodied Agents Modularly with Large Language Models". arXiv Website Code
  • RoCo: "RoCo: Dialectic Multi-Robot Collaboration with Large Language Models". arXiv Website Code
  • SMART-LLM: "SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models". arXiv Website
  • COHERENT: "COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models". arXiv Code
  • EMOS: "EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents". arXiv
  • HMCF: "HMCF: A Human-in-the-loop Multi-Robot Collaboration Framework Based on Large Language Models". arXiv
  • RAI: "RAI: Flexible Agent Framework for Embodied AI". arXiv
  • Embodied World-Model Alignment: "Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue". arXiv
  • LLawCo: "LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior". arXiv

Adjacent Multi-Agent References

  • CAMEL: "CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society". arXiv
  • ChatDev: "ChatDev: Communicative Agents for Software Development". arXiv Paper
  • MetaGPT: "MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework". OpenReview
  • AutoGen: "AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation". arXiv
  • Agent2Agent Protocol: "Agent2Agent Protocol". Website

Environments and Evaluation

Resources are grouped by the system properties they expose.

Platforms and Task Suites

  • AI2-THOR: "AI2-THOR: An Interactive 3D Environment for Visual AI". arXiv Website
  • VirtualHome: "VirtualHome: Simulating Household Activities via Programs". arXiv
  • Habitat: "Habitat: A Platform for Embodied AI Research". arXiv Website
  • RLBench: "RLBench: The Robot Learning Benchmark & Learning Environment". arXiv
  • ManiSkill: "ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations". arXiv
  • BEHAVIOR: "BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments". arXiv Website
  • Habitat 3.0: "Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots". arXiv
  • BEHAVIOR-1K / OmniGibson: "BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation". arXiv Website
  • RoboCasa: "RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots". arXiv Website
  • RoboVerse / MetaSim: "RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning". arXiv Website Code
  • RoboTwin 2.0: "RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation". arXiv Website Plus Code
  • MolmoSpaces: "MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation". arXiv Website Code
  • RoboCasa365: "RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots". arXiv

Evaluation Infrastructure and Real-World Arenas

Infrastructure for evaluating complete policies and agent systems across tasks, embodiments, and physical sites.

  • SIMPLER: "Evaluating Real-World Robot Manipulation Policies in Simulation". arXiv Website Code
  • RoboArena: "RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies". arXiv Website Code
  • RoboChallenge / Table30: "RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies". arXiv Website
  • Isaac Lab-Arena: "NVIDIA Isaac Lab-Arena". Website Code
  • vla-eval: "vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models". arXiv Leaderboard Code
  • CaP-X: "CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation". arXiv Website Code
  • ManipArena: "ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation". arXiv Code
  • SimFoundry: "SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation". arXiv Website Code
  • RoboDojo: "RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies". arXiv Website Code
  • Embody / Claude Plays Robotics: "How Claude performs on robotics tasks". Website

Long-Horizon, Language, Memory, and Collaboration

VoLo is listed under Architectures and Harnesses for its primary system contribution.

  • ALFRED: "ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks". arXiv
  • ALFWorld: "ALFWorld: Aligning Text and Embodied Environments for Interactive Learning". arXiv
  • TEACh: "TEACh: Task-driven Embodied Agents that Chat". arXiv
  • CALVIN: "CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks". arXiv
  • MineDojo: "MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge". arXiv Website
  • LIBERO: "LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning". arXiv Website
  • PARTNR: "PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks". arXiv
  • VLABench: "VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks". arXiv
  • MIKASA-Robo: "Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning". arXiv Website Code
  • RoboCerebra: "RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation". arXiv Website
  • LongBench: "LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks". arXiv
  • RoboMemArena: "RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark". arXiv Website Code
  • ESI-Bench: "ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop". arXiv
  • RoboGraph / Task-State Horizons: "Compiling and Benchmarking Task-State Horizons for Embodied Agents". arXiv

Robustness and System Evaluation

Evaluations of system behavior under distribution shift, interface changes, and cross-site variation.

Embodied System Evaluation

  • Embodied Agent Interface: "Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making". arXiv
  • EmbodiedBench: "EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents". arXiv Website
  • LIBERO-PRO: "LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization". arXiv Code
  • LIBERO-Plus: "LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models". arXiv Code
  • RobotArena ∞: "RobotArena $∞$: Scalable Robot Benchmarking via Real-to-Sim Translation". arXiv Website Code
  • VLA-Arena: "VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models". arXiv Website Code
  • SafeVLA-Bench: "SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models". arXiv Website
  • ASIMOV-Agentic: "Asimov Agentic Safety Evaluation". Dataset

Adjacent Digital-Agent Evaluation

  • WebArena: "WebArena: A Realistic Web Environment for Building Autonomous Agents". OpenReview
  • AgentBench: "AgentBench: Evaluating LLMs as Agents". OpenReview
  • OSWorld: "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments". OpenReview

Challenges and Open Questions

Embodied agents require reliable state, execution, verification, and recovery under uncertain physical interaction. The table summarizes the main harness-level challenges.

ChallengeWhy it matters for an embodied harnessUseful research direction
State and partial observationPhysical state is partial, noisy, and viewpoint-dependent.Persistent scene/world models, active perception, uncertainty-aware context
Action interfacesNatural-language plans omit safety and dynamic feasibility.Typed tools, skill contracts, preconditions/effects, capability discovery
Verification and terminationController completion provides weak evidence of physical goal success.Independent success detectors, progress monitors, diagnostic exit states
Timing and latencyReasoning and communication add latency as the world changes.Hierarchical rates, asynchronous execution, latency-aware planning, safe interruption
Reset and repeatabilityPhysical resets are costly, imperfect, and sometimes impossible.Auto-reset, reset verification, counterbalancing, simulation and digital twins
Safety and permissionsActions can injure people, damage hardware, or create irreversible states.Permission boundaries, runtime shields, human escalation, rollback-aware planning
Memory and contextRaw multimodal history quickly exceeds context and obscures causal evidence.Event abstraction, trace indexing, multimodal episodic/semantic memory
Regression-safe learningPolicy and skill updates can regress established behaviors.Held-out regression suites, versioned skills, gated promotion, reproducible artifacts
Embodiment variationTools and policies expose different kinematics, sensors, timing, and failure modes.Semantic capability descriptions, adapters, cross-embodiment skill transfer
EvaluationFinal success alone hides retries, unsafe actions, resets, human labor, and resource use.Process metrics, perturbation suites, audit trails, real-world distributed evaluation
Fleet orchestrationMore robots and agents add contention, communication cost, and inconsistent state.Resource-aware scheduling, shared evidence, branch comparison, fault tolerance
Foundation-model limitsModel scale leaves grounding, control, and verification as separate system requirements.Joint model–harness scaling studies and failure-aware system design
  • Awesome World Models: "Awesome World Models". Code
  • Awesome Robotics Foundation Models: "Awesome Robotics Foundation Models". Code
  • Awesome LLM Robotics: "Awesome LLM Robotics". Code
  • Embodied AI Paper List: "Embodied AI Paper List". Code
  • NVIDIA GEAR: "Generalist Embodied Agent Research". Website

Contributing

Contributions are welcome. Please read CONTRIBUTING.md before opening a pull request. A good addition verifies the official title and links, checks aliases and duplicates, chooses one primary section, and preserves chronological order.

Acknowledgements

This project takes presentation inspiration from Awesome World Models and builds on the open work of the robotics, embodied AI, machine learning, computer vision, language, agents, and systems communities.

This project was developed with extensive use of GPT-5.6 Sol Ultra and Codex for literature discovery, verification, organization, and editing.

If this list helps your research, consider starring the eventual public repository and contributing missing or corrected work.

Contributors

dongqianyu99

1 commits

dongqianyu99/Awesome-Agentic-Embodided-Systems

A curated research map of agentic embodied systems, covering foundation models, embodied harnesses, in-context adaptation, self-improvement, and evaluation.

4

1 commits

updated Aug 29, 2026

See the code

README

🤖 Awesome Agentic Embodied Systems

Awesome License: BSD-3-Clause PRs Welcome

A curated research map of the models, runtimes, and learning loops behind goal-directed embodied agents.

Last literature sweep: 2026-08-30


News and Updates

  • 2026-08 — Initial release. Published the initial system-oriented research map.
  • Ongoing — Contributions welcome. Missing papers, corrected links, and better placements are welcome through focused issues or pull requests; see CONTRIBUTING.md.

Overview

Motivation: Why Agentic Embodied Systems, Why Now?

A useful way to read recent language and vision-language AI is:

scaling -> foundation models -> agents / harnesses

Empirical scaling laws linked model performance to data, parameters, and compute. Broad pretraining produced adaptable foundation models. As multimodal reasoning, code generation, and tool use improved, research expanded toward reliable multi-step goal completion. ReAct interleaved reasoning with action and observation, and SWE-agent showed that agent interfaces materially shape behavior. A harness supplies runtime interfaces, persistent state, and outcome-driven control around a model. All three stages continue to co-evolve.

Embodied AI can draw on semantic and multimodal priors that are costly to learn from robot trajectories alone. RT-2 transferred Internet-scale vision-language training to robotic control, Open X-Embodiment found positive transfer across embodiments, and π₀.₅ combined heterogeneous robot, semantic, and web data for manipulation in unseen homes. Together, these works provide evidence for reusable cross-task transfer and control within the studied settings.

Physical interaction makes state uncertain, action constrained, and recovery costly. System interfaces and outcome verification therefore become central to long-horizon performance: SayCan grounds plans in executable affordances, and LIBERO-PRO measures degradation under task and environment perturbations.

Reusable models and skills now make harness design a visible research target. Thea, ENPIRE, and ASPIRE study closed-loop state management, recovery, and improvement in embodied systems. The central question is how model and runtime capabilities should scale together in grounded interaction.

Scope and System Map

This repository maps agentic embodied research across enabling models, system design, and evaluation. It prioritizes verified primary links, concise synthesis, and one primary section per work.

Papers are the default evidence unit. Technically substantive non-paper sources are included when they materially shape the field; source types and company reports are labeled explicitly.

Agentic embodied systems organize embodied models and policies through a runtime that supports goal-directed interaction, feedback, and improvement.

The scope covers physical and simulated environments, within and across episodes.

The collection moves from scaling evidence and reusable models through the agentic system stack to evaluation. Historical robotics and adjacent software-agent precedents appear within the corresponding module. Entries follow earliest public release, use one primary home, and cross-reference other relationships.

Tutorials, Surveys, and Starter Resources

  • Embodied AI Survey: "A Survey of Embodied AI: From Simulators to Research Tasks". arXiv
  • Large Language Models for Robotics: "Large Language Models for Robotics: A Survey". arXiv
  • Foundation Models in Robotics: "Foundation Models in Robotics: Applications, Challenges, and the Future". arXiv Code
  • LLMs for Robotics: Opportunities and Challenges: "Large Language Models for Robotics: Opportunities, Challenges, and Perspectives". arXiv
  • Robotics with Foundation Models: "A Survey on Robotics with Foundation Models: toward Embodied AI". arXiv
  • VLA Survey: "A Survey on Vision-Language-Action Models for Embodied AI". arXiv
  • Building Effective Agents: "Building effective agents". Website
  • Harness Engineering with Codex: "Harness engineering: leveraging Codex in an agent-first world". Website
  • Harness Engineering for Physical AI: "Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer". arXiv
  • Scaling Laws, Carefully: "Scaling Laws, Carefully". Website
  • Embodied Collective Intelligence: "When Multi-Robot Systems Meet Agentic AI: Towards Embodied Collective Intelligence". arXiv
  • Harness Engineering for Self-Improvement: "Harness Engineering for Self-Improvement". Website

Scaling Laws and Scaling Evidence

These works connect scaling research in language, vision-language, and embodied AI.

Language and Vision-Language Models

  • Kaplan Scaling Laws: "Scaling Laws for Neural Language Models". arXiv
  • Chinchilla: "Training Compute-Optimal Large Language Models". arXiv
  • OpenCLIP Scaling: "Reproducible scaling laws for contrastive language-image learning". arXiv Code
  • Zero-Shot Data Limit: "No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance". arXiv

Embodied AI Scaling Evidence

Embodied scaling evidence remains less mature than language-model scaling laws. Proprietary company reports are labeled to reflect their limited reproducibility.

  • Manipulation Data Scaling: "Data Scaling Laws in Imitation Learning for Robotic Manipulation". arXiv Website
  • Agents and World Models: "Scaling Laws for Pre-training Agents and World Models". arXiv
  • GEN-0: "GEN-0 / Embodied Foundation Models That Scale with Physical Interaction". Website
  • EgoScale: "EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data". arXiv Website
  • Precision Scaling: "The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation". arXiv
  • Dyna-2: "Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models". Website

Foundation Models, Policies, and Data

Foundation models, generalist policies, and shared datasets supply reusable capabilities for agentic runtimes.

Generalist Embodied Models and Policies

This section selects models and policies that materially advance reusable embodied capability or its integration into agentic systems. Checkpoint updates are grouped by model family.

  • CLIPort: "CLIPort: What and Where Pathways for Robotic Manipulation". arXiv Website
  • BC-Z: "BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning". arXiv
  • Gato: "A Generalist Agent". arXiv
  • VIMA: "VIMA: General Robot Manipulation with Multimodal Prompts". arXiv Website
  • RT-1: "RT-1: Robotics Transformer for Real-World Control at Scale". arXiv
  • PaLM-E: "PaLM-E: An Embodied Multimodal Language Model". arXiv
  • RT-2: "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control". arXiv
  • Open X-Embodiment / RT-X: "Open X-Embodiment: Robotic Learning Datasets and RT-X Models". arXiv Website
  • RT-H: "RT-H: Action Hierarchies Using Language". arXiv Website
  • Octo: "Octo: An Open-Source Generalist Robot Policy". arXiv Website
  • OpenVLA: "OpenVLA: An Open-Source Vision-Language-Action Model". arXiv Website
  • RDT-1B: "RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation". arXiv
  • π₀: "$π_0$: A Vision-Language-Action Flow Model for General Robot Control". arXiv Website
  • Magma: "Magma: A Foundation Model for Multimodal AI Agents". arXiv Website Code
  • Helix family: "Helix: A Vision-Language-Action Model for Generalist Humanoid Control". Helix Helix 02
  • GR00T N family: "GR00T N1: An Open Foundation Model for Generalist Humanoid Robots". arXiv N1.6 N1.7 Code
  • Gemini Robotics: "Gemini Robotics: Bringing AI into the Physical World". arXiv Website
  • π₀.₅: "$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization". arXiv
  • SmolVLA: "SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics". arXiv
  • TRI LBM 1.0: "A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation". arXiv Website
  • GR-3: "GR-3 Technical Report". arXiv Website
  • Gemini Robotics 1.5: "Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer". arXiv
  • X-VLA: "X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model". arXiv Website Code
  • LingBot-VLA: "A Pragmatic VLA Foundation Model". arXiv
  • DreamZero: "World Action Models are Zero-shot Policies". arXiv Website Code
  • GEN-1: "GEN-1: Scaling Embodied Foundation Models to Mastery". Website
  • π₀.₇: "$π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities". arXiv Website
  • MolmoAct 2: "MolmoAct2: Action Reasoning Models for Real-world Deployment". arXiv Website Code
  • Embodied-R1.5: "Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models". arXiv Website Code
  • Vesta: "Vesta: A Generalist Embodied Reasoning Model". arXiv Website
  • LingBot-VLA 2.0: "From Foundation to Application: Improving VLA Models in Practice". arXiv Website Code
  • Xiaomi-Robotics-1: "Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories". arXiv Website Code
  • Gemini Robotics 2: "Gemini Robotics 2 brings whole body intelligence to robots". Website

Promptable and In-Context Robot Policies

Policies that adapt at inference time through demonstrations, interaction history, or persistent context.

  • ICRT: "In-Context Imitation Learning via Next-Token Prediction". arXiv Website
  • Instant Policy: "Instant Policy: In-Context Imitation Learning via Graph Diffusion". arXiv Website
  • LocoFormer: "LocoFormer: Generalist Locomotion via Long-context Adaptation". arXiv Website
  • BPP: "Behavior Prompting Policy: Demonstrations as Prompts for Manipulation". arXiv Website
  • RoboTTT: "RoboTTT: Context Scaling for Robot Policies". arXiv Website
  • Skild S1: "Introducing S1: In-Context Learning for Robotics". Website
  • GEN-1.5: "GEN-1.5: Embodied Foundation Models are One-Shot Learners". Website

Data and Open Infrastructure

  • RoboNet: "RoboNet: Large-Scale Multi-Robot Learning". arXiv Website
  • BridgeData V2: "BridgeData V2: A Dataset for Robot Learning at Scale". arXiv Website
  • DROID: "DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset". arXiv Website
  • LeRobot: "LeRobot". Website Code
  • OpenVLA-OFT: "Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success". arXiv Website
  • AgiBot World / GO-1: "AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems". arXiv Website Code
  • DreamGen: "DreamGen: Unlocking Generalization in Robot Learning through Video World Models". arXiv Website Code

Agentic Embodied System Stack

The primary classification axis groups each work by its main system responsibility, with embodied systems foregrounded and relevant precedents placed alongside them.

Architectures and Harnesses

Architectures, middleware, and runtimes that span multiple stages of the interaction loop.

  • Subsumption Architecture: "A Robust Layered Control System for a Mobile Robot". Paper
  • ATLANTIS: "Integrating Planning and Reacting in a Heterogeneous Asynchronous Architecture for Controlling Real-World Mobile Robots". Paper
  • 3T: "Experiences with an Architecture for Intelligent, Reactive Agents". Paper
  • Remote Agent: "Remote Agent: To Boldly Go Where No AI System Has Gone Before". Paper
  • CRAM: "CRAM—A Cognitive Robot Abstract Machine for Everyday Manipulation in Human Environments". Paper Website
  • OK-Robot: "OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics". arXiv Website
  • ROS-LLM: "ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning". arXiv Code
  • Being-0: "Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills". arXiv Website
  • RoboNeuron: "RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI". arXiv Code
  • VoLo / RoboVoLo: "VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation". arXiv Website Code
  • Harness VLA: "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents". arXiv
  • OpenETA: "ETA: A New Agentic Paradigm for Embodied Tasks". arXiv Website Code
  • Thea: "Towards the Harness of Embodied Agents". arXiv Website Code

State, Context, and Memory

Methods for constructing world state and retaining knowledge across steps or episodes.

Embodied Systems and Robotics Foundations

  • KnowRob: "KnowRob—Knowledge Processing for Autonomous Personal Robots". Website
  • RoboEarth: "RoboEarth—A World Wide Web for Robots". Website
  • RoboBrain: "RoboBrain: Large-Scale Knowledge Engine for Robots". arXiv
  • Statler: "Statler: State-Maintaining Language Models for Embodied Reasoning". arXiv
  • HELPER: "Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language Models". arXiv Website
  • HELPER-X: "HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models". arXiv
  • KARMA: "KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems". arXiv
  • ViReSkill: "ViReSkill: Vision-Grounded Replanning with Skill Memory for LLM-Based Planning in Lifelong Robot Learning". arXiv
  • MEM: "MEM: Multi-Scale Embodied Memory for Vision Language Action Models". arXiv Website

Adjacent Memory Architectures

  • Reflexion: "Reflexion: Language Agents with Verbal Reinforcement Learning". OpenReview
  • Generative Agents: "Generative Agents: Interactive Simulacra of Human Behavior". arXiv
  • ExpeL: "ExpeL: LLM Agents Are Experiential Learners". arXiv Paper
  • MemGPT: "MemGPT: Towards LLMs as Operating Systems". arXiv

Planning and Orchestration

Goal decomposition and high-level coordination of plans, actions, and skills.

  • STRIPS: "STRIPS: A New Approach to the Application of Theorem Proving to Problem Solving". Paper
  • Generalized Robot Plans: "Learning and Executing Generalized Robot Plans". Paper
  • PRS: "Reactive Reasoning and Planning". Paper
  • PDDL: "PDDL—The Planning Domain Definition Language". Paper
  • HTAMP: "Hierarchical Task and Motion Planning in the Now". Paper
  • ROSPlan: "ROSPlan: Planning in the Robot Operating System". Paper Code
  • Logic-Geometric Programming: "Logic-Geometric Programming: An Optimization-Based Approach to Combined Task and Motion Planning". Paper
  • PDDLStream: "PDDLStream: Integrating Symbolic Planners and Blackbox Samplers via Optimistic Adaptive Planning". arXiv
  • Zero-Shot Planner: "Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents". arXiv
  • SayCan: "Do As I Can, Not As I Say: Grounding Language in Robotic Affordances". arXiv Website
  • ProgPrompt: "ProgPrompt: Generating Situated Robot Task Plans using Large Language Models". arXiv
  • LLM-Planner: "LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models". arXiv
  • ASC: "ASC: Adaptive Skill Coordination for Robotic Mobile Manipulation". arXiv
  • LLM+P: "LLM+P: Empowering Large Language Models with Optimal Planning Proficiency". arXiv Code
  • SayPlan: "SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning". arXiv Website
  • Reflective Planning: "Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation". arXiv Website
  • Hi Robot: "Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models". arXiv Website
  • VLA-Reasoner: "VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search". arXiv
  • Goal2Skill: "Goal2Skill: Long-Horizon Manipulation with Adaptive Planning and Reflection". arXiv
  • τ₀-VLA: "$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation". arXiv

Tools, Programs, and Action Grounding

Representations and interfaces that turn selected capabilities into executable actions.

Embodied Systems and Robotics Foundations

  • RAPs: "An Investigation into Reactive Planning in Complex Domains". Paper
  • Teleo-Reactive Programs: "Teleo-Reactive Programs for Agent Control". arXiv
  • Options: "Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning". Paper
  • Walk the Talk / MARCO: "Walk the Talk: Connecting Language, Knowledge, and Action in Route Instructions". Paper
  • : "Understanding Natural Language Commands for Robotic Navigation and Mobile Manipulation". Paper
  • Navigation from Observation: "Learning to Interpret Natural Language Navigation Instructions from Observations". Paper
  • Grounded Semantic Parsing: "Weakly Supervised Learning of Semantic Parsers for Mapping Instructions to Actions". Paper
  • Behavior Trees: "Behavior Trees in Robotics and AI: An Introduction". arXiv
  • Code as Policies: "Code as Policies: Language Model Programs for Embodied Control". arXiv Website
  • ChatGPT for Robotics: "ChatGPT for Robotics: Design Principles and Model Abilities". Paper Code
  • VoxPoser: "VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models". arXiv Website
  • ReKep: "ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation". arXiv Website
  • VLAs-as-Tools: "Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models". arXiv
  • Project Fetch: Phase Two: "Project Fetch: Phase two". Website

Adjacent Agent Interfaces

  • MRKL Systems: "MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning". arXiv
  • Toolformer: "Toolformer: Language Models Can Teach Themselves to Use Tools". arXiv
  • MM-REACT: "MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action". arXiv
  • HuggingGPT: "HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face". Paper
  • CodeAct: "Executable Code Actions Elicit Better LLM Agents". Paper
  • SWE-agent: "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering". arXiv
  • Model Context Protocol: "Model Context Protocol". Website

Monitoring, Verification, and Recovery

Mechanisms for tracking progress, verifying outcomes, and choosing recovery or termination.

Embodied Systems and Robotics Foundations

  • IPEM: "Integrating Planning, Execution and Monitoring". Paper
  • Inner Monologue: "Inner Monologue: Embodied Reasoning through Planning with Language Models". arXiv Website
  • REFLECT: "REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction". arXiv
  • DoReMi: "DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment". arXiv Website
  • BrainBody-LLM: "Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback". arXiv
  • VASO: "VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents". arXiv

Adjacent Verification and Feedback

  • Training Verifiers: "Training Verifiers to Solve Math Word Problems". arXiv
  • ReAct: "ReAct: Synergizing Reasoning and Acting in Language Models". OpenReview
  • Let's Verify Step by Step: "Let's Verify Step by Step". arXiv
  • LLM Critics: "LLM Critics Help Catch LLM Bugs". Paper

Skill Learning and Lifelong Adaptation

Methods for acquiring and reusing skills across tasks, environments, and embodiments.

  • Lifelong Robot Learning: "Lifelong Robot Learning". Paper
  • Portable Options: "Building Portable Options: Skill Transfer in Reinforcement Learning". Paper
  • Robot Learning from Demonstration Survey: "A Survey of Robot Learning from Demonstration". Paper
  • Neural Task Programming: "Neural Task Programming: Learning to Generalize Across Hierarchical Tasks". arXiv
  • Voyager: "Voyager: An Open-Ended Embodied Agent with Large Language Models". arXiv Website
  • RoboCat: "RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation". arXiv
  • LRLL: "Lifelong Robot Library Learning: Bootstrapping Composable and Generalizable Skills for Embodied Control with Language Models". arXiv Website
  • ASPIRE: "ASPIRE: Agentic /Skills Discovery for Robotics". arXiv Website Code

Autonomous System and Policy Improvement

Feedback-driven loops that improve policies and their supporting data, tools, or runtimes.

Embodied Data, Policy, and Harness Loops

  • GenSim: "GenSim: Generating Robotic Simulation Tasks via Large Language Models". arXiv
  • Eureka: "Eureka: Human-Level Reward Design via Coding Large Language Models". arXiv Website
  • RoboGen: "RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation". arXiv
  • AutoRT: "AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents". arXiv
  • DrEureka: "DrEureka: Language Model Guided Sim-To-Real Transfer". arXiv Website
  • π*₀.₆ / RECAP: "$π^{*}_{0.6}$: a VLA That Learns From Experience". arXiv Website
  • RoboClaw: "RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks". arXiv
  • HARBOR: "HARBOR: A Harness Framework for Agentic Robot Reinforcement Learning". arXiv
  • RHO: "RHO: Your Coding Agent is Secretly a Roboticist". arXiv Website Code
  • RATs: "Playful Agentic Robot Learning". arXiv Website Code
  • ENPIRE: "ENPIRE: Agentic Robot Policy Self-Improvement in the Real World". arXiv Website
  • GaP: "GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks". arXiv Website Code
  • SHAPER: "Self-Evolving Embodied Agents via Skill-Harness Evolution". arXiv
  • Zetta ζ: "Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence". arXiv Website
  • Agentic Push-T: "Revisiting the “Push-T” Robot Manipulation Task with Agentic Robotics". arXiv

Adjacent Autonomous-Agent Precedents

  • Self-Refine: "Self-Refine: Iterative Refinement with Self-Feedback". Paper
  • Coscientist: "Autonomous Chemical Research with Large Language Models". Paper
  • ChemCrow: "ChemCrow: Augmenting large-language models with chemistry tools". arXiv Paper
  • DSPy: "DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines". arXiv
  • A-Lab: "An autonomous laboratory for the accelerated synthesis of inorganic materials". Paper

Multi-Agent, Multi-Robot, and Fleet Coordination

Embodied and Multi-Robot Systems

  • ALLIANCE: "ALLIANCE: An Architecture for Fault Tolerant Multirobot Cooperation". Paper
  • CoELA: "Building Cooperative Embodied Agents Modularly with Large Language Models". arXiv Website Code
  • RoCo: "RoCo: Dialectic Multi-Robot Collaboration with Large Language Models". arXiv Website Code
  • SMART-LLM: "SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models". arXiv Website
  • COHERENT: "COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models". arXiv Code
  • EMOS: "EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents". arXiv
  • HMCF: "HMCF: A Human-in-the-loop Multi-Robot Collaboration Framework Based on Large Language Models". arXiv
  • RAI: "RAI: Flexible Agent Framework for Embodied AI". arXiv
  • Embodied World-Model Alignment: "Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue". arXiv
  • LLawCo: "LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior". arXiv

Adjacent Multi-Agent References

  • CAMEL: "CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society". arXiv
  • ChatDev: "ChatDev: Communicative Agents for Software Development". arXiv Paper
  • MetaGPT: "MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework". OpenReview
  • AutoGen: "AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation". arXiv
  • Agent2Agent Protocol: "Agent2Agent Protocol". Website

Environments and Evaluation

Resources are grouped by the system properties they expose.

Platforms and Task Suites

  • AI2-THOR: "AI2-THOR: An Interactive 3D Environment for Visual AI". arXiv Website
  • VirtualHome: "VirtualHome: Simulating Household Activities via Programs". arXiv
  • Habitat: "Habitat: A Platform for Embodied AI Research". arXiv Website
  • RLBench: "RLBench: The Robot Learning Benchmark & Learning Environment". arXiv
  • ManiSkill: "ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations". arXiv
  • BEHAVIOR: "BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments". arXiv Website
  • Habitat 3.0: "Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots". arXiv
  • BEHAVIOR-1K / OmniGibson: "BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation". arXiv Website
  • RoboCasa: "RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots". arXiv Website
  • RoboVerse / MetaSim: "RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning". arXiv Website Code
  • RoboTwin 2.0: "RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation". arXiv Website Plus Code
  • MolmoSpaces: "MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation". arXiv Website Code
  • RoboCasa365: "RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots". arXiv

Evaluation Infrastructure and Real-World Arenas

Infrastructure for evaluating complete policies and agent systems across tasks, embodiments, and physical sites.

  • SIMPLER: "Evaluating Real-World Robot Manipulation Policies in Simulation". arXiv Website Code
  • RoboArena: "RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies". arXiv Website Code
  • RoboChallenge / Table30: "RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies". arXiv Website
  • Isaac Lab-Arena: "NVIDIA Isaac Lab-Arena". Website Code
  • vla-eval: "vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models". arXiv Leaderboard Code
  • CaP-X: "CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation". arXiv Website Code
  • ManipArena: "ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation". arXiv Code
  • SimFoundry: "SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation". arXiv Website Code
  • RoboDojo: "RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies". arXiv Website Code
  • Embody / Claude Plays Robotics: "How Claude performs on robotics tasks". Website

Long-Horizon, Language, Memory, and Collaboration

VoLo is listed under Architectures and Harnesses for its primary system contribution.

  • ALFRED: "ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks". arXiv
  • ALFWorld: "ALFWorld: Aligning Text and Embodied Environments for Interactive Learning". arXiv
  • TEACh: "TEACh: Task-driven Embodied Agents that Chat". arXiv
  • CALVIN: "CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks". arXiv
  • MineDojo: "MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge". arXiv Website
  • LIBERO: "LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning". arXiv Website
  • PARTNR: "PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks". arXiv
  • VLABench: "VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks". arXiv
  • MIKASA-Robo: "Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning". arXiv Website Code
  • RoboCerebra: "RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation". arXiv Website
  • LongBench: "LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks". arXiv
  • RoboMemArena: "RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark". arXiv Website Code
  • ESI-Bench: "ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop". arXiv
  • RoboGraph / Task-State Horizons: "Compiling and Benchmarking Task-State Horizons for Embodied Agents". arXiv

Robustness and System Evaluation

Evaluations of system behavior under distribution shift, interface changes, and cross-site variation.

Embodied System Evaluation

  • Embodied Agent Interface: "Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making". arXiv
  • EmbodiedBench: "EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents". arXiv Website
  • LIBERO-PRO: "LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization". arXiv Code
  • LIBERO-Plus: "LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models". arXiv Code
  • RobotArena ∞: "RobotArena $∞$: Scalable Robot Benchmarking via Real-to-Sim Translation". arXiv Website Code
  • VLA-Arena: "VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models". arXiv Website Code
  • SafeVLA-Bench: "SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models". arXiv Website
  • ASIMOV-Agentic: "Asimov Agentic Safety Evaluation". Dataset

Adjacent Digital-Agent Evaluation

  • WebArena: "WebArena: A Realistic Web Environment for Building Autonomous Agents". OpenReview
  • AgentBench: "AgentBench: Evaluating LLMs as Agents". OpenReview
  • OSWorld: "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments". OpenReview

Challenges and Open Questions

Embodied agents require reliable state, execution, verification, and recovery under uncertain physical interaction. The table summarizes the main harness-level challenges.

ChallengeWhy it matters for an embodied harnessUseful research direction
State and partial observationPhysical state is partial, noisy, and viewpoint-dependent.Persistent scene/world models, active perception, uncertainty-aware context
Action interfacesNatural-language plans omit safety and dynamic feasibility.Typed tools, skill contracts, preconditions/effects, capability discovery
Verification and terminationController completion provides weak evidence of physical goal success.Independent success detectors, progress monitors, diagnostic exit states
Timing and latencyReasoning and communication add latency as the world changes.Hierarchical rates, asynchronous execution, latency-aware planning, safe interruption
Reset and repeatabilityPhysical resets are costly, imperfect, and sometimes impossible.Auto-reset, reset verification, counterbalancing, simulation and digital twins
Safety and permissionsActions can injure people, damage hardware, or create irreversible states.Permission boundaries, runtime shields, human escalation, rollback-aware planning
Memory and contextRaw multimodal history quickly exceeds context and obscures causal evidence.Event abstraction, trace indexing, multimodal episodic/semantic memory
Regression-safe learningPolicy and skill updates can regress established behaviors.Held-out regression suites, versioned skills, gated promotion, reproducible artifacts
Embodiment variationTools and policies expose different kinematics, sensors, timing, and failure modes.Semantic capability descriptions, adapters, cross-embodiment skill transfer
EvaluationFinal success alone hides retries, unsafe actions, resets, human labor, and resource use.Process metrics, perturbation suites, audit trails, real-world distributed evaluation
Fleet orchestrationMore robots and agents add contention, communication cost, and inconsistent state.Resource-aware scheduling, shared evidence, branch comparison, fault tolerance
Foundation-model limitsModel scale leaves grounding, control, and verification as separate system requirements.Joint model–harness scaling studies and failure-aware system design
  • Awesome World Models: "Awesome World Models". Code
  • Awesome Robotics Foundation Models: "Awesome Robotics Foundation Models". Code
  • Awesome LLM Robotics: "Awesome LLM Robotics". Code
  • Embodied AI Paper List: "Embodied AI Paper List". Code
  • NVIDIA GEAR: "Generalist Embodied Agent Research". Website

Contributing

Contributions are welcome. Please read CONTRIBUTING.md before opening a pull request. A good addition verifies the official title and links, checks aliases and duplicates, chooses one primary section, and preserves chronological order.

Acknowledgements

This project takes presentation inspiration from Awesome World Models and builds on the open work of the robotics, embodied AI, machine learning, computer vision, language, agents, and systems communities.

This project was developed with extensive use of GPT-5.6 Sol Ultra and Codex for literature discovery, verification, organization, and editing.

If this list helps your research, consider starring the eventual public repository and contributing missing or corrected work.

Contributors

dongqianyu99

1 commits