Orlando-CS/Awesome-VLA

✨✨latest advancements in VLA models(VIsion Language Action)

128

25 commits

updated Feb 27, 2026

See the code

README

Awesome Vision-Language-Action (VLA) Models

Awesome Badge GitHub stars

This is a collection of research papers about Embodied Multimodal Large Language Models (VLA models).

If you would like to include your paper or update any details (e.g., code URLs, conference information), please feel free to submit a pull request. Any advice is also welcome!

📑 Table of Contents

Awesome VLA Models


🔥🔥🔥 ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

📖 Paper | 🌟 Project | 💻 Code

Largest unified open-source robotic manipulation dataset (6M+ trajectories) with Action Manifold Learning (AML) for direct clean action prediction.
Supports modular 3D perception for plug-and-play spatial enhancement and improved precision in complex tasks. ✨

🔥🔥🔥 RT-2: Robotics Transformer 2 — End-to-End Vision-Language-Action Model

📖 Paper | 🌟 Project

RT-2 Image

Integrates vision-language models trained on internet-scale data directly into robotic control pipelines. ✨


🔥🔥🔥 Helix: Generalist VLA Model for Full-Body Humanoid Control

🌟 Project

Helix Image

First VLA model achieving full upper-body humanoid control including fingers, wrists, torso, and head. ✨


🔥🔥🔥 π0 (Pi-Zero): Generalist VLA Across Diverse Robots

🌟 Project

Generalist control across various robot embodiments, utilizing large-scale pretraining and flow matching action generation. ✨


🔥🔥🔥 OpenVLA: Open-Source Large-Scale Vision-Language-Action Model

📖 Paper | 🌟 Project | 🤖 Hugging Face

Pretrained on 970k+ robotic episodes, setting a new benchmark for generalist robotic policies. ✨


🔥🔥🔥 Gemini Robotics: Multimodal Generalization to Physical Action

🌟 Project

Gemini Robotics Image

Built on Gemini 2.0, enabling complex real-world manipulation without task-specific training. ✨


Awesome Papers

TitleIntroductionDateCode
PaLM-E: An Embodied Multimodal Language ModelIntegrates perception, language, and action for embodied AI.2023-03-06-
EmbodiedGPTVision-language models with embodied CoT reasoning.2023-05-24Github
Co-LLM-AgentsCooperative embodied agents via modular LLMs.2023-07-05Github
RT-2Transfers VLM internet knowledge to robotic control.2023-07-28-
LLM as PoliciesApple: LLMs for embodied tasks as policies.2023-10-26Github
Embodied Generalist Agent 3DGeneralist agent in 3D worlds.2023-11-18Github
LL3DAOmni-3D understanding via instruction tuning.2023-11-30Github
NaviLLMGeneralist navigation models.2023-12-04Github
MP5Open-ended embodied agent in Minecraft.2023-12-12Github
ManipLLMObject-centric robotic manipulation via LLMs.2023-12-24Github
MultiPLYMultisensory 3D embodied LLMs.2024-01-16Github
NaVidNext-step planning in navigation.2024-02-24-
ShapeLLM3D object understanding for embodied agents.2024-02-27Github
3D-VLAGenerative 3D world model for VLA learning.2024-03-14Github
RoboMP²Multimodal robotic perception-planning.2024-04-07-
HelixFull-body humanoid control model.2024-04Project
Embodied CoT DistillationDistilling embodied CoT into agents.2024-05-02-
Gemini RoboticsReal-world manipulation by Gemini.2024-05Project
A3VLMActionable articulation-aware VLMs.2024-06-11Github
OpenVLAOpen-sourced 7B vision-language-action model.2024-06-13Github
TinyVLACompact and efficient VLA models.2024-09Paper
VLA Expert CollaborationImproves VLA via expert actions.2025-03Paper
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold LearningLargest unified 6M+ manipulation dataset with Action Manifold Learning for stable clean action prediction.2026-02Github

Awesome Datasets and Benchmarks

🏗️ Scene / Environment Generation

TitleIntroductionDateCode
HolodeckLLMs generate interactive 3D simulation environments.2023-12-14Github
PhyScenePhysically interactive 3D scenes for embodied training.2024-04-15Github

🧠 Question Answering / Language Interaction Benchmarks

TitleIntroductionDateCode
OpenEQAVisual QA benchmark for real-world scenes.2024-06-17Github
EQA-REALReal-world EmbodiedQA for indoor settings.2024-04Github
TEAChHuman-human embodied task dialogues.2023 update (original 2021)Github

👀 Multi-Modal Perception Datasets

TitleIntroductionDateCode
EmbodiedScanReal-world RGB-D + language 3D scans.2023-12-26Github

🕹️ End-to-End Embodied Decision Making / Simulators

TitleIntroductionDateCode
PCA-EVALDecision-making via GPT-4V evaluation.2023-10-03Github
UniSimInteractive real-world simulator learning.2023-10-09-

🏡 Household Activities / Task Benchmarks

TitleIntroductionDateCode
BEHAVIOR-1K1,000 household activity programs and scenes.2023-07-11Project
large-language-models
large-vision-language-model
multi-modality

Orlando-CS/Awesome-VLA

✨✨latest advancements in VLA models(VIsion Language Action)

128

25 commits

updated Feb 27, 2026

See the code

README

Awesome Vision-Language-Action (VLA) Models

Awesome Badge GitHub stars

This is a collection of research papers about Embodied Multimodal Large Language Models (VLA models).

If you would like to include your paper or update any details (e.g., code URLs, conference information), please feel free to submit a pull request. Any advice is also welcome!

📑 Table of Contents

Awesome VLA Models


🔥🔥🔥 ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

📖 Paper | 🌟 Project | 💻 Code

Largest unified open-source robotic manipulation dataset (6M+ trajectories) with Action Manifold Learning (AML) for direct clean action prediction.
Supports modular 3D perception for plug-and-play spatial enhancement and improved precision in complex tasks. ✨

🔥🔥🔥 RT-2: Robotics Transformer 2 — End-to-End Vision-Language-Action Model

📖 Paper | 🌟 Project

RT-2 Image

Integrates vision-language models trained on internet-scale data directly into robotic control pipelines. ✨


🔥🔥🔥 Helix: Generalist VLA Model for Full-Body Humanoid Control

🌟 Project

Helix Image

First VLA model achieving full upper-body humanoid control including fingers, wrists, torso, and head. ✨


🔥🔥🔥 π0 (Pi-Zero): Generalist VLA Across Diverse Robots

🌟 Project

Generalist control across various robot embodiments, utilizing large-scale pretraining and flow matching action generation. ✨


🔥🔥🔥 OpenVLA: Open-Source Large-Scale Vision-Language-Action Model

📖 Paper | 🌟 Project | 🤖 Hugging Face

Pretrained on 970k+ robotic episodes, setting a new benchmark for generalist robotic policies. ✨


🔥🔥🔥 Gemini Robotics: Multimodal Generalization to Physical Action

🌟 Project

Gemini Robotics Image

Built on Gemini 2.0, enabling complex real-world manipulation without task-specific training. ✨


Awesome Papers

TitleIntroductionDateCode
PaLM-E: An Embodied Multimodal Language ModelIntegrates perception, language, and action for embodied AI.2023-03-06-
EmbodiedGPTVision-language models with embodied CoT reasoning.2023-05-24Github
Co-LLM-AgentsCooperative embodied agents via modular LLMs.2023-07-05Github
RT-2Transfers VLM internet knowledge to robotic control.2023-07-28-
LLM as PoliciesApple: LLMs for embodied tasks as policies.2023-10-26Github
Embodied Generalist Agent 3DGeneralist agent in 3D worlds.2023-11-18Github
LL3DAOmni-3D understanding via instruction tuning.2023-11-30Github
NaviLLMGeneralist navigation models.2023-12-04Github
MP5Open-ended embodied agent in Minecraft.2023-12-12Github
ManipLLMObject-centric robotic manipulation via LLMs.2023-12-24Github
MultiPLYMultisensory 3D embodied LLMs.2024-01-16Github
NaVidNext-step planning in navigation.2024-02-24-
ShapeLLM3D object understanding for embodied agents.2024-02-27Github
3D-VLAGenerative 3D world model for VLA learning.2024-03-14Github
RoboMP²Multimodal robotic perception-planning.2024-04-07-
HelixFull-body humanoid control model.2024-04Project
Embodied CoT DistillationDistilling embodied CoT into agents.2024-05-02-
Gemini RoboticsReal-world manipulation by Gemini.2024-05Project
A3VLMActionable articulation-aware VLMs.2024-06-11Github
OpenVLAOpen-sourced 7B vision-language-action model.2024-06-13Github
TinyVLACompact and efficient VLA models.2024-09Paper
VLA Expert CollaborationImproves VLA via expert actions.2025-03Paper
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold LearningLargest unified 6M+ manipulation dataset with Action Manifold Learning for stable clean action prediction.2026-02Github

Awesome Datasets and Benchmarks

🏗️ Scene / Environment Generation

TitleIntroductionDateCode
HolodeckLLMs generate interactive 3D simulation environments.2023-12-14Github
PhyScenePhysically interactive 3D scenes for embodied training.2024-04-15Github

🧠 Question Answering / Language Interaction Benchmarks

TitleIntroductionDateCode
OpenEQAVisual QA benchmark for real-world scenes.2024-06-17Github
EQA-REALReal-world EmbodiedQA for indoor settings.2024-04Github
TEAChHuman-human embodied task dialogues.2023 update (original 2021)Github

👀 Multi-Modal Perception Datasets

TitleIntroductionDateCode
EmbodiedScanReal-world RGB-D + language 3D scans.2023-12-26Github

🕹️ End-to-End Embodied Decision Making / Simulators

TitleIntroductionDateCode
PCA-EVALDecision-making via GPT-4V evaluation.2023-10-03Github
UniSimInteractive real-world simulator learning.2023-10-09-

🏡 Household Activities / Task Benchmarks

TitleIntroductionDateCode
BEHAVIOR-1K1,000 household activity programs and scenes.2023-07-11Project
large-language-models
large-vision-language-model
multi-modality