HyperbolicCurve/Awesome-World-Action-Model

A curated list of academic papers and resources on Vision-Language-Action (VLA) and World Action Models (WAM)

Python

35

202 commits

updated Oct 1, 2026

See the code

README

πŸ€– Awesome World Action Models Awesome

A curated, continuously-updated reading list of World Action Models (WAM), Vision-Language-Action (VLA) models, and Embodied AI β€” organized by a survey-grounded taxonomy.

Last Update PRs Welcome License: MIT Stars

🌐 Website: hyperboliccurve.github.io/Awesome-World-Action-Model


Overview

The push toward general-purpose robots has produced two converging families of foundation models:

  • Vision-Language-Action (VLA) models inherit the language grounding and visual understanding of pretrained Vision-Language Models (VLMs) and adapt them to emit actions β€” a scalable route to language-conditioned policies.
  • World Action Models (WAM) start from a world model / video backbone that predicts how a scene evolves, and adapt that predictive prior to emit actions β€” trading the "languageβ†’motion" grounding gap for a "dynamicsβ†’action" one.

These two families overlap: a WAM built on a pretrained VLM is simultaneously a VLA and a WAM. This list maps that landscape with a taxonomy grounded in the recent survey literature (see Surveys), so each category has a clear, defensible scope rather than an ad-hoc label.

[!NOTE] Legend β€” πŸ“„ arXiv Β· 🌐 project page Β· πŸ’» code Β· πŸ“Š dataset/benchmark. Tables are sorted newest-first within each category. The πŸ†• Latest Papers section is refreshed daily from arXiv by a GitHub Action; everything else is hand-curated.


Taxonomy at a Glance

flowchart TD
    A[Robot Foundation Models] --> B[Vision-Language-Action<br/>VLA]
    A --> C[World &amp; World-Action Models<br/>WM / WAM]
    A --> R[Action Representations]
    A --> P[Foundational Policies]

    B --> B1[By Action Representation:<br/>Autoregressive Β· Diffusion Β· Flow-Matching]
    B --> B2[By Capability:<br/>Reasoning/Dual-System Β· 3D-4D Β· Efficient Β· RL Fine-Tuning]

    C --> C1[Foundation / General World Models]
    C --> C2[WAM from Video Generation]
    C --> C3[WAM from VLMs]
    C --> C4[WAM from Scratch Β· Latent / JEPA]
    C --> C5[Domain: Driving Β· Navigation]

    R --> R1[Discrete / Autoregressive Tokenizers]
    R --> R2[Diffusion &amp; Flow-Matching Policies]

Table of Contents


πŸ”‘ Key Definitions

TermDefinitionCanonical reference
Vision-Language-Action (VLA)A robot policy that adapts a pretrained VLM to map images + language instructions to actions.RT-2 (Brohan et al., 2023)
World Model (WM)A learned model that predicts future states of an environment (in pixels, latents, or 3D/4D), used for planning, simulation, or representation.World Models (Ha & Schmidhuber, 2018)
World Action Model (WAM)A policy that leverages world-modeling capability (predicting future states) for action prediction β€” typically by adapting a video / world-model backbone to emit actions.GR-1 (Wu et al., 2023)

[!IMPORTANT] VLA ∩ WAM. The families intersect: a WAM built on a pretrained VLM is both. The split in this list is by what prior the model starts from β€” VLM-style vision-language priors (VLA) vs. video/dynamics priors (WAM) β€” and, within VLA, by how actions are represented, the axis most surveys agree is the field's clearest discriminator.


πŸ†• Latest Papers (Auto-updated)

Papers are automatically fetched daily from arXiv. Last updated: 2026-09-30

VLA

World Model

Policy


πŸ“š Surveys

Recent surveys that define the field and motivate the taxonomy used here.

World & Embodied World Models

TitleAuthorsYearLinks
Understanding World or Predicting Future? A Comprehensive Survey of World ModelsDing et al.2024πŸ“„
A Comprehensive Survey on World Models for Embodied AILi et al.2025πŸ“„
3D and 4D World Modeling: A SurveyKong et al.2025πŸ“„
Learning Embodied Intelligence from Physical Simulators and World ModelsLong et al.2025πŸ“„
Embodied AI: From LLMs to World ModelsFeng et al.2025πŸ“„
World Model for Robot Learning: A Comprehensive SurveyHou et al.2026πŸ“„
Modeling the Mental World for Embodied AI: A Comprehensive ReviewLiu et al.2026πŸ“„
The Role of World Models in Shaping Autonomous Driving: A SurveyTu et al.2025πŸ“„

Vision-Language-Action

TitleAuthorsYearLinks
A Survey on Vision-Language-Action Models for Embodied AIMa et al.2024πŸ“„
A Survey on VLA Models: An Action Tokenization PerspectiveZhong et al.2025πŸ“„
VLA Models: Concepts, Progress, Applications and ChallengesSapkota et al.2025πŸ“„
Large VLM-based VLA Models for Robotic Manipulation: A SurveyShao et al.2025πŸ“„
Efficient VLA Models for Embodied Manipulation: A Systematic SurveyGuan et al.2025πŸ“„
VLA Models for Robotics: A Review Towards Real-World ApplicationsKawaharazuka et al.2025πŸ“„ Β· 🌐
An Anatomy of VLA Models: From Modules to Milestones and Challengesβ€”2025πŸ“„
Pure Vision-Language-Action Models: A Comprehensive Surveyβ€”2025πŸ“„
A Survey on Efficient Vision-Language-Action ModelsYu et al.2025πŸ“„
A Survey on VLA Models for Autonomous DrivingJiang et al.2025πŸ“„
VLA in Robotics: A Survey of Datasets, Benchmarks, and Data EnginesWang et al.2026πŸ“„

Foundation Models & Embodied AI

TitleAuthorsYearLinks
Foundation Models in Robotics: Applications, Challenges, and the FutureFiroozi et al.2023πŸ“„
Toward General-Purpose Robots via Foundation Models: A SurveyHu et al.2023πŸ“„
Aligning Cyber Space with Physical World: A Survey on Embodied AILiu et al.2024πŸ“„
What Foundation Models Can Bring for Robot Learning in Manipulation: A SurveyLi et al.2024πŸ“„
Deep Reinforcement Learning for Robotics: A Survey of Real-World SuccessesTang et al.2024πŸ“„
Generative AI in Robotic Manipulation: A SurveyZhang et al.2025πŸ“„
A Survey of Sim-to-Real Methods in RL with Foundation ModelsDa et al.2025πŸ“„
Behavior Foundation Model: Towards Next-Generation Whole-Body Control of HumanoidsYuan et al.2025πŸ“„
Towards a Unified Understanding of Robot Manipulation: A Comprehensive SurveyBai et al.2025πŸ“„
Robotic Foundation Models for Industrial Control: A Survey & Readiness AssessmentKube et al.2026πŸ“„

πŸ€– Vision-Language-Action (VLA) Models

Following the action-tokenization view (Zhong et al., 2025), the primary split is by how actions are represented; capability-oriented subsections (reasoning, 3D/4D, efficiency, RL) cut across it. A few pre-/non-VLM generalist policies (e.g., RT-1, Octo) are listed alongside their successors to show lineage β€” see Foundational Robot Policies for the strictly non-VLA baselines.

By Action Representation

Autoregressive / Discrete-Token VLA

Actions are binned into discrete tokens and decoded like text. Simple and VLM-native; high-frequency dexterity needs better tokenizers (see FAST).

ModelTitleYearLinks
VLA-0Building SOTA VLAs with Zero Modification2025πŸ“„ Β· 🌐
UniVLAUnified Vision-Language-Action Model (native multimodal tokens)2025πŸ“„ Β· 🌐
Ο€0-FASTAutoregressive Ο€0 variant using the FAST action tokenizer2025πŸ“„ Β· 🌐
OpenVLAAn Open-Source Vision-Language-Action Model2024πŸ“„ Β· 🌐 Β· πŸ’»
RT-2VLA Models Transfer Web Knowledge to Robotic Control2023πŸ“„ Β· 🌐
RT-1Robotics Transformer for Real-World Control at Scale2022πŸ“„ Β· 🌐 Β· πŸ’»

Diffusion-based VLA

A diffusion action head denoises continuous action chunks conditioned on vision-language features.

ModelTitleYearLinks
RoboVLMsTowards Generalist Robot Policies: What Matters in Building VLAs2024πŸ“„ Β· 🌐
CogACTA Foundational VLA Model for Synergizing Cognition and Action2024πŸ“„
TinyVLAFast, Data-Efficient VLA Models for Manipulation2024πŸ“„ Β· 🌐
OctoAn Open-Source Generalist Robot Policy2024πŸ“„ Β· 🌐 Β· πŸ’»

Flow-Matching VLA

A conditional flow/vector field transports noise to action chunks β€” the dominant head for current SOTA generalist VLAs.

ModelTitleYearLinks
Ο€*0.6A VLA That Learns From Experience2025πŸ“„ Β· 🌐
X-VLASoft-Prompted Transformer as a Scalable Cross-Embodiment VLA2025πŸ“„ Β· 🌐 Β· πŸ’»
SmolVLAA VLA for Affordable and Efficient Robotics2025πŸ“„ Β· πŸ’»
Ο€0.5A VLA with Open-World Generalization2025πŸ“„ Β· 🌐
Gemini RoboticsBringing AI into the Physical World2025πŸ“„ Β· 🌐
GR00T N1An Open Foundation Model for Generalist Humanoid Robots2025πŸ“„ Β· πŸ’»
EO-1An Open Unified Embodied Foundation Model (interleaved reasoning + acting)2025πŸ“„ Β· 🌐
GR-3Large-Scale Vision-Language-Action Model (Technical Report)2025πŸ“„
FLOWERDemocratizing Generalist Robot Policies with Efficient VLA Flow Policies2025πŸ“„
Ο€0A Vision-Language-Action Flow Model for General Robot Control2024πŸ“„ Β· 🌐

By Capability

Reasoning & Dual-System (Fast–Slow) VLA

Explicit chain-of-thought / embodied reasoning, or a slow System-2 planner paired with a fast System-1 controller.

ModelTitleYearLinks
ACoT-VLAAction Chain-of-Thought for VLA Models2026πŸ“„ Β· πŸ’»
Gemini Robotics 1.5Embodied Reasoning & Motion Transfer2025πŸ“„
ThinkActVLA Reasoning via Reinforced Visual Latent Planning2025πŸ“„
OpenHelixA Short Survey & Open-Source Dual-System VLA2025πŸ“„
FiS-VLAFast-in-Slow: A Dual-System Foundation Model for Unified Fast–Slow Reasoning2025πŸ“„
WALL-OSSIgniting VLMs toward the Embodied Space2025πŸ“„ Β· πŸ’»
CoT-VLAVisual Chain-of-Thought Reasoning for VLA2025πŸ“„

3D / 4D-Aware VLA

Policies that reason over explicit 3D/4D structure (point clouds, occupancy, predicted future frames) rather than 2D images alone. (VoxPoser, a zero-shot 3D value-map planner, lives under Foundational Robot Policies.)

ModelTitleYearLinks
3D-VLAA 3D Vision-Language-Action Generative World Model2024πŸ“„

Efficient & Real-Time VLA

Compression, caching, parallel decoding, and distillation to make VLAs small and fast enough for real-time / edge control (Guan et al., 2025).

ModelTitleYearLinks
FASTERRethinking Real-Time Flow VLAs2026πŸ“„
RTCReal-Time Chunking: Running VLAs at Real-Time Speed2025πŸ“„
NanoVLARouting-Decoupled VLA for Nano-Sized Generalist Policies2025πŸ“„
VLA-AdapterA Tiny-Scale VLA Paradigm2025πŸ“„
OpenVLA-OFTFine-Tuning VLAs: Optimizing Speed and Success2025πŸ“„ Β· 🌐
TinyVLAFast, Data-Efficient VLA Models2024πŸ“„ Β· 🌐

RL Fine-Tuning for VLA

Reinforcement learning (often on top of flow-/diffusion-based VLAs) to improve over imitation-only training.

ModelTitleYearLinks
Ο€_RLOnline RL Fine-Tuning for Flow-based VLAs2025πŸ“„
VLA-RFTRL Fine-Tuning with Verified Rewards in World Simulators2025πŸ“„
SimpleVLA-RLScaling VLA Training via Reinforcement Learning2025πŸ“„
ConRFTA Reinforced Fine-Tuning Method for VLA via Consistency Policy2025πŸ“„

🌎 World & World-Action Models

Organized by what the model predicts and how it is built, following the embodied-world-model taxonomy of Li et al., 2025 and the WAM split popularized by awesome-vla-wam.

General World Models

General-purpose models of environment dynamics β€” spanning classical latent world models for model-based RL (World Models, DreamerV3) and modern large-scale video / foundation world models β€” used for planning, neural simulation, or as backbones for WAMs.

ModelTitleYearLinks
Cosmos-Predict2.5World Simulation with Video Foundation Models for Physical AI2025πŸ“„ Β· πŸ’»
Cosmos-Reason1From Physical Common Sense to Embodied Reasoning2025πŸ“„
CosmosWorld Foundation Model Platform for Physical AI2025πŸ“„ Β· 🌐
V-JEPA 2Self-Supervised Video Models Enable Understanding, Prediction & Planning2025πŸ“„
iVideoGPTInteractive VideoGPTs are Scalable World Models2024πŸ“„
GenieGenerative Interactive Environments2024πŸ“„
DreamerV3Mastering Diverse Domains through World Models2023πŸ“„ Β· πŸ’»
UniSimLearning Interactive Real-World Simulators2023πŸ“„
World ModelsRecurrent latent world model + controller (origin of the term)2018πŸ“„

WAM from Video Generation

A (text-/image-conditioned) video generator imagines future frames; actions are recovered via an inverse-dynamics / action head.

ModelTitleYearLinks
DreamZeroWorld Action Models are Zero-shot Policies2026πŸ“„ Β· 🌐
DiT4DiTJointly Modeling Video Dynamics and Actions2026πŸ“„
Cosmos PolicyFine-Tuning Video Models for Visuomotor Control & Planning2026πŸ“„ Β· 🌐
Video2ActA Dual-System Video Diffusion Policy2025πŸ“„
GR-2A Generative Video-Language-Action Model with Web-Scale Knowledge2024πŸ“„
GR-1Large-Scale Video Generative Pre-training for Visual Robot Manipulation2023πŸ“„

WAM from VLMs

A pretrained VLM is turned into a world model (e.g., predicting goal images / object-centric futures) that then drives action.

ModelTitleYearLinks
DreamVLAA VLA Model Dreamed with Comprehensive World Knowledge2025πŸ“„
Goal-VLAImage-Generative VLMs as Object-Centric World Models for VLA2025πŸ“„

Unified VLA–World Models

Single architectures that jointly learn to act and to predict world dynamics, blurring the VLA/WAM boundary.

ModelTitleYearLinks
RynnVLA-002A Unified Vision-Language-Action and World Model2025πŸ“„ Β· πŸ’»
WholeBodyVLAUnified Latent VLA for Whole-Body Loco-Manipulation2025πŸ“„ Β· πŸ’»
WorldVLATowards an Autoregressive Action World Model2025πŸ“„ Β· πŸ’»

Latent & JEPA World Models

Self-supervised latent predictive models (non-reconstructive joint-embedding / JEPA). The JEPA foundations (I-JEPA) learn to predict in representation space; the action-conditioned variant (V-JEPA 2-AC) turns that prior into a world model for planning.

ModelTitleYearLinks
V-JEPA 2-ACAction-Conditioned Latent World Model for Zero-Shot Planning2025πŸ“„
I-JEPAImage-based Joint-Embedding Predictive Architecture (representation foundation)2023πŸ“„

Domain World Models (Driving & Navigation)

ModelTitleYearLinks
GAIA-2A Controllable Multi-View Generative World Model for Autonomous Driving2025πŸ“„
Navigation World ModelsConditional Diffusion Transformer for Navigation2024πŸ“„
GAIA-1A Generative World Model for Autonomous Driving2023πŸ“„

🧩 Action Representations & Tokenization

Building blocks shared across VLA and WAM policies β€” how continuous actions become learnable targets.

Discrete / Autoregressive Tokenizers

MethodTitleYearLinks
FASTEfficient (DCT-based) Action Tokenization for VLAs2025πŸ“„ Β· 🌐
BeTBehavior Transformers: Cloning k Modes with One Stone2022πŸ“„

Continuous & Chunked Action Policies

Heads that emit continuous action chunks β€” by denoising diffusion (Diffusion Policy) or by chunked sequence prediction with a CVAE (ACT). Flow-matching heads (Ο€0, SmolVLA, …) are listed with their models under Flow-Matching VLA.

MethodTitleYearLinks
Diffusion PolicyVisuomotor Policy Learning via Action Diffusion2023πŸ“„ Β· 🌐
ACT / ALOHAAction Chunking with Transformers2023πŸ“„ Β· 🌐

🦾 Foundational Robot Policies

Non-VLA policies and planners that remain standard baselines in the experimental tables of the papers above. (Diffusion Policy, ACT, and BeT are described under Action Representations.)

MethodTitleYearLinks
CrossFormerScaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion & Flight2024πŸ“„ Β· πŸ’»
RoboFlamingoVision-Language Foundation Models as Effective Robot Imitators2023πŸ“„ Β· πŸ’»
VoxPoserComposable 3D Value Maps for Robotic Manipulation (zero-shot LLM + 3D planner)2023πŸ“„ Β· 🌐
RT-1Robotics Transformer for Real-World Control at Scale2022πŸ“„

πŸ“¦ Resources

Datasets

NameDescriptionScaleLinks
Open X-EmbodimentCross-embodiment aggregation behind the RT-X models1M+ traj Β· 22 embodimentsπŸ“„ Β· 🌐
AgiBot WorldLarge-scale real-world manipulation (Colosseo)1M+ traj Β· 217 tasksπŸ“„ Β· 🌐
EgoScaleScaling dexterous manipulation with diverse egocentric human dataegocentric Β· 2026πŸ“„
DexCanvasHuman demos ↔ robot learning for dexterous manipulationdexterousπŸ“„
Galaxea Open-WorldMobile-bimanual dataset paired with the G0 dual-system VLA500 hrs Β· 150 tasksπŸ“„ Β· πŸ’»
DROIDIn-the-wild Franka manipulation across 3 continents76K traj Β· 564 scenesπŸ“„ Β· 🌐
RoboMINDMulti-embodiment teleop incl. labeled failures107K traj Β· 479 tasksπŸ“„ Β· 🌐
BridgeData V2WidowX manipulation w/ language + goal images60K traj Β· 24 envsπŸ“„ Β· 🌐
RH20TContact-rich skills w/ paired human demos110K+ seq Β· 147 tasksπŸ“„ Β· 🌐
Ego-Exo4DSimultaneous ego + exo video of skilled activity1,286 hrsπŸ“„ Β· 🌐
Ego4DMassive egocentric daily-life video3,670 hrsπŸ“„ Β· 🌐

Benchmarks

NameDescriptionLinks
LIBEROLifelong robot-learning, 130 manipulation tasks (de-facto VLA eval)πŸ“„ Β· πŸ’»
CALVINLong-horizon language-conditioned manipulationπŸ“„ Β· πŸ’»
SimplerEnvReal-to-sim evaluation for manipulation policiesπŸ“„ Β· 🌐
RoboCasaLarge-scale kitchen simulation (100 tasks)πŸ“„ Β· 🌐
VLABenchWorld-knowledge & long-horizon language tasksπŸ“„ Β· 🌐
ManiSkill3GPU-parallel manipulation (30K+ FPS)πŸ“„ Β· 🌐
THE COLOSSEUMRobustness under 14 environmental perturbationsπŸ“„ Β· 🌐
RoboArenaDistributed crowd-sourced real-world policy evalπŸ“„ Β· 🌐
RoboChallengeLarge-scale real-robot evaluation of embodied policiesπŸ“„
RobotArena ∞Scalable robot benchmarking via real-to-sim translationπŸ“„ Β· 🌐
WorldArenaPerception & functional-utility benchmark for embodied world modelsπŸ“„
Meta-World50 tabletop tasks for multi-task / meta-RLπŸ“„ Β· πŸ’»
RLBench100 hand-designed manipulation tasksπŸ“„ Β· πŸ’»

Simulation Platforms

NameDescriptionLinks
Isaac Sim / Isaac LabGPU-native robotics sim + RL/IL framework (Omniverse/USD)🌐
MuJoCo / MJXStandard rigid-body engine + JAX/XLA parallel variant🌐
GenesisGenerative, multi-solver physics platform (up to ~43M FPS)🌐
ManiSkillGPU-parallel manipulation simulator on SAPIEN🌐
SAPIENPart-level articulated-object simulator (PartNet-Mobility)πŸ“„ Β· 🌐
HabitatPhotorealistic indoor navigation & rearrangement🌐
ThreeDWorldMultimodal Unity3D sim (vision + audio + physics)πŸ“„ Β· 🌐
NewtonOpen, differentiable GPU physics engine (NVIDIA + DeepMind + Disney)🌐

Tools & Frameworks

NameDescriptionLinks
LeRobotEnd-to-end PyTorch robot-learning library + datasets + low-cost HWπŸ“„ Β· πŸ’»
openpiOpen models & training/inference for Ο€0, Ο€0-FAST, Ο€0.5πŸ’»
Isaac GR00TOpen humanoid foundation-model framework + checkpointsπŸ’»
OpenVLATraining / LoRA fine-tuning for the 7B OpenVLA modelπŸ’»
OctoJAX/Flax generalist transformer policy on OXEπŸ’»
robomimic / robosuiteLearning-from-demonstration framework + MuJoCo manipulation simπŸ’»
HIL-SERLHuman-in-the-loop, sample-efficient real-world RLπŸ’»

πŸ—‚οΈ Extended Paper Index (Auto-Curated, Newest First)

A broader, continuously-mined index of recent arXiv work that complements the curated highlights above β€” 182 additional papers, newest first. Last updated: 2026-09-30. Auto-generated from data/*.json by scripts/expand_papers.py; papers already highlighted above are omitted here to avoid duplication.

VLA β€” General & Manipulation Β· 45 papers
PaperAuthorsDateLinks
Online Evolution Strategy for Flow-Matching VLA Policies via Self-Supervised Trajectory Distribution OptimizationGongxin Yao, Yongsheng Zhao et al.2026-09-30
Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential GuidanceYanyan Zhang, Disheng Liu et al.2026-09-29
Taming VLAs under Robot Execution Errors: Self-Compensation and Stress TestingSohyun Lee, Yoonjae Baek et al.2026-09-29
Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent LookaheadJunghyun Kim, Ngseo Kim et al.2026-09-29
CoRe-VLA: Preserving Cross-View Coordination in VLAs under Camera ShiftsTianhang Pan, Xuanhao Wang et al.2026-09-29
VLALight: A Vision-Language-Action Model for Traffic Signal ControlPan Zhang, Siqi Lai et al.2026-09-29
Where Predictive Supervision Goes Shapes What VLA Policies LearnHanseul Kim, Jewon Yeom et al.2026-09-29
The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent InterfaceYuxiang Liu, Lizhi Yang et al.2026-09-28
FocusVLA: Focused Visual Utilization for Vision-Language-Action ModelsYichi Zhang, Weihao Yuan et al.2026-03-30
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic ManipulationHongyu Yan, Qiwei Li et al.2026-03-29
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and GenerationYang Liu, Pengxiang Ding et al.2026-03-26
ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-MakingYoung-Chae Son, Dae-Kwan Ko et al.2026-03-26
$Ο€$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial ManipulationJohnathan Tucker, Denis Liu et al.2026-03-26
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action ModelsJiaying Zhou, Zhihao Zhan et al.2026-03-25
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAsHaoran Yuan, Weigang Yi et al.2026-03-24
Gaze-Regularized Vision-Language-Action Models for Robotic ManipulationAnupam Pani, Yanchao Yang2026-03-24
CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action ModelsYouzhi Liu, Li Gao et al.2026-03-24
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action ModelsZhou Fang, Jiaqi Wang et al.2026-03-18
The Steep-spectrum Radio-loud AGN Luminosity Function and Its Implications for Black Hole Growth and Star FormationWenjie Wang, Zunli Yuan et al.2026-03-16
Building Explicit World Model for Zero-Shot Open-World Object ManipulationXiaotong Li, Gang Chen et al.2026-03-14
Beyond Dense Futures: World Models as Structured Planners for Robotic ManipulationMinghao Jin, Mozheng Liao et al.2026-03-13
Adaptive Capacity Allocation for Vision Language Action Fine-tuningDonghoon Kim, Minji Bae et al.2026-03-08
HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the WildZiyang Zhao, Shuheng Wang et al.2026-03-06
CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuningYike Zhang, Yaonan Wang et al.2026-02-13
HoloBrain-0 Technical ReportHorizon Robotics2026-02-07🌐
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and ExpansionRalf RΓΆmer, Yi Zhang et al.2026-01-14
Diverse stages of star formation in the IRAS 18162-2048 region. Emergence of UV FeedbackR. Fedriani, G. Anglada et al.2025-12-08
Mixture of Horizons in Action ChunkingDong Jing, Gang Wang et al.2025-11-24
A scaling relationship for non-thermal radio emission from ordered magnetospheres - II. Investigating the efficiency of relativistic electron production in magnetospheres of BA-type starsP. Leto, S. Owocki et al.2025-11-07
First X-ray and radio polarimetry of the neutron star low-mass X-ray binary GX 17+2Unnati Kashyap, Thomas J. Maccarone et al.2025-10-06
Deciphering the radio-star formation correlation on kpc scales. IV. Radio halos of highly-inclined Virgo cluster spiral galaxiesB. Vollmer, M. Soida et al.2025-10-03
Masses, Star-Formation Efficiencies, and Dynamical Evolution of 18,000 HII RegionsDebosmita Pathak, Adam K. Leroy et al.2025-09-26
X-ray and radio polarimetry of the neutron star low mass X-ray binary GX 13+1Unnati Kashyap, Thomas J. Maccarone et al.2025-08-07
Protostellar Outflows at the EarliesT Stages (POETS). VIII. The jets in the intermediate-mass star-forming region G105.42+9.88 (alias LkHΞ± 234)Luca Moscadelli, Fabrizio Massi et al.2025-08-05
Star formation histories and gas content limits of three ultra-faint dwarfs on the periphery of M31Michael G. Jones, David J. Sand et al.2025-08-01
Quenching Through Tidal Gas Removal: Molecular Gas and Star Formation in Tidal Tails of z ~ 0.7 Post-Starburst GalaxiesVincenzo R. D'Onofrio, Justin S. Spilker et al.2025-07-28
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile ManipulationSixiang Chen, Jiaming Liu et al.2025-07-02
The Radio Spectral Energy Distribution and Star Formation Calibration in MIGHTEE-COSMOS Highly Star-Forming Galaxies at 1.5 < z < 3.5Fatemeh Tabatabaei, Maryam Khademi et al.2025-06-19
Semi-empirical constraints on the HI mass function of star-forming galaxies and $Ξ©_{\rm HI}$ at $z\sim 0.37$ from interferometric surveysFrancesco Sinigaglia, Alessandro Bianchetti et al.2025-06-12
A persistent disk wind and variable jet outflow in the neutron-star low-mass X-ray binary GX 13+1Daniele Rogantini, Jeroen Homan et al.2025-04-07
X-ray and radio data obtained by XMM-Newton and VLA constrain the stellar wind of the magnetic quasi-Wolf-Rayet star in HD45166P. Leto, L. M. Oskinova et al.2025-03-10
The Arp 240 Galaxy Merger: A Detailed Look at the Molecular Kennicutt-Schmidt Star Formation Law on Sub-kpc ScalesAlejandro Saravia, Eduardo Rodas-Quito et al.2024-12-10
Runaway O and Be stars found using Gaia DR3, new stellar bow shocks and search for binariesM. Carretero-Castrillo, M. RibΓ³ et al.2024-12-10
A-VL: Adaptive Attention for Large Vision-Language ModelsJunyang Zhang, Mu Yuan et al.2024-09-23
HiRT: Enhancing Robotic Control with Hierarchical Robot TransformersJianke Zhang, Yanjiang Guo et al.2024-09-12
VLA β€” Reasoning, Planning & Dual-System Β· 5 papers
VLA β€” Autonomous Driving Β· 13 papers
PaperAuthorsDateLinks
Vision-Language-Action Autonomous Driving Agent with Language-based MemoryKai Yan, Xiangyu Chen et al.2026-09-29
Data-Efficient Adaptation of a Driving VLA to Class 8 TrucksSatyajeet Das, Aaron Buxbaum et al.2026-09-29
Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action PoliciesYaxin Zhao, Dianye Huang et al.2026-09-29
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early ObservationYiran Shi, Dongqi Guo et al.2026-03-30
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous DrivingQiqi Liu, Huan Xu et al.2026-03-28
Vega: Learning to Drive with Natural Language InstructionsSicheng Zuo, Yuxuan Li et al.2026-03-26
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized DrivingZehao Wang, Huaide Jiang et al.2026-03-26
ETA-VLA: Efficient Token Adaptation via Temporal Fusion and Intra-LLM Sparsification for Vision-Language-Action ModelsYiru Wang, Anqing Jiang et al.2026-03-26
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action ModelsJintao Cheng, Haozhe Wang et al.2026-03-24
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous DrivingZihan You, Hongwei Liu et al.2026-03-09
VLANeXt: Recipes for Building Strong VLA ModelsLiu, Xiangyu et al.2026-02-18🌐
A global view on star formation: The GLOSTAR Galactic plane survey XII. Effelsberg's continuum view and data releaseY. Gong, W. Reich et al.2025-12-17
Low-frequency spectra of neutron star + OB supergiant binaries: Does wind density drive persistent and flaring modes of accretion?J. van den Eijnden, L. Sidoli et al.2025-08-06
VLA β€” Dexterous & Humanoid Β· 2 papers
VLA β€” 3D / 4D & Spatial Β· 4 papers
VLA β€” RL & Post-Training Β· 7 papers
VLA β€” Efficient & Real-Time Β· 19 papers
PaperAuthorsDateLinks
Rho: A Foundation for Efficiently Adaptable VLA ModelsRho Team, Simran Bagaria et al.2026-09-29
Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA ControlZibo Wang, Haochen Han et al.2026-09-29
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLAYi Chen, Yuying Ge et al.2026-03-31
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and AccurateChen Yang, Yucheng Hu et al.2026-03-27
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow MatchingJiayi Chen, Wenxuan Song et al.2026-03-27
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time PerformanceWenxuan Song, Jiayi Chen et al.2026-03-26
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action ModelsPeiju Liu, Jinming Liu et al.2026-03-26
Agile-VLA: Few-Shot Industrial Pose Rectification via Implicit Affordance AnchoringTeng Yan, Zhengyang Pei et al.2026-03-24
Fast-WAM: Do World Action Models Need Test-time Future Imagination?Tianyuan Yuan, Zibin Dong et al.2026-03-17
FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic ManipulationYao Li, Peiyuan Tang et al.2026-02-27
Learning Native Continuation for Action Chunking Flow PoliciesYufeng Liu, Hang Yu et al.2026-02-13
Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action ModelsYuting Huang, Leilei Ding et al.2026-01-31
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token CachingYujie Wei, Jiahan Fan et al.2026-01-31
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic ManipulationWenda Yu, Tianshi Wang et al.2026-01-27
U-DiT Policy: U-shaped Diffusion Transformers for Robotic ManipulationLinzhi Wu, Aoran Mei et al.2025-09-29
Leave No Observation Behind: Real-time Correction for VLA Action ChunksKohei Sendai, Maxime Alvarez et al.2025-09-27
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel DecodingWenxuan Song, Jiayi Chen et al.2025-03-04
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language TuningZhiwei Hao, Jianyuan Guo et al.2024-10-23
VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language TasksYi-Lin Sung, Jaemin Cho et al.2021-12-13
VLA β€” Safety, Robustness & Evaluation Β· 12 papers
PaperAuthorsDateLinks
Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action AccelerationQiwei Chen, Kaijun Zhou et al.2026-09-29
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action ModelsXiyang Wu, Guangyao Shi et al.2026-03-26
SOMA: Strategic Orchestration and Memory-Augmented System for Vision-Language-Action Model Robustness via In-Context AdaptationZhuoran Li, Zhiyang Li et al.2026-03-25
ROBOGATE: Adaptive Failure Discovery for Safe Robot Policy Deployment via Two-Stage Boundary-Focused SamplingAzuki Kim2026-03-23
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action ModelsZhilong Zhang, Haoxiang Ren et al.2026-03-21
Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic ControlZunzhe Zhang, Runhan Huang et al.2026-03-18
World2Act: Latent Action Post-Training via Skill-Compositional World ModelsAn Dinh Vuong, Tuan Van Vo et al.2026-03-11
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action ModelYuanjie Lu, Beichen Wang et al.2026-03-09
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World ModelsXiaoquan Sun, Zetian Xu et al.2026-03-09
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action ModelsHyeongjun Heo, Seungyeon Woo et al.2026-03-06
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action ModelsHyeonbeom Choi, Daechul Ahn et al.2026-02-04
SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action ModelsBingxin Xu, Yuzhang Shang et al.2026-01-20
World Models β€” General & Foundation Β· 10 papers
World Models β€” Video Generation & WAM Β· 8 papers
World Models β€” Driving & Navigation Β· 6 papers
Policies β€” Diffusion & Flow Β· 27 papers
PaperAuthorsDateLinks
Encoding Predictability and Legibility for Style-Conditioned Diffusion PolicyAdrien Jacquet CrΓ©tides, Mouad Abrini et al.2026-03-17
ReMAP-DP: Reprojected Multi-view Aligned PointMaps for Diffusion PolicyXinzhang Yang, Renjun Wu et al.2026-03-16
REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement LearningZhaoyuan Gu, Yipu Chen et al.2026-03-14
PPGuide: Steering Diffusion Policies with Performance Predictive GuidanceZixing Wang, Devesh K. Jha et al.2026-03-11
SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot ManipulationYouqiang Gui, Yuxuan Zhou et al.2026-03-05
Diffusion Policy through Conditional Proximal Policy OptimizationBen Liu, Shunpeng Yang et al.2026-03-05
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion PolicyPengyuan Wu, Pingrui Zhang et al.2026-03-02
ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent ManipulationEnyi Wang, Wen Fan et al.2026-02-25
Preference Aligned Visuomotor Diffusion Policies for Deformable Object ManipulationMarco Moletta, Michael C. Welle et al.2026-02-10
SERFN: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing FlowsChenyu Yang, Denis Tarasov et al.2026-02-10
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic ManipulationYuxuan Hu, Xiangyu Chen et al.2026-02-07
Moving On, Even When You're Broken: Fail-Active Trajectory Generation via Diffusion Policies Conditioned on Embodiment and TaskGilberto G. Briscoe-Martinez, Yaashia Gautam et al.2026-02-02
RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human FeedbackAmitesh Vatsa, Zhixian Xie et al.2026-01-31
Self-Imitated Diffusion Policy for Efficient and Robust Visual NavigationRunhua Zhang, Junyi Hou et al.2026-01-30
Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion PoliciesCe Hao, Xuanran Zhai et al.2026-01-29
ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot ManipulationWeize Xie, Yi Ding et al.2026-01-19
Sparse ActionGen: Accelerating Diffusion Policy with Real-time PruningKangye Ji, Yuan Meng et al.2026-01-19
CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action SpaceBingyi Liu, Jinbo He et al.2026-01-09
Learning Diffusion Policy from Primitive Skills for Robot ManipulationZhihao Gu, Ming Yang et al.2026-01-05
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic ControlWonhyeok Choi, Shutong Ding et al.2026-01-05
Flexible Multitask Learning with Factorized Diffusion PolicyChaoqi Liu, Haonan Chen et al.2025-12-26
Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic ManipulationKangchen Lv, Mingrui Yu et al.2025-12-19
ISS Policy : Scalable Diffusion Policy with Implicit Scene SupervisionWenlong Xia, Jinhao Zhang et al.2025-12-17
Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic TasksAileen Liao, Dong-Ki Kim et al.2025-12-08
CAPE: Context-Aware Diffusion Policy Via Proximal Mode Expansion for Collision AvoidanceRui Heng Yang, Xuan Zhao et al.2025-11-27
Learning Diffusion Policies for Robotic Manipulation of Timber Joinery under Fabrication UncertaintySalma Mozaffari, Daniel Ruan et al.2025-11-21
UltraDP: Generalizable Carotid Ultrasound Scanning with Force-Aware Diffusion PolicyRuoqu Chen, Xiangjie Yan et al.2025-11-19
Policies β€” Imitation & Behavior Learning Β· 10 papers
Policies β€” Robot Learning & Manipulation Β· 14 papers
PaperAuthorsDateLinks
Learning Multi-View Spatial Reasoning from Cross-View RelationsSuchae Jeong, Jaehwi Song et al.2026-03-30
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory GenerationMotonari Kambara, Koki Seno et al.2026-03-26
Chunk-Boundary Artifact in Action-Chunked Generative Policies: A Noise-Sensitive Failure MechanismRui Wang2026-03-12
Real-Time Robot Execution with Masked Action ChunkingHaoxuan Wang, Gengyu Zhang et al.2026-01-27
Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse RewardJiarui Yang, Bin Zhu et al.2025-08-15
Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou et al.2025-07-10
Real-Time Execution of Action Chunking Flow PoliciesKevin Black, Manuel Y. Galliker et al.2025-06-09
Learning Bimanual Manipulation via Action Chunking and Inter-Arm Coordination with TransformersTomohiro Motoda, Ryo Hanai et al.2025-03-18
MissionGPT: Mission Planner for Mobile Robot based on Robotics Transformer ModelVladimir Berman, Artem Bazhenov et al.2024-11-07
VQ-ACE: Efficient Policy Search for Dexterous Robotic Manipulation via Action Chunking EmbeddingChenyu Yang, Davide Liconti et al.2024-11-05
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual ManipulationAndrew Lee, Ian Chuang et al.2024-09-12
Bringing the RT-1-X Foundation Model to a SCARA robotJonathan Salzer, Arnoud Visser2024-09-05
Logically Constrained Robotics Transformers for Enhanced Perception-Action PlanningParv Kapoor, Sai Vemprala et al.2024-08-09
SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust AttentionIsabel Leal, Krzysztof Choromanski et al.2023-12-04

πŸ“‹ Full Paper Index & Baselines

πŸ“Š Click to expand the complete paper list and baseline methods

The curated tables above highlight landmark and representative work. For the exhaustive, auto-maintained index and the baseline methods extracted from experimental tables, see:

Quick Reference (common baselines)

FamilyKey baselines
VLART-1, RT-2, OpenVLA, Octo, Ο€0, Ο€0.5, X-VLA, UniVLA, SmolVLA
PolicyDiffusion Policy, ACT, BeT, RoboFlamingo, CrossFormer
World ModelDreamerV3, I-JEPA, V-JEPA 2, Genie, Cosmos, GR-1/GR-2

🀝 Contributing

Contributions are very welcome! To add or fix a paper:

  1. Add a paper β€” open a PR placing it in the appropriate category (keep tables sorted newest-first), or open an issue with the arXiv link.
  2. Fix an error β€” submit a PR with the correction.
  3. New papers appear automatically β€” the πŸ†• Latest Papers section and the πŸ—‚οΈ Extended Paper Index are regenerated daily by the scraper; do not hand-edit content between the auto markers.

To run the discovery pipeline locally:

pip install -r requirements.txt
python scripts/arxiv_scraper.py --max-results 50 --days-back 30   # writes data/papers.json
python scripts/update_readme.py                                  # refreshes the πŸ†• auto section
python scripts/expand_papers.py                                  # refreshes the πŸ—‚οΈ extended index
python scripts/build_site.py                                     # rebuilds the GitHub Pages site (index.html)

License

Released under the MIT License.

Acknowledgments

Inspired by awesome-vla-wam, awesome-physical-ai, and awesome-vla-study. Taxonomy grounded in the surveys listed above.


If you find this repository useful, please consider giving it a ⭐
awesone
emobodied-ai
robotics
video-action-models
vla
world-action-model

HyperbolicCurve/Awesome-World-Action-Model

A curated list of academic papers and resources on Vision-Language-Action (VLA) and World Action Models (WAM)

Python

35

202 commits

updated Oct 1, 2026

See the code

README

πŸ€– Awesome World Action Models Awesome

A curated, continuously-updated reading list of World Action Models (WAM), Vision-Language-Action (VLA) models, and Embodied AI β€” organized by a survey-grounded taxonomy.

Last Update PRs Welcome License: MIT Stars

🌐 Website: hyperboliccurve.github.io/Awesome-World-Action-Model


Overview

The push toward general-purpose robots has produced two converging families of foundation models:

  • Vision-Language-Action (VLA) models inherit the language grounding and visual understanding of pretrained Vision-Language Models (VLMs) and adapt them to emit actions β€” a scalable route to language-conditioned policies.
  • World Action Models (WAM) start from a world model / video backbone that predicts how a scene evolves, and adapt that predictive prior to emit actions β€” trading the "languageβ†’motion" grounding gap for a "dynamicsβ†’action" one.

These two families overlap: a WAM built on a pretrained VLM is simultaneously a VLA and a WAM. This list maps that landscape with a taxonomy grounded in the recent survey literature (see Surveys), so each category has a clear, defensible scope rather than an ad-hoc label.

[!NOTE] Legend β€” πŸ“„ arXiv Β· 🌐 project page Β· πŸ’» code Β· πŸ“Š dataset/benchmark. Tables are sorted newest-first within each category. The πŸ†• Latest Papers section is refreshed daily from arXiv by a GitHub Action; everything else is hand-curated.


Taxonomy at a Glance

flowchart TD
    A[Robot Foundation Models] --> B[Vision-Language-Action<br/>VLA]
    A --> C[World &amp; World-Action Models<br/>WM / WAM]
    A --> R[Action Representations]
    A --> P[Foundational Policies]

    B --> B1[By Action Representation:<br/>Autoregressive Β· Diffusion Β· Flow-Matching]
    B --> B2[By Capability:<br/>Reasoning/Dual-System Β· 3D-4D Β· Efficient Β· RL Fine-Tuning]

    C --> C1[Foundation / General World Models]
    C --> C2[WAM from Video Generation]
    C --> C3[WAM from VLMs]
    C --> C4[WAM from Scratch Β· Latent / JEPA]
    C --> C5[Domain: Driving Β· Navigation]

    R --> R1[Discrete / Autoregressive Tokenizers]
    R --> R2[Diffusion &amp; Flow-Matching Policies]

Table of Contents


πŸ”‘ Key Definitions

TermDefinitionCanonical reference
Vision-Language-Action (VLA)A robot policy that adapts a pretrained VLM to map images + language instructions to actions.RT-2 (Brohan et al., 2023)
World Model (WM)A learned model that predicts future states of an environment (in pixels, latents, or 3D/4D), used for planning, simulation, or representation.World Models (Ha & Schmidhuber, 2018)
World Action Model (WAM)A policy that leverages world-modeling capability (predicting future states) for action prediction β€” typically by adapting a video / world-model backbone to emit actions.GR-1 (Wu et al., 2023)

[!IMPORTANT] VLA ∩ WAM. The families intersect: a WAM built on a pretrained VLM is both. The split in this list is by what prior the model starts from β€” VLM-style vision-language priors (VLA) vs. video/dynamics priors (WAM) β€” and, within VLA, by how actions are represented, the axis most surveys agree is the field's clearest discriminator.


πŸ†• Latest Papers (Auto-updated)

Papers are automatically fetched daily from arXiv. Last updated: 2026-09-30

VLA

World Model

Policy


πŸ“š Surveys

Recent surveys that define the field and motivate the taxonomy used here.

World & Embodied World Models

TitleAuthorsYearLinks
Understanding World or Predicting Future? A Comprehensive Survey of World ModelsDing et al.2024πŸ“„
A Comprehensive Survey on World Models for Embodied AILi et al.2025πŸ“„
3D and 4D World Modeling: A SurveyKong et al.2025πŸ“„
Learning Embodied Intelligence from Physical Simulators and World ModelsLong et al.2025πŸ“„
Embodied AI: From LLMs to World ModelsFeng et al.2025πŸ“„
World Model for Robot Learning: A Comprehensive SurveyHou et al.2026πŸ“„
Modeling the Mental World for Embodied AI: A Comprehensive ReviewLiu et al.2026πŸ“„
The Role of World Models in Shaping Autonomous Driving: A SurveyTu et al.2025πŸ“„

Vision-Language-Action

TitleAuthorsYearLinks
A Survey on Vision-Language-Action Models for Embodied AIMa et al.2024πŸ“„
A Survey on VLA Models: An Action Tokenization PerspectiveZhong et al.2025πŸ“„
VLA Models: Concepts, Progress, Applications and ChallengesSapkota et al.2025πŸ“„
Large VLM-based VLA Models for Robotic Manipulation: A SurveyShao et al.2025πŸ“„
Efficient VLA Models for Embodied Manipulation: A Systematic SurveyGuan et al.2025πŸ“„
VLA Models for Robotics: A Review Towards Real-World ApplicationsKawaharazuka et al.2025πŸ“„ Β· 🌐
An Anatomy of VLA Models: From Modules to Milestones and Challengesβ€”2025πŸ“„
Pure Vision-Language-Action Models: A Comprehensive Surveyβ€”2025πŸ“„
A Survey on Efficient Vision-Language-Action ModelsYu et al.2025πŸ“„
A Survey on VLA Models for Autonomous DrivingJiang et al.2025πŸ“„
VLA in Robotics: A Survey of Datasets, Benchmarks, and Data EnginesWang et al.2026πŸ“„

Foundation Models & Embodied AI

TitleAuthorsYearLinks
Foundation Models in Robotics: Applications, Challenges, and the FutureFiroozi et al.2023πŸ“„
Toward General-Purpose Robots via Foundation Models: A SurveyHu et al.2023πŸ“„
Aligning Cyber Space with Physical World: A Survey on Embodied AILiu et al.2024πŸ“„
What Foundation Models Can Bring for Robot Learning in Manipulation: A SurveyLi et al.2024πŸ“„
Deep Reinforcement Learning for Robotics: A Survey of Real-World SuccessesTang et al.2024πŸ“„
Generative AI in Robotic Manipulation: A SurveyZhang et al.2025πŸ“„
A Survey of Sim-to-Real Methods in RL with Foundation ModelsDa et al.2025πŸ“„
Behavior Foundation Model: Towards Next-Generation Whole-Body Control of HumanoidsYuan et al.2025πŸ“„
Towards a Unified Understanding of Robot Manipulation: A Comprehensive SurveyBai et al.2025πŸ“„
Robotic Foundation Models for Industrial Control: A Survey & Readiness AssessmentKube et al.2026πŸ“„

πŸ€– Vision-Language-Action (VLA) Models

Following the action-tokenization view (Zhong et al., 2025), the primary split is by how actions are represented; capability-oriented subsections (reasoning, 3D/4D, efficiency, RL) cut across it. A few pre-/non-VLM generalist policies (e.g., RT-1, Octo) are listed alongside their successors to show lineage β€” see Foundational Robot Policies for the strictly non-VLA baselines.

By Action Representation

Autoregressive / Discrete-Token VLA

Actions are binned into discrete tokens and decoded like text. Simple and VLM-native; high-frequency dexterity needs better tokenizers (see FAST).

ModelTitleYearLinks
VLA-0Building SOTA VLAs with Zero Modification2025πŸ“„ Β· 🌐
UniVLAUnified Vision-Language-Action Model (native multimodal tokens)2025πŸ“„ Β· 🌐
Ο€0-FASTAutoregressive Ο€0 variant using the FAST action tokenizer2025πŸ“„ Β· 🌐
OpenVLAAn Open-Source Vision-Language-Action Model2024πŸ“„ Β· 🌐 Β· πŸ’»
RT-2VLA Models Transfer Web Knowledge to Robotic Control2023πŸ“„ Β· 🌐
RT-1Robotics Transformer for Real-World Control at Scale2022πŸ“„ Β· 🌐 Β· πŸ’»

Diffusion-based VLA

A diffusion action head denoises continuous action chunks conditioned on vision-language features.

ModelTitleYearLinks
RoboVLMsTowards Generalist Robot Policies: What Matters in Building VLAs2024πŸ“„ Β· 🌐
CogACTA Foundational VLA Model for Synergizing Cognition and Action2024πŸ“„
TinyVLAFast, Data-Efficient VLA Models for Manipulation2024πŸ“„ Β· 🌐
OctoAn Open-Source Generalist Robot Policy2024πŸ“„ Β· 🌐 Β· πŸ’»

Flow-Matching VLA

A conditional flow/vector field transports noise to action chunks β€” the dominant head for current SOTA generalist VLAs.

ModelTitleYearLinks
Ο€*0.6A VLA That Learns From Experience2025πŸ“„ Β· 🌐
X-VLASoft-Prompted Transformer as a Scalable Cross-Embodiment VLA2025πŸ“„ Β· 🌐 Β· πŸ’»
SmolVLAA VLA for Affordable and Efficient Robotics2025πŸ“„ Β· πŸ’»
Ο€0.5A VLA with Open-World Generalization2025πŸ“„ Β· 🌐
Gemini RoboticsBringing AI into the Physical World2025πŸ“„ Β· 🌐
GR00T N1An Open Foundation Model for Generalist Humanoid Robots2025πŸ“„ Β· πŸ’»
EO-1An Open Unified Embodied Foundation Model (interleaved reasoning + acting)2025πŸ“„ Β· 🌐
GR-3Large-Scale Vision-Language-Action Model (Technical Report)2025πŸ“„
FLOWERDemocratizing Generalist Robot Policies with Efficient VLA Flow Policies2025πŸ“„
Ο€0A Vision-Language-Action Flow Model for General Robot Control2024πŸ“„ Β· 🌐

By Capability

Reasoning & Dual-System (Fast–Slow) VLA

Explicit chain-of-thought / embodied reasoning, or a slow System-2 planner paired with a fast System-1 controller.

ModelTitleYearLinks
ACoT-VLAAction Chain-of-Thought for VLA Models2026πŸ“„ Β· πŸ’»
Gemini Robotics 1.5Embodied Reasoning & Motion Transfer2025πŸ“„
ThinkActVLA Reasoning via Reinforced Visual Latent Planning2025πŸ“„
OpenHelixA Short Survey & Open-Source Dual-System VLA2025πŸ“„
FiS-VLAFast-in-Slow: A Dual-System Foundation Model for Unified Fast–Slow Reasoning2025πŸ“„
WALL-OSSIgniting VLMs toward the Embodied Space2025πŸ“„ Β· πŸ’»
CoT-VLAVisual Chain-of-Thought Reasoning for VLA2025πŸ“„

3D / 4D-Aware VLA

Policies that reason over explicit 3D/4D structure (point clouds, occupancy, predicted future frames) rather than 2D images alone. (VoxPoser, a zero-shot 3D value-map planner, lives under Foundational Robot Policies.)

ModelTitleYearLinks
3D-VLAA 3D Vision-Language-Action Generative World Model2024πŸ“„

Efficient & Real-Time VLA

Compression, caching, parallel decoding, and distillation to make VLAs small and fast enough for real-time / edge control (Guan et al., 2025).

ModelTitleYearLinks
FASTERRethinking Real-Time Flow VLAs2026πŸ“„
RTCReal-Time Chunking: Running VLAs at Real-Time Speed2025πŸ“„
NanoVLARouting-Decoupled VLA for Nano-Sized Generalist Policies2025πŸ“„
VLA-AdapterA Tiny-Scale VLA Paradigm2025πŸ“„
OpenVLA-OFTFine-Tuning VLAs: Optimizing Speed and Success2025πŸ“„ Β· 🌐
TinyVLAFast, Data-Efficient VLA Models2024πŸ“„ Β· 🌐

RL Fine-Tuning for VLA

Reinforcement learning (often on top of flow-/diffusion-based VLAs) to improve over imitation-only training.

ModelTitleYearLinks
Ο€_RLOnline RL Fine-Tuning for Flow-based VLAs2025πŸ“„
VLA-RFTRL Fine-Tuning with Verified Rewards in World Simulators2025πŸ“„
SimpleVLA-RLScaling VLA Training via Reinforcement Learning2025πŸ“„
ConRFTA Reinforced Fine-Tuning Method for VLA via Consistency Policy2025πŸ“„

🌎 World & World-Action Models

Organized by what the model predicts and how it is built, following the embodied-world-model taxonomy of Li et al., 2025 and the WAM split popularized by awesome-vla-wam.

General World Models

General-purpose models of environment dynamics β€” spanning classical latent world models for model-based RL (World Models, DreamerV3) and modern large-scale video / foundation world models β€” used for planning, neural simulation, or as backbones for WAMs.

ModelTitleYearLinks
Cosmos-Predict2.5World Simulation with Video Foundation Models for Physical AI2025πŸ“„ Β· πŸ’»
Cosmos-Reason1From Physical Common Sense to Embodied Reasoning2025πŸ“„
CosmosWorld Foundation Model Platform for Physical AI2025πŸ“„ Β· 🌐
V-JEPA 2Self-Supervised Video Models Enable Understanding, Prediction & Planning2025πŸ“„
iVideoGPTInteractive VideoGPTs are Scalable World Models2024πŸ“„
GenieGenerative Interactive Environments2024πŸ“„
DreamerV3Mastering Diverse Domains through World Models2023πŸ“„ Β· πŸ’»
UniSimLearning Interactive Real-World Simulators2023πŸ“„
World ModelsRecurrent latent world model + controller (origin of the term)2018πŸ“„

WAM from Video Generation

A (text-/image-conditioned) video generator imagines future frames; actions are recovered via an inverse-dynamics / action head.

ModelTitleYearLinks
DreamZeroWorld Action Models are Zero-shot Policies2026πŸ“„ Β· 🌐
DiT4DiTJointly Modeling Video Dynamics and Actions2026πŸ“„
Cosmos PolicyFine-Tuning Video Models for Visuomotor Control & Planning2026πŸ“„ Β· 🌐
Video2ActA Dual-System Video Diffusion Policy2025πŸ“„
GR-2A Generative Video-Language-Action Model with Web-Scale Knowledge2024πŸ“„
GR-1Large-Scale Video Generative Pre-training for Visual Robot Manipulation2023πŸ“„

WAM from VLMs

A pretrained VLM is turned into a world model (e.g., predicting goal images / object-centric futures) that then drives action.

ModelTitleYearLinks
DreamVLAA VLA Model Dreamed with Comprehensive World Knowledge2025πŸ“„
Goal-VLAImage-Generative VLMs as Object-Centric World Models for VLA2025πŸ“„

Unified VLA–World Models

Single architectures that jointly learn to act and to predict world dynamics, blurring the VLA/WAM boundary.

ModelTitleYearLinks
RynnVLA-002A Unified Vision-Language-Action and World Model2025πŸ“„ Β· πŸ’»
WholeBodyVLAUnified Latent VLA for Whole-Body Loco-Manipulation2025πŸ“„ Β· πŸ’»
WorldVLATowards an Autoregressive Action World Model2025πŸ“„ Β· πŸ’»

Latent & JEPA World Models

Self-supervised latent predictive models (non-reconstructive joint-embedding / JEPA). The JEPA foundations (I-JEPA) learn to predict in representation space; the action-conditioned variant (V-JEPA 2-AC) turns that prior into a world model for planning.

ModelTitleYearLinks
V-JEPA 2-ACAction-Conditioned Latent World Model for Zero-Shot Planning2025πŸ“„
I-JEPAImage-based Joint-Embedding Predictive Architecture (representation foundation)2023πŸ“„

Domain World Models (Driving & Navigation)

ModelTitleYearLinks
GAIA-2A Controllable Multi-View Generative World Model for Autonomous Driving2025πŸ“„
Navigation World ModelsConditional Diffusion Transformer for Navigation2024πŸ“„
GAIA-1A Generative World Model for Autonomous Driving2023πŸ“„

🧩 Action Representations & Tokenization

Building blocks shared across VLA and WAM policies β€” how continuous actions become learnable targets.

Discrete / Autoregressive Tokenizers

MethodTitleYearLinks
FASTEfficient (DCT-based) Action Tokenization for VLAs2025πŸ“„ Β· 🌐
BeTBehavior Transformers: Cloning k Modes with One Stone2022πŸ“„

Continuous & Chunked Action Policies

Heads that emit continuous action chunks β€” by denoising diffusion (Diffusion Policy) or by chunked sequence prediction with a CVAE (ACT). Flow-matching heads (Ο€0, SmolVLA, …) are listed with their models under Flow-Matching VLA.

MethodTitleYearLinks
Diffusion PolicyVisuomotor Policy Learning via Action Diffusion2023πŸ“„ Β· 🌐
ACT / ALOHAAction Chunking with Transformers2023πŸ“„ Β· 🌐

🦾 Foundational Robot Policies

Non-VLA policies and planners that remain standard baselines in the experimental tables of the papers above. (Diffusion Policy, ACT, and BeT are described under Action Representations.)

MethodTitleYearLinks
CrossFormerScaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion & Flight2024πŸ“„ Β· πŸ’»
RoboFlamingoVision-Language Foundation Models as Effective Robot Imitators2023πŸ“„ Β· πŸ’»
VoxPoserComposable 3D Value Maps for Robotic Manipulation (zero-shot LLM + 3D planner)2023πŸ“„ Β· 🌐
RT-1Robotics Transformer for Real-World Control at Scale2022πŸ“„

πŸ“¦ Resources

Datasets

NameDescriptionScaleLinks
Open X-EmbodimentCross-embodiment aggregation behind the RT-X models1M+ traj Β· 22 embodimentsπŸ“„ Β· 🌐
AgiBot WorldLarge-scale real-world manipulation (Colosseo)1M+ traj Β· 217 tasksπŸ“„ Β· 🌐
EgoScaleScaling dexterous manipulation with diverse egocentric human dataegocentric Β· 2026πŸ“„
DexCanvasHuman demos ↔ robot learning for dexterous manipulationdexterousπŸ“„
Galaxea Open-WorldMobile-bimanual dataset paired with the G0 dual-system VLA500 hrs Β· 150 tasksπŸ“„ Β· πŸ’»
DROIDIn-the-wild Franka manipulation across 3 continents76K traj Β· 564 scenesπŸ“„ Β· 🌐
RoboMINDMulti-embodiment teleop incl. labeled failures107K traj Β· 479 tasksπŸ“„ Β· 🌐
BridgeData V2WidowX manipulation w/ language + goal images60K traj Β· 24 envsπŸ“„ Β· 🌐
RH20TContact-rich skills w/ paired human demos110K+ seq Β· 147 tasksπŸ“„ Β· 🌐
Ego-Exo4DSimultaneous ego + exo video of skilled activity1,286 hrsπŸ“„ Β· 🌐
Ego4DMassive egocentric daily-life video3,670 hrsπŸ“„ Β· 🌐

Benchmarks

NameDescriptionLinks
LIBEROLifelong robot-learning, 130 manipulation tasks (de-facto VLA eval)πŸ“„ Β· πŸ’»
CALVINLong-horizon language-conditioned manipulationπŸ“„ Β· πŸ’»
SimplerEnvReal-to-sim evaluation for manipulation policiesπŸ“„ Β· 🌐
RoboCasaLarge-scale kitchen simulation (100 tasks)πŸ“„ Β· 🌐
VLABenchWorld-knowledge & long-horizon language tasksπŸ“„ Β· 🌐
ManiSkill3GPU-parallel manipulation (30K+ FPS)πŸ“„ Β· 🌐
THE COLOSSEUMRobustness under 14 environmental perturbationsπŸ“„ Β· 🌐
RoboArenaDistributed crowd-sourced real-world policy evalπŸ“„ Β· 🌐
RoboChallengeLarge-scale real-robot evaluation of embodied policiesπŸ“„
RobotArena ∞Scalable robot benchmarking via real-to-sim translationπŸ“„ Β· 🌐
WorldArenaPerception & functional-utility benchmark for embodied world modelsπŸ“„
Meta-World50 tabletop tasks for multi-task / meta-RLπŸ“„ Β· πŸ’»
RLBench100 hand-designed manipulation tasksπŸ“„ Β· πŸ’»

Simulation Platforms

NameDescriptionLinks
Isaac Sim / Isaac LabGPU-native robotics sim + RL/IL framework (Omniverse/USD)🌐
MuJoCo / MJXStandard rigid-body engine + JAX/XLA parallel variant🌐
GenesisGenerative, multi-solver physics platform (up to ~43M FPS)🌐
ManiSkillGPU-parallel manipulation simulator on SAPIEN🌐
SAPIENPart-level articulated-object simulator (PartNet-Mobility)πŸ“„ Β· 🌐
HabitatPhotorealistic indoor navigation & rearrangement🌐
ThreeDWorldMultimodal Unity3D sim (vision + audio + physics)πŸ“„ Β· 🌐
NewtonOpen, differentiable GPU physics engine (NVIDIA + DeepMind + Disney)🌐

Tools & Frameworks

NameDescriptionLinks
LeRobotEnd-to-end PyTorch robot-learning library + datasets + low-cost HWπŸ“„ Β· πŸ’»
openpiOpen models & training/inference for Ο€0, Ο€0-FAST, Ο€0.5πŸ’»
Isaac GR00TOpen humanoid foundation-model framework + checkpointsπŸ’»
OpenVLATraining / LoRA fine-tuning for the 7B OpenVLA modelπŸ’»
OctoJAX/Flax generalist transformer policy on OXEπŸ’»
robomimic / robosuiteLearning-from-demonstration framework + MuJoCo manipulation simπŸ’»
HIL-SERLHuman-in-the-loop, sample-efficient real-world RLπŸ’»

πŸ—‚οΈ Extended Paper Index (Auto-Curated, Newest First)

A broader, continuously-mined index of recent arXiv work that complements the curated highlights above β€” 182 additional papers, newest first. Last updated: 2026-09-30. Auto-generated from data/*.json by scripts/expand_papers.py; papers already highlighted above are omitted here to avoid duplication.

VLA β€” General & Manipulation Β· 45 papers
PaperAuthorsDateLinks
Online Evolution Strategy for Flow-Matching VLA Policies via Self-Supervised Trajectory Distribution OptimizationGongxin Yao, Yongsheng Zhao et al.2026-09-30
Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential GuidanceYanyan Zhang, Disheng Liu et al.2026-09-29
Taming VLAs under Robot Execution Errors: Self-Compensation and Stress TestingSohyun Lee, Yoonjae Baek et al.2026-09-29
Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent LookaheadJunghyun Kim, Ngseo Kim et al.2026-09-29
CoRe-VLA: Preserving Cross-View Coordination in VLAs under Camera ShiftsTianhang Pan, Xuanhao Wang et al.2026-09-29
VLALight: A Vision-Language-Action Model for Traffic Signal ControlPan Zhang, Siqi Lai et al.2026-09-29
Where Predictive Supervision Goes Shapes What VLA Policies LearnHanseul Kim, Jewon Yeom et al.2026-09-29
The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent InterfaceYuxiang Liu, Lizhi Yang et al.2026-09-28
FocusVLA: Focused Visual Utilization for Vision-Language-Action ModelsYichi Zhang, Weihao Yuan et al.2026-03-30
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic ManipulationHongyu Yan, Qiwei Li et al.2026-03-29
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and GenerationYang Liu, Pengxiang Ding et al.2026-03-26
ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-MakingYoung-Chae Son, Dae-Kwan Ko et al.2026-03-26
$Ο€$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial ManipulationJohnathan Tucker, Denis Liu et al.2026-03-26
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action ModelsJiaying Zhou, Zhihao Zhan et al.2026-03-25
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAsHaoran Yuan, Weigang Yi et al.2026-03-24
Gaze-Regularized Vision-Language-Action Models for Robotic ManipulationAnupam Pani, Yanchao Yang2026-03-24
CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action ModelsYouzhi Liu, Li Gao et al.2026-03-24
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action ModelsZhou Fang, Jiaqi Wang et al.2026-03-18
The Steep-spectrum Radio-loud AGN Luminosity Function and Its Implications for Black Hole Growth and Star FormationWenjie Wang, Zunli Yuan et al.2026-03-16
Building Explicit World Model for Zero-Shot Open-World Object ManipulationXiaotong Li, Gang Chen et al.2026-03-14
Beyond Dense Futures: World Models as Structured Planners for Robotic ManipulationMinghao Jin, Mozheng Liao et al.2026-03-13
Adaptive Capacity Allocation for Vision Language Action Fine-tuningDonghoon Kim, Minji Bae et al.2026-03-08
HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the WildZiyang Zhao, Shuheng Wang et al.2026-03-06
CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuningYike Zhang, Yaonan Wang et al.2026-02-13
HoloBrain-0 Technical ReportHorizon Robotics2026-02-07🌐
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and ExpansionRalf RΓΆmer, Yi Zhang et al.2026-01-14
Diverse stages of star formation in the IRAS 18162-2048 region. Emergence of UV FeedbackR. Fedriani, G. Anglada et al.2025-12-08
Mixture of Horizons in Action ChunkingDong Jing, Gang Wang et al.2025-11-24
A scaling relationship for non-thermal radio emission from ordered magnetospheres - II. Investigating the efficiency of relativistic electron production in magnetospheres of BA-type starsP. Leto, S. Owocki et al.2025-11-07
First X-ray and radio polarimetry of the neutron star low-mass X-ray binary GX 17+2Unnati Kashyap, Thomas J. Maccarone et al.2025-10-06
Deciphering the radio-star formation correlation on kpc scales. IV. Radio halos of highly-inclined Virgo cluster spiral galaxiesB. Vollmer, M. Soida et al.2025-10-03
Masses, Star-Formation Efficiencies, and Dynamical Evolution of 18,000 HII RegionsDebosmita Pathak, Adam K. Leroy et al.2025-09-26
X-ray and radio polarimetry of the neutron star low mass X-ray binary GX 13+1Unnati Kashyap, Thomas J. Maccarone et al.2025-08-07
Protostellar Outflows at the EarliesT Stages (POETS). VIII. The jets in the intermediate-mass star-forming region G105.42+9.88 (alias LkHΞ± 234)Luca Moscadelli, Fabrizio Massi et al.2025-08-05
Star formation histories and gas content limits of three ultra-faint dwarfs on the periphery of M31Michael G. Jones, David J. Sand et al.2025-08-01
Quenching Through Tidal Gas Removal: Molecular Gas and Star Formation in Tidal Tails of z ~ 0.7 Post-Starburst GalaxiesVincenzo R. D'Onofrio, Justin S. Spilker et al.2025-07-28
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile ManipulationSixiang Chen, Jiaming Liu et al.2025-07-02
The Radio Spectral Energy Distribution and Star Formation Calibration in MIGHTEE-COSMOS Highly Star-Forming Galaxies at 1.5 < z < 3.5Fatemeh Tabatabaei, Maryam Khademi et al.2025-06-19
Semi-empirical constraints on the HI mass function of star-forming galaxies and $Ξ©_{\rm HI}$ at $z\sim 0.37$ from interferometric surveysFrancesco Sinigaglia, Alessandro Bianchetti et al.2025-06-12
A persistent disk wind and variable jet outflow in the neutron-star low-mass X-ray binary GX 13+1Daniele Rogantini, Jeroen Homan et al.2025-04-07
X-ray and radio data obtained by XMM-Newton and VLA constrain the stellar wind of the magnetic quasi-Wolf-Rayet star in HD45166P. Leto, L. M. Oskinova et al.2025-03-10
The Arp 240 Galaxy Merger: A Detailed Look at the Molecular Kennicutt-Schmidt Star Formation Law on Sub-kpc ScalesAlejandro Saravia, Eduardo Rodas-Quito et al.2024-12-10
Runaway O and Be stars found using Gaia DR3, new stellar bow shocks and search for binariesM. Carretero-Castrillo, M. RibΓ³ et al.2024-12-10
A-VL: Adaptive Attention for Large Vision-Language ModelsJunyang Zhang, Mu Yuan et al.2024-09-23
HiRT: Enhancing Robotic Control with Hierarchical Robot TransformersJianke Zhang, Yanjiang Guo et al.2024-09-12
VLA β€” Reasoning, Planning & Dual-System Β· 5 papers
VLA β€” Autonomous Driving Β· 13 papers
PaperAuthorsDateLinks
Vision-Language-Action Autonomous Driving Agent with Language-based MemoryKai Yan, Xiangyu Chen et al.2026-09-29
Data-Efficient Adaptation of a Driving VLA to Class 8 TrucksSatyajeet Das, Aaron Buxbaum et al.2026-09-29
Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action PoliciesYaxin Zhao, Dianye Huang et al.2026-09-29
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early ObservationYiran Shi, Dongqi Guo et al.2026-03-30
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous DrivingQiqi Liu, Huan Xu et al.2026-03-28
Vega: Learning to Drive with Natural Language InstructionsSicheng Zuo, Yuxuan Li et al.2026-03-26
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized DrivingZehao Wang, Huaide Jiang et al.2026-03-26
ETA-VLA: Efficient Token Adaptation via Temporal Fusion and Intra-LLM Sparsification for Vision-Language-Action ModelsYiru Wang, Anqing Jiang et al.2026-03-26
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action ModelsJintao Cheng, Haozhe Wang et al.2026-03-24
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous DrivingZihan You, Hongwei Liu et al.2026-03-09
VLANeXt: Recipes for Building Strong VLA ModelsLiu, Xiangyu et al.2026-02-18🌐
A global view on star formation: The GLOSTAR Galactic plane survey XII. Effelsberg's continuum view and data releaseY. Gong, W. Reich et al.2025-12-17
Low-frequency spectra of neutron star + OB supergiant binaries: Does wind density drive persistent and flaring modes of accretion?J. van den Eijnden, L. Sidoli et al.2025-08-06
VLA β€” Dexterous & Humanoid Β· 2 papers
VLA β€” 3D / 4D & Spatial Β· 4 papers
VLA β€” RL & Post-Training Β· 7 papers
VLA β€” Efficient & Real-Time Β· 19 papers
PaperAuthorsDateLinks
Rho: A Foundation for Efficiently Adaptable VLA ModelsRho Team, Simran Bagaria et al.2026-09-29
Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA ControlZibo Wang, Haochen Han et al.2026-09-29
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLAYi Chen, Yuying Ge et al.2026-03-31
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and AccurateChen Yang, Yucheng Hu et al.2026-03-27
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow MatchingJiayi Chen, Wenxuan Song et al.2026-03-27
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time PerformanceWenxuan Song, Jiayi Chen et al.2026-03-26
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action ModelsPeiju Liu, Jinming Liu et al.2026-03-26
Agile-VLA: Few-Shot Industrial Pose Rectification via Implicit Affordance AnchoringTeng Yan, Zhengyang Pei et al.2026-03-24
Fast-WAM: Do World Action Models Need Test-time Future Imagination?Tianyuan Yuan, Zibin Dong et al.2026-03-17
FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic ManipulationYao Li, Peiyuan Tang et al.2026-02-27
Learning Native Continuation for Action Chunking Flow PoliciesYufeng Liu, Hang Yu et al.2026-02-13
Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action ModelsYuting Huang, Leilei Ding et al.2026-01-31
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token CachingYujie Wei, Jiahan Fan et al.2026-01-31
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic ManipulationWenda Yu, Tianshi Wang et al.2026-01-27
U-DiT Policy: U-shaped Diffusion Transformers for Robotic ManipulationLinzhi Wu, Aoran Mei et al.2025-09-29
Leave No Observation Behind: Real-time Correction for VLA Action ChunksKohei Sendai, Maxime Alvarez et al.2025-09-27
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel DecodingWenxuan Song, Jiayi Chen et al.2025-03-04
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language TuningZhiwei Hao, Jianyuan Guo et al.2024-10-23
VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language TasksYi-Lin Sung, Jaemin Cho et al.2021-12-13
VLA β€” Safety, Robustness & Evaluation Β· 12 papers
PaperAuthorsDateLinks
Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action AccelerationQiwei Chen, Kaijun Zhou et al.2026-09-29
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action ModelsXiyang Wu, Guangyao Shi et al.2026-03-26
SOMA: Strategic Orchestration and Memory-Augmented System for Vision-Language-Action Model Robustness via In-Context AdaptationZhuoran Li, Zhiyang Li et al.2026-03-25
ROBOGATE: Adaptive Failure Discovery for Safe Robot Policy Deployment via Two-Stage Boundary-Focused SamplingAzuki Kim2026-03-23
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action ModelsZhilong Zhang, Haoxiang Ren et al.2026-03-21
Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic ControlZunzhe Zhang, Runhan Huang et al.2026-03-18
World2Act: Latent Action Post-Training via Skill-Compositional World ModelsAn Dinh Vuong, Tuan Van Vo et al.2026-03-11
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action ModelYuanjie Lu, Beichen Wang et al.2026-03-09
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World ModelsXiaoquan Sun, Zetian Xu et al.2026-03-09
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action ModelsHyeongjun Heo, Seungyeon Woo et al.2026-03-06
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action ModelsHyeonbeom Choi, Daechul Ahn et al.2026-02-04
SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action ModelsBingxin Xu, Yuzhang Shang et al.2026-01-20
World Models β€” General & Foundation Β· 10 papers
World Models β€” Video Generation & WAM Β· 8 papers
World Models β€” Driving & Navigation Β· 6 papers
Policies β€” Diffusion & Flow Β· 27 papers
PaperAuthorsDateLinks
Encoding Predictability and Legibility for Style-Conditioned Diffusion PolicyAdrien Jacquet CrΓ©tides, Mouad Abrini et al.2026-03-17
ReMAP-DP: Reprojected Multi-view Aligned PointMaps for Diffusion PolicyXinzhang Yang, Renjun Wu et al.2026-03-16
REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement LearningZhaoyuan Gu, Yipu Chen et al.2026-03-14
PPGuide: Steering Diffusion Policies with Performance Predictive GuidanceZixing Wang, Devesh K. Jha et al.2026-03-11
SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot ManipulationYouqiang Gui, Yuxuan Zhou et al.2026-03-05
Diffusion Policy through Conditional Proximal Policy OptimizationBen Liu, Shunpeng Yang et al.2026-03-05
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion PolicyPengyuan Wu, Pingrui Zhang et al.2026-03-02
ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent ManipulationEnyi Wang, Wen Fan et al.2026-02-25
Preference Aligned Visuomotor Diffusion Policies for Deformable Object ManipulationMarco Moletta, Michael C. Welle et al.2026-02-10
SERFN: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing FlowsChenyu Yang, Denis Tarasov et al.2026-02-10
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic ManipulationYuxuan Hu, Xiangyu Chen et al.2026-02-07
Moving On, Even When You're Broken: Fail-Active Trajectory Generation via Diffusion Policies Conditioned on Embodiment and TaskGilberto G. Briscoe-Martinez, Yaashia Gautam et al.2026-02-02
RoDiF: Robust Direct Fine-Tuning of Diffusion Policies with Corrupted Human FeedbackAmitesh Vatsa, Zhixian Xie et al.2026-01-31
Self-Imitated Diffusion Policy for Efficient and Robust Visual NavigationRunhua Zhang, Junyi Hou et al.2026-01-30
Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion PoliciesCe Hao, Xuanran Zhai et al.2026-01-29
ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot ManipulationWeize Xie, Yi Ding et al.2026-01-19
Sparse ActionGen: Accelerating Diffusion Policy with Real-time PruningKangye Ji, Yuan Meng et al.2026-01-19
CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action SpaceBingyi Liu, Jinbo He et al.2026-01-09
Learning Diffusion Policy from Primitive Skills for Robot ManipulationZhihao Gu, Ming Yang et al.2026-01-05
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic ControlWonhyeok Choi, Shutong Ding et al.2026-01-05
Flexible Multitask Learning with Factorized Diffusion PolicyChaoqi Liu, Haonan Chen et al.2025-12-26
Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic ManipulationKangchen Lv, Mingrui Yu et al.2025-12-19
ISS Policy : Scalable Diffusion Policy with Implicit Scene SupervisionWenlong Xia, Jinhao Zhang et al.2025-12-17
Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic TasksAileen Liao, Dong-Ki Kim et al.2025-12-08
CAPE: Context-Aware Diffusion Policy Via Proximal Mode Expansion for Collision AvoidanceRui Heng Yang, Xuan Zhao et al.2025-11-27
Learning Diffusion Policies for Robotic Manipulation of Timber Joinery under Fabrication UncertaintySalma Mozaffari, Daniel Ruan et al.2025-11-21
UltraDP: Generalizable Carotid Ultrasound Scanning with Force-Aware Diffusion PolicyRuoqu Chen, Xiangjie Yan et al.2025-11-19
Policies β€” Imitation & Behavior Learning Β· 10 papers
Policies β€” Robot Learning & Manipulation Β· 14 papers
PaperAuthorsDateLinks
Learning Multi-View Spatial Reasoning from Cross-View RelationsSuchae Jeong, Jaehwi Song et al.2026-03-30
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory GenerationMotonari Kambara, Koki Seno et al.2026-03-26
Chunk-Boundary Artifact in Action-Chunked Generative Policies: A Noise-Sensitive Failure MechanismRui Wang2026-03-12
Real-Time Robot Execution with Masked Action ChunkingHaoxuan Wang, Gengyu Zhang et al.2026-01-27
Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse RewardJiarui Yang, Bin Zhu et al.2025-08-15
Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou et al.2025-07-10
Real-Time Execution of Action Chunking Flow PoliciesKevin Black, Manuel Y. Galliker et al.2025-06-09
Learning Bimanual Manipulation via Action Chunking and Inter-Arm Coordination with TransformersTomohiro Motoda, Ryo Hanai et al.2025-03-18
MissionGPT: Mission Planner for Mobile Robot based on Robotics Transformer ModelVladimir Berman, Artem Bazhenov et al.2024-11-07
VQ-ACE: Efficient Policy Search for Dexterous Robotic Manipulation via Action Chunking EmbeddingChenyu Yang, Davide Liconti et al.2024-11-05
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual ManipulationAndrew Lee, Ian Chuang et al.2024-09-12
Bringing the RT-1-X Foundation Model to a SCARA robotJonathan Salzer, Arnoud Visser2024-09-05
Logically Constrained Robotics Transformers for Enhanced Perception-Action PlanningParv Kapoor, Sai Vemprala et al.2024-08-09
SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust AttentionIsabel Leal, Krzysztof Choromanski et al.2023-12-04

πŸ“‹ Full Paper Index & Baselines

πŸ“Š Click to expand the complete paper list and baseline methods

The curated tables above highlight landmark and representative work. For the exhaustive, auto-maintained index and the baseline methods extracted from experimental tables, see:

Quick Reference (common baselines)

FamilyKey baselines
VLART-1, RT-2, OpenVLA, Octo, Ο€0, Ο€0.5, X-VLA, UniVLA, SmolVLA
PolicyDiffusion Policy, ACT, BeT, RoboFlamingo, CrossFormer
World ModelDreamerV3, I-JEPA, V-JEPA 2, Genie, Cosmos, GR-1/GR-2

🀝 Contributing

Contributions are very welcome! To add or fix a paper:

  1. Add a paper β€” open a PR placing it in the appropriate category (keep tables sorted newest-first), or open an issue with the arXiv link.
  2. Fix an error β€” submit a PR with the correction.
  3. New papers appear automatically β€” the πŸ†• Latest Papers section and the πŸ—‚οΈ Extended Paper Index are regenerated daily by the scraper; do not hand-edit content between the auto markers.

To run the discovery pipeline locally:

pip install -r requirements.txt
python scripts/arxiv_scraper.py --max-results 50 --days-back 30   # writes data/papers.json
python scripts/update_readme.py                                  # refreshes the πŸ†• auto section
python scripts/expand_papers.py                                  # refreshes the πŸ—‚οΈ extended index
python scripts/build_site.py                                     # rebuilds the GitHub Pages site (index.html)

License

Released under the MIT License.

Acknowledgments

Inspired by awesome-vla-wam, awesome-physical-ai, and awesome-vla-study. Taxonomy grounded in the surveys listed above.


If you find this repository useful, please consider giving it a ⭐
awesone
emobodied-ai
robotics
video-action-models
vla
world-action-model

Languages

Python

100.0%