A curated list of academic papers and resources on Vision-Language-Action (VLA) and World Action Models (WAM)
Python
35
202 commits
updated Oct 1, 2026
A curated, continuously-updated reading list of World Action Models (WAM), Vision-Language-Action (VLA) models, and Embodied AI β organized by a survey-grounded taxonomy.
π Website: hyperboliccurve.github.io/Awesome-World-Action-Model
The push toward general-purpose robots has produced two converging families of foundation models:
These two families overlap: a WAM built on a pretrained VLM is simultaneously a VLA and a WAM. This list maps that landscape with a taxonomy grounded in the recent survey literature (see Surveys), so each category has a clear, defensible scope rather than an ad-hoc label.
[!NOTE] Legend β π arXiv Β· π project page Β· π» code Β· π dataset/benchmark. Tables are sorted newest-first within each category. The π Latest Papers section is refreshed daily from arXiv by a GitHub Action; everything else is hand-curated.
flowchart TD
A[Robot Foundation Models] --> B[Vision-Language-Action<br/>VLA]
A --> C[World & World-Action Models<br/>WM / WAM]
A --> R[Action Representations]
A --> P[Foundational Policies]
B --> B1[By Action Representation:<br/>Autoregressive Β· Diffusion Β· Flow-Matching]
B --> B2[By Capability:<br/>Reasoning/Dual-System Β· 3D-4D Β· Efficient Β· RL Fine-Tuning]
C --> C1[Foundation / General World Models]
C --> C2[WAM from Video Generation]
C --> C3[WAM from VLMs]
C --> C4[WAM from Scratch Β· Latent / JEPA]
C --> C5[Domain: Driving Β· Navigation]
R --> R1[Discrete / Autoregressive Tokenizers]
R --> R2[Diffusion & Flow-Matching Policies]
| Term | Definition | Canonical reference |
|---|---|---|
| Vision-Language-Action (VLA) | A robot policy that adapts a pretrained VLM to map images + language instructions to actions. | RT-2 (Brohan et al., 2023) |
| World Model (WM) | A learned model that predicts future states of an environment (in pixels, latents, or 3D/4D), used for planning, simulation, or representation. | World Models (Ha & Schmidhuber, 2018) |
| World Action Model (WAM) | A policy that leverages world-modeling capability (predicting future states) for action prediction β typically by adapting a video / world-model backbone to emit actions. | GR-1 (Wu et al., 2023) |
[!IMPORTANT] VLA β© WAM. The families intersect: a WAM built on a pretrained VLM is both. The split in this list is by what prior the model starts from β VLM-style vision-language priors (VLA) vs. video/dynamics priors (WAM) β and, within VLA, by how actions are represented, the axis most surveys agree is the field's clearest discriminator.
Papers are automatically fetched daily from arXiv. Last updated: 2026-09-30
| Paper | Date | Code |
|---|---|---|
| UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision Wei Xue, Keliang Liu et al. | 2026-09-30 | |
| SteerQuant: Steering Quantization Error with Action-Guided Scaling in World-Action Models Yunhan Wang, Haodong Wang et al. | 2026-09-30 | |
| Rethinking Representations for World-Action Modeling Haoyi Jiang, Liu Liu et al. | 2026-09-29 | |
| MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation Wenbo Chen, Tianfu Li et al. | 2026-09-29 | |
| Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks Guoheng Sun, Chen Chen et al. | 2026-09-29 | |
| Revision, Not Restart: Revisable Visual Plans for Closed-Loop World-Action Models Pengyiang Liu, Junbo Niu et al. | 2026-09-28 |
Recent surveys that define the field and motivate the taxonomy used here.
| Title | Authors | Year | Links |
|---|---|---|---|
| Understanding World or Predicting Future? A Comprehensive Survey of World Models | Ding et al. | 2024 | π |
| A Comprehensive Survey on World Models for Embodied AI | Li et al. | 2025 | π |
| 3D and 4D World Modeling: A Survey | Kong et al. | 2025 | π |
| Learning Embodied Intelligence from Physical Simulators and World Models | Long et al. | 2025 | π |
| Embodied AI: From LLMs to World Models | Feng et al. | 2025 | π |
| World Model for Robot Learning: A Comprehensive Survey | Hou et al. | 2026 | π |
| Modeling the Mental World for Embodied AI: A Comprehensive Review | Liu et al. | 2026 | π |
| The Role of World Models in Shaping Autonomous Driving: A Survey | Tu et al. | 2025 | π |
| Title | Authors | Year | Links |
|---|---|---|---|
| A Survey on Vision-Language-Action Models for Embodied AI | Ma et al. | 2024 | π |
| A Survey on VLA Models: An Action Tokenization Perspective | Zhong et al. | 2025 | π |
| VLA Models: Concepts, Progress, Applications and Challenges | Sapkota et al. | 2025 | π |
| Large VLM-based VLA Models for Robotic Manipulation: A Survey | Shao et al. | 2025 | π |
| Efficient VLA Models for Embodied Manipulation: A Systematic Survey | Guan et al. | 2025 | π |
| VLA Models for Robotics: A Review Towards Real-World Applications | Kawaharazuka et al. | 2025 | π Β· π |
| An Anatomy of VLA Models: From Modules to Milestones and Challenges | β | 2025 | π |
| Pure Vision-Language-Action Models: A Comprehensive Survey | β | 2025 | π |
| A Survey on Efficient Vision-Language-Action Models | Yu et al. | 2025 | π |
| A Survey on VLA Models for Autonomous Driving | Jiang et al. | 2025 | π |
| VLA in Robotics: A Survey of Datasets, Benchmarks, and Data Engines | Wang et al. | 2026 | π |
| Title | Authors | Year | Links |
|---|---|---|---|
| Foundation Models in Robotics: Applications, Challenges, and the Future | Firoozi et al. | 2023 | π |
| Toward General-Purpose Robots via Foundation Models: A Survey | Hu et al. | 2023 | π |
| Aligning Cyber Space with Physical World: A Survey on Embodied AI | Liu et al. | 2024 | π |
| What Foundation Models Can Bring for Robot Learning in Manipulation: A Survey | Li et al. | 2024 | π |
| Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes | Tang et al. | 2024 | π |
| Generative AI in Robotic Manipulation: A Survey | Zhang et al. | 2025 | π |
| A Survey of Sim-to-Real Methods in RL with Foundation Models | Da et al. | 2025 | π |
| Behavior Foundation Model: Towards Next-Generation Whole-Body Control of Humanoids | Yuan et al. | 2025 | π |
| Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey | Bai et al. | 2025 | π |
| Robotic Foundation Models for Industrial Control: A Survey & Readiness Assessment | Kube et al. | 2026 | π |
Following the action-tokenization view (Zhong et al., 2025), the primary split is by how actions are represented; capability-oriented subsections (reasoning, 3D/4D, efficiency, RL) cut across it. A few pre-/non-VLM generalist policies (e.g., RT-1, Octo) are listed alongside their successors to show lineage β see Foundational Robot Policies for the strictly non-VLA baselines.
Actions are binned into discrete tokens and decoded like text. Simple and VLM-native; high-frequency dexterity needs better tokenizers (see FAST).
| Model | Title | Year | Links |
|---|---|---|---|
| VLA-0 | Building SOTA VLAs with Zero Modification | 2025 | π Β· π |
| UniVLA | Unified Vision-Language-Action Model (native multimodal tokens) | 2025 | π Β· π |
| Ο0-FAST | Autoregressive Ο0 variant using the FAST action tokenizer | 2025 | π Β· π |
| OpenVLA | An Open-Source Vision-Language-Action Model | 2024 | π Β· π Β· π» |
| RT-2 | VLA Models Transfer Web Knowledge to Robotic Control | 2023 | π Β· π |
| RT-1 | Robotics Transformer for Real-World Control at Scale | 2022 | π Β· π Β· π» |
A diffusion action head denoises continuous action chunks conditioned on vision-language features.
| Model | Title | Year | Links |
|---|---|---|---|
| RoboVLMs | Towards Generalist Robot Policies: What Matters in Building VLAs | 2024 | π Β· π |
| CogACT | A Foundational VLA Model for Synergizing Cognition and Action | 2024 | π |
| TinyVLA | Fast, Data-Efficient VLA Models for Manipulation | 2024 | π Β· π |
| Octo | An Open-Source Generalist Robot Policy | 2024 | π Β· π Β· π» |
A conditional flow/vector field transports noise to action chunks β the dominant head for current SOTA generalist VLAs.
| Model | Title | Year | Links |
|---|---|---|---|
| Ο*0.6 | A VLA That Learns From Experience | 2025 | π Β· π |
| X-VLA | Soft-Prompted Transformer as a Scalable Cross-Embodiment VLA | 2025 | π Β· π Β· π» |
| SmolVLA | A VLA for Affordable and Efficient Robotics | 2025 | π Β· π» |
| Ο0.5 | A VLA with Open-World Generalization | 2025 | π Β· π |
| Gemini Robotics | Bringing AI into the Physical World | 2025 | π Β· π |
| GR00T N1 | An Open Foundation Model for Generalist Humanoid Robots | 2025 | π Β· π» |
| EO-1 | An Open Unified Embodied Foundation Model (interleaved reasoning + acting) | 2025 | π Β· π |
| GR-3 | Large-Scale Vision-Language-Action Model (Technical Report) | 2025 | π |
| FLOWER | Democratizing Generalist Robot Policies with Efficient VLA Flow Policies | 2025 | π |
| Ο0 | A Vision-Language-Action Flow Model for General Robot Control | 2024 | π Β· π |
Explicit chain-of-thought / embodied reasoning, or a slow System-2 planner paired with a fast System-1 controller.
| Model | Title | Year | Links |
|---|---|---|---|
| ACoT-VLA | Action Chain-of-Thought for VLA Models | 2026 | π Β· π» |
| Gemini Robotics 1.5 | Embodied Reasoning & Motion Transfer | 2025 | π |
| ThinkAct | VLA Reasoning via Reinforced Visual Latent Planning | 2025 | π |
| OpenHelix | A Short Survey & Open-Source Dual-System VLA | 2025 | π |
| FiS-VLA | Fast-in-Slow: A Dual-System Foundation Model for Unified FastβSlow Reasoning | 2025 | π |
| WALL-OSS | Igniting VLMs toward the Embodied Space | 2025 | π Β· π» |
| CoT-VLA | Visual Chain-of-Thought Reasoning for VLA | 2025 | π |
Policies that reason over explicit 3D/4D structure (point clouds, occupancy, predicted future frames) rather than 2D images alone. (VoxPoser, a zero-shot 3D value-map planner, lives under Foundational Robot Policies.)
| Model | Title | Year | Links |
|---|---|---|---|
| 3D-VLA | A 3D Vision-Language-Action Generative World Model | 2024 | π |
Compression, caching, parallel decoding, and distillation to make VLAs small and fast enough for real-time / edge control (Guan et al., 2025).
| Model | Title | Year | Links |
|---|---|---|---|
| FASTER | Rethinking Real-Time Flow VLAs | 2026 | π |
| RTC | Real-Time Chunking: Running VLAs at Real-Time Speed | 2025 | π |
| NanoVLA | Routing-Decoupled VLA for Nano-Sized Generalist Policies | 2025 | π |
| VLA-Adapter | A Tiny-Scale VLA Paradigm | 2025 | π |
| OpenVLA-OFT | Fine-Tuning VLAs: Optimizing Speed and Success | 2025 | π Β· π |
| TinyVLA | Fast, Data-Efficient VLA Models | 2024 | π Β· π |
Reinforcement learning (often on top of flow-/diffusion-based VLAs) to improve over imitation-only training.
| Model | Title | Year | Links |
|---|---|---|---|
| Ο_RL | Online RL Fine-Tuning for Flow-based VLAs | 2025 | π |
| VLA-RFT | RL Fine-Tuning with Verified Rewards in World Simulators | 2025 | π |
| SimpleVLA-RL | Scaling VLA Training via Reinforcement Learning | 2025 | π |
| ConRFT | A Reinforced Fine-Tuning Method for VLA via Consistency Policy | 2025 | π |
Organized by what the model predicts and how it is built, following the embodied-world-model taxonomy of Li et al., 2025 and the WAM split popularized by awesome-vla-wam.
General-purpose models of environment dynamics β spanning classical latent world models for model-based RL (World Models, DreamerV3) and modern large-scale video / foundation world models β used for planning, neural simulation, or as backbones for WAMs.
| Model | Title | Year | Links |
|---|---|---|---|
| Cosmos-Predict2.5 | World Simulation with Video Foundation Models for Physical AI | 2025 | π Β· π» |
| Cosmos-Reason1 | From Physical Common Sense to Embodied Reasoning | 2025 | π |
| Cosmos | World Foundation Model Platform for Physical AI | 2025 | π Β· π |
| V-JEPA 2 | Self-Supervised Video Models Enable Understanding, Prediction & Planning | 2025 | π |
| iVideoGPT | Interactive VideoGPTs are Scalable World Models | 2024 | π |
| Genie | Generative Interactive Environments | 2024 | π |
| DreamerV3 | Mastering Diverse Domains through World Models | 2023 | π Β· π» |
| UniSim | Learning Interactive Real-World Simulators | 2023 | π |
| World Models | Recurrent latent world model + controller (origin of the term) | 2018 | π |
A (text-/image-conditioned) video generator imagines future frames; actions are recovered via an inverse-dynamics / action head.
| Model | Title | Year | Links |
|---|---|---|---|
| DreamZero | World Action Models are Zero-shot Policies | 2026 | π Β· π |
| DiT4DiT | Jointly Modeling Video Dynamics and Actions | 2026 | π |
| Cosmos Policy | Fine-Tuning Video Models for Visuomotor Control & Planning | 2026 | π Β· π |
| Video2Act | A Dual-System Video Diffusion Policy | 2025 | π |
| GR-2 | A Generative Video-Language-Action Model with Web-Scale Knowledge | 2024 | π |
| GR-1 | Large-Scale Video Generative Pre-training for Visual Robot Manipulation | 2023 | π |
A pretrained VLM is turned into a world model (e.g., predicting goal images / object-centric futures) that then drives action.
| Model | Title | Year | Links |
|---|---|---|---|
| DreamVLA | A VLA Model Dreamed with Comprehensive World Knowledge | 2025 | π |
| Goal-VLA | Image-Generative VLMs as Object-Centric World Models for VLA | 2025 | π |
Single architectures that jointly learn to act and to predict world dynamics, blurring the VLA/WAM boundary.
| Model | Title | Year | Links |
|---|---|---|---|
| RynnVLA-002 | A Unified Vision-Language-Action and World Model | 2025 | π Β· π» |
| WholeBodyVLA | Unified Latent VLA for Whole-Body Loco-Manipulation | 2025 | π Β· π» |
| WorldVLA | Towards an Autoregressive Action World Model | 2025 | π Β· π» |
Self-supervised latent predictive models (non-reconstructive joint-embedding / JEPA). The JEPA foundations (I-JEPA) learn to predict in representation space; the action-conditioned variant (V-JEPA 2-AC) turns that prior into a world model for planning.
| Model | Title | Year | Links |
|---|---|---|---|
| V-JEPA 2-AC | Action-Conditioned Latent World Model for Zero-Shot Planning | 2025 | π |
| I-JEPA | Image-based Joint-Embedding Predictive Architecture (representation foundation) | 2023 | π |
| Model | Title | Year | Links |
|---|---|---|---|
| GAIA-2 | A Controllable Multi-View Generative World Model for Autonomous Driving | 2025 | π |
| Navigation World Models | Conditional Diffusion Transformer for Navigation | 2024 | π |
| GAIA-1 | A Generative World Model for Autonomous Driving | 2023 | π |
Building blocks shared across VLA and WAM policies β how continuous actions become learnable targets.
| Method | Title | Year | Links |
|---|---|---|---|
| FAST | Efficient (DCT-based) Action Tokenization for VLAs | 2025 | π Β· π |
| BeT | Behavior Transformers: Cloning k Modes with One Stone | 2022 | π |
Heads that emit continuous action chunks β by denoising diffusion (Diffusion Policy) or by chunked sequence prediction with a CVAE (ACT). Flow-matching heads (Ο0, SmolVLA, β¦) are listed with their models under Flow-Matching VLA.
| Method | Title | Year | Links |
|---|---|---|---|
| Diffusion Policy | Visuomotor Policy Learning via Action Diffusion | 2023 | π Β· π |
| ACT / ALOHA | Action Chunking with Transformers | 2023 | π Β· π |
Non-VLA policies and planners that remain standard baselines in the experimental tables of the papers above. (Diffusion Policy, ACT, and BeT are described under Action Representations.)
| Method | Title | Year | Links |
|---|---|---|---|
| CrossFormer | Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion & Flight | 2024 | π Β· π» |
| RoboFlamingo | Vision-Language Foundation Models as Effective Robot Imitators | 2023 | π Β· π» |
| VoxPoser | Composable 3D Value Maps for Robotic Manipulation (zero-shot LLM + 3D planner) | 2023 | π Β· π |
| RT-1 | Robotics Transformer for Real-World Control at Scale | 2022 | π |
| Name | Description | Scale | Links |
|---|---|---|---|
| Open X-Embodiment | Cross-embodiment aggregation behind the RT-X models | 1M+ traj Β· 22 embodiments | π Β· π |
| AgiBot World | Large-scale real-world manipulation (Colosseo) | 1M+ traj Β· 217 tasks | π Β· π |
| EgoScale | Scaling dexterous manipulation with diverse egocentric human data | egocentric Β· 2026 | π |
| DexCanvas | Human demos β robot learning for dexterous manipulation | dexterous | π |
| Galaxea Open-World | Mobile-bimanual dataset paired with the G0 dual-system VLA | 500 hrs Β· 150 tasks | π Β· π» |
| DROID | In-the-wild Franka manipulation across 3 continents | 76K traj Β· 564 scenes | π Β· π |
| RoboMIND | Multi-embodiment teleop incl. labeled failures | 107K traj Β· 479 tasks | π Β· π |
| BridgeData V2 | WidowX manipulation w/ language + goal images | 60K traj Β· 24 envs | π Β· π |
| RH20T | Contact-rich skills w/ paired human demos | 110K+ seq Β· 147 tasks | π Β· π |
| Ego-Exo4D | Simultaneous ego + exo video of skilled activity | 1,286 hrs | π Β· π |
| Ego4D | Massive egocentric daily-life video | 3,670 hrs | π Β· π |
| Name | Description | Links |
|---|---|---|
| LIBERO | Lifelong robot-learning, 130 manipulation tasks (de-facto VLA eval) | π Β· π» |
| CALVIN | Long-horizon language-conditioned manipulation | π Β· π» |
| SimplerEnv | Real-to-sim evaluation for manipulation policies | π Β· π |
| RoboCasa | Large-scale kitchen simulation (100 tasks) | π Β· π |
| VLABench | World-knowledge & long-horizon language tasks | π Β· π |
| ManiSkill3 | GPU-parallel manipulation (30K+ FPS) | π Β· π |
| THE COLOSSEUM | Robustness under 14 environmental perturbations | π Β· π |
| RoboArena | Distributed crowd-sourced real-world policy eval | π Β· π |
| RoboChallenge | Large-scale real-robot evaluation of embodied policies | π |
| RobotArena β | Scalable robot benchmarking via real-to-sim translation | π Β· π |
| WorldArena | Perception & functional-utility benchmark for embodied world models | π |
| Meta-World | 50 tabletop tasks for multi-task / meta-RL | π Β· π» |
| RLBench | 100 hand-designed manipulation tasks | π Β· π» |
| Name | Description | Links |
|---|---|---|
| Isaac Sim / Isaac Lab | GPU-native robotics sim + RL/IL framework (Omniverse/USD) | π |
| MuJoCo / MJX | Standard rigid-body engine + JAX/XLA parallel variant | π |
| Genesis | Generative, multi-solver physics platform (up to ~43M FPS) | π |
| ManiSkill | GPU-parallel manipulation simulator on SAPIEN | π |
| SAPIEN | Part-level articulated-object simulator (PartNet-Mobility) | π Β· π |
| Habitat | Photorealistic indoor navigation & rearrangement | π |
| ThreeDWorld | Multimodal Unity3D sim (vision + audio + physics) | π Β· π |
| Newton | Open, differentiable GPU physics engine (NVIDIA + DeepMind + Disney) | π |
| Name | Description | Links |
|---|---|---|
| LeRobot | End-to-end PyTorch robot-learning library + datasets + low-cost HW | π Β· π» |
| openpi | Open models & training/inference for Ο0, Ο0-FAST, Ο0.5 | π» |
| Isaac GR00T | Open humanoid foundation-model framework + checkpoints | π» |
| OpenVLA | Training / LoRA fine-tuning for the 7B OpenVLA model | π» |
| Octo | JAX/Flax generalist transformer policy on OXE | π» |
| robomimic / robosuite | Learning-from-demonstration framework + MuJoCo manipulation sim | π» |
| HIL-SERL | Human-in-the-loop, sample-efficient real-world RL | π» |
A broader, continuously-mined index of recent arXiv work that complements the curated highlights above β 182 additional papers, newest first. Last updated: 2026-09-30. Auto-generated from
data/*.jsonbyscripts/expand_papers.py; papers already highlighted above are omitted here to avoid duplication.
| Paper | Authors | Date | Links |
|---|---|---|---|
| Do World Action Models Generalize Better than VLAs? A Robustness Study | Zhanguang Zhang, Zhiyuan Li et al. | 2026-03-23 | |
| Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models | Riccardo Andrea Izzo, Gianluca Bardaro et al. | 2026-03-05 | |
| Chain of World: World Model Thinking in Latent Motion | Fuxiang Yang, Donglin Di et al. | 2026-03-03 | |
| FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment | Han Zhao, Jingbo Wang et al. | 2026-02-19 | |
| VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation | Changhua Xu, Jie Lu et al. | 2026-02-07 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation | Jade Choghari, Pepijn Kooijmans et al. | 2026-09-29 | |
| Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models | Ruixing Jin, Zicheng Zhu et al. | 2026-03-24 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| Beyond Token Importance: Preserving Spatial Scaffolds for Efficient Vision-Language-Action Inference | Jiayu Chen, Shuyong Gao et al. | 2026-09-29 | |
| AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations | Rui Huang, Yanlin Mu et al. | 2026-09-29 | |
| LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior | Xinkai Wang, Chenyi Wang et al. | 2026-03-26 | |
| 3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models | Bin Yu, Shijie Lian et al. | 2026-03-25 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning | Ruixiao Xu, Wong Lik Hang Kenny et al. | 2026-09-29 | |
| StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks | Ziyi Yin, Sangmin Woo et al. | 2026-09-28 | |
| VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation | Zhide Zhong, Haodong Yan et al. | 2026-03-27 | |
| On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning | Changyu Liu, Yiyang Liu et al. | 2026-01-11 | |
| VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning | Si-Cheng Wang, Tian-Yu Xiang et al. | 2025-09-30 | |
| The arc-shaped radio source at the center of NGC 6334A: Is it a colliding wind region of two young massive stars or the bow shock of a runaway star? | Vanessa Yanza, Sergio A. Dzib et al. | 2025-02-24 | |
| VLA 22 GHz Imaging of Massive Star Formation in Local Wolf-Rayet Galaxies | Nicholas G. Ferraro, Jean L. Turner et al. | 2024-11-09 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving | Junwei You, Weizhe Tang et al. | 2026-09-29 | |
| EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning | Yichao Liang, Amber Li et al. | 2026-09-28 | |
| Enhancing Policy Learning with World-Action Model | Yuci Han, Alper Yilmaz | 2026-03-30 | |
| Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving | Linbo Wang, Yupeng Zheng et al. | 2026-03-25 | |
| NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation | Tianshuai Hu, Zeying Gong et al. | 2026-03-16 | |
| AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation | Ge Yuan, Qiyuan Qiao et al. | 2026-02-23 |
The curated tables above highlight landmark and representative work. For the exhaustive, auto-maintained index and the baseline methods extracted from experimental tables, see:
| Family | Key baselines |
|---|---|
| VLA | RT-1, RT-2, OpenVLA, Octo, Ο0, Ο0.5, X-VLA, UniVLA, SmolVLA |
| Policy | Diffusion Policy, ACT, BeT, RoboFlamingo, CrossFormer |
| World Model | DreamerV3, I-JEPA, V-JEPA 2, Genie, Cosmos, GR-1/GR-2 |
Contributions are very welcome! To add or fix a paper:
To run the discovery pipeline locally:
pip install -r requirements.txt
python scripts/arxiv_scraper.py --max-results 50 --days-back 30 # writes data/papers.json
python scripts/update_readme.py # refreshes the π auto section
python scripts/expand_papers.py # refreshes the ποΈ extended index
python scripts/build_site.py # rebuilds the GitHub Pages site (index.html)
Released under the MIT License.
Inspired by awesome-vla-wam, awesome-physical-ai, and awesome-vla-study. Taxonomy grounded in the surveys listed above.
Python
100.0%
A curated list of academic papers and resources on Vision-Language-Action (VLA) and World Action Models (WAM)
Python
35
202 commits
updated Oct 1, 2026
A curated, continuously-updated reading list of World Action Models (WAM), Vision-Language-Action (VLA) models, and Embodied AI β organized by a survey-grounded taxonomy.
π Website: hyperboliccurve.github.io/Awesome-World-Action-Model
The push toward general-purpose robots has produced two converging families of foundation models:
These two families overlap: a WAM built on a pretrained VLM is simultaneously a VLA and a WAM. This list maps that landscape with a taxonomy grounded in the recent survey literature (see Surveys), so each category has a clear, defensible scope rather than an ad-hoc label.
[!NOTE] Legend β π arXiv Β· π project page Β· π» code Β· π dataset/benchmark. Tables are sorted newest-first within each category. The π Latest Papers section is refreshed daily from arXiv by a GitHub Action; everything else is hand-curated.
flowchart TD
A[Robot Foundation Models] --> B[Vision-Language-Action<br/>VLA]
A --> C[World & World-Action Models<br/>WM / WAM]
A --> R[Action Representations]
A --> P[Foundational Policies]
B --> B1[By Action Representation:<br/>Autoregressive Β· Diffusion Β· Flow-Matching]
B --> B2[By Capability:<br/>Reasoning/Dual-System Β· 3D-4D Β· Efficient Β· RL Fine-Tuning]
C --> C1[Foundation / General World Models]
C --> C2[WAM from Video Generation]
C --> C3[WAM from VLMs]
C --> C4[WAM from Scratch Β· Latent / JEPA]
C --> C5[Domain: Driving Β· Navigation]
R --> R1[Discrete / Autoregressive Tokenizers]
R --> R2[Diffusion & Flow-Matching Policies]
| Term | Definition | Canonical reference |
|---|---|---|
| Vision-Language-Action (VLA) | A robot policy that adapts a pretrained VLM to map images + language instructions to actions. | RT-2 (Brohan et al., 2023) |
| World Model (WM) | A learned model that predicts future states of an environment (in pixels, latents, or 3D/4D), used for planning, simulation, or representation. | World Models (Ha & Schmidhuber, 2018) |
| World Action Model (WAM) | A policy that leverages world-modeling capability (predicting future states) for action prediction β typically by adapting a video / world-model backbone to emit actions. | GR-1 (Wu et al., 2023) |
[!IMPORTANT] VLA β© WAM. The families intersect: a WAM built on a pretrained VLM is both. The split in this list is by what prior the model starts from β VLM-style vision-language priors (VLA) vs. video/dynamics priors (WAM) β and, within VLA, by how actions are represented, the axis most surveys agree is the field's clearest discriminator.
Papers are automatically fetched daily from arXiv. Last updated: 2026-09-30
| Paper | Date | Code |
|---|---|---|
| UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision Wei Xue, Keliang Liu et al. | 2026-09-30 | |
| SteerQuant: Steering Quantization Error with Action-Guided Scaling in World-Action Models Yunhan Wang, Haodong Wang et al. | 2026-09-30 | |
| Rethinking Representations for World-Action Modeling Haoyi Jiang, Liu Liu et al. | 2026-09-29 | |
| MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation Wenbo Chen, Tianfu Li et al. | 2026-09-29 | |
| Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks Guoheng Sun, Chen Chen et al. | 2026-09-29 | |
| Revision, Not Restart: Revisable Visual Plans for Closed-Loop World-Action Models Pengyiang Liu, Junbo Niu et al. | 2026-09-28 |
Recent surveys that define the field and motivate the taxonomy used here.
| Title | Authors | Year | Links |
|---|---|---|---|
| Understanding World or Predicting Future? A Comprehensive Survey of World Models | Ding et al. | 2024 | π |
| A Comprehensive Survey on World Models for Embodied AI | Li et al. | 2025 | π |
| 3D and 4D World Modeling: A Survey | Kong et al. | 2025 | π |
| Learning Embodied Intelligence from Physical Simulators and World Models | Long et al. | 2025 | π |
| Embodied AI: From LLMs to World Models | Feng et al. | 2025 | π |
| World Model for Robot Learning: A Comprehensive Survey | Hou et al. | 2026 | π |
| Modeling the Mental World for Embodied AI: A Comprehensive Review | Liu et al. | 2026 | π |
| The Role of World Models in Shaping Autonomous Driving: A Survey | Tu et al. | 2025 | π |
| Title | Authors | Year | Links |
|---|---|---|---|
| A Survey on Vision-Language-Action Models for Embodied AI | Ma et al. | 2024 | π |
| A Survey on VLA Models: An Action Tokenization Perspective | Zhong et al. | 2025 | π |
| VLA Models: Concepts, Progress, Applications and Challenges | Sapkota et al. | 2025 | π |
| Large VLM-based VLA Models for Robotic Manipulation: A Survey | Shao et al. | 2025 | π |
| Efficient VLA Models for Embodied Manipulation: A Systematic Survey | Guan et al. | 2025 | π |
| VLA Models for Robotics: A Review Towards Real-World Applications | Kawaharazuka et al. | 2025 | π Β· π |
| An Anatomy of VLA Models: From Modules to Milestones and Challenges | β | 2025 | π |
| Pure Vision-Language-Action Models: A Comprehensive Survey | β | 2025 | π |
| A Survey on Efficient Vision-Language-Action Models | Yu et al. | 2025 | π |
| A Survey on VLA Models for Autonomous Driving | Jiang et al. | 2025 | π |
| VLA in Robotics: A Survey of Datasets, Benchmarks, and Data Engines | Wang et al. | 2026 | π |
| Title | Authors | Year | Links |
|---|---|---|---|
| Foundation Models in Robotics: Applications, Challenges, and the Future | Firoozi et al. | 2023 | π |
| Toward General-Purpose Robots via Foundation Models: A Survey | Hu et al. | 2023 | π |
| Aligning Cyber Space with Physical World: A Survey on Embodied AI | Liu et al. | 2024 | π |
| What Foundation Models Can Bring for Robot Learning in Manipulation: A Survey | Li et al. | 2024 | π |
| Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes | Tang et al. | 2024 | π |
| Generative AI in Robotic Manipulation: A Survey | Zhang et al. | 2025 | π |
| A Survey of Sim-to-Real Methods in RL with Foundation Models | Da et al. | 2025 | π |
| Behavior Foundation Model: Towards Next-Generation Whole-Body Control of Humanoids | Yuan et al. | 2025 | π |
| Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey | Bai et al. | 2025 | π |
| Robotic Foundation Models for Industrial Control: A Survey & Readiness Assessment | Kube et al. | 2026 | π |
Following the action-tokenization view (Zhong et al., 2025), the primary split is by how actions are represented; capability-oriented subsections (reasoning, 3D/4D, efficiency, RL) cut across it. A few pre-/non-VLM generalist policies (e.g., RT-1, Octo) are listed alongside their successors to show lineage β see Foundational Robot Policies for the strictly non-VLA baselines.
Actions are binned into discrete tokens and decoded like text. Simple and VLM-native; high-frequency dexterity needs better tokenizers (see FAST).
| Model | Title | Year | Links |
|---|---|---|---|
| VLA-0 | Building SOTA VLAs with Zero Modification | 2025 | π Β· π |
| UniVLA | Unified Vision-Language-Action Model (native multimodal tokens) | 2025 | π Β· π |
| Ο0-FAST | Autoregressive Ο0 variant using the FAST action tokenizer | 2025 | π Β· π |
| OpenVLA | An Open-Source Vision-Language-Action Model | 2024 | π Β· π Β· π» |
| RT-2 | VLA Models Transfer Web Knowledge to Robotic Control | 2023 | π Β· π |
| RT-1 | Robotics Transformer for Real-World Control at Scale | 2022 | π Β· π Β· π» |
A diffusion action head denoises continuous action chunks conditioned on vision-language features.
| Model | Title | Year | Links |
|---|---|---|---|
| RoboVLMs | Towards Generalist Robot Policies: What Matters in Building VLAs | 2024 | π Β· π |
| CogACT | A Foundational VLA Model for Synergizing Cognition and Action | 2024 | π |
| TinyVLA | Fast, Data-Efficient VLA Models for Manipulation | 2024 | π Β· π |
| Octo | An Open-Source Generalist Robot Policy | 2024 | π Β· π Β· π» |
A conditional flow/vector field transports noise to action chunks β the dominant head for current SOTA generalist VLAs.
| Model | Title | Year | Links |
|---|---|---|---|
| Ο*0.6 | A VLA That Learns From Experience | 2025 | π Β· π |
| X-VLA | Soft-Prompted Transformer as a Scalable Cross-Embodiment VLA | 2025 | π Β· π Β· π» |
| SmolVLA | A VLA for Affordable and Efficient Robotics | 2025 | π Β· π» |
| Ο0.5 | A VLA with Open-World Generalization | 2025 | π Β· π |
| Gemini Robotics | Bringing AI into the Physical World | 2025 | π Β· π |
| GR00T N1 | An Open Foundation Model for Generalist Humanoid Robots | 2025 | π Β· π» |
| EO-1 | An Open Unified Embodied Foundation Model (interleaved reasoning + acting) | 2025 | π Β· π |
| GR-3 | Large-Scale Vision-Language-Action Model (Technical Report) | 2025 | π |
| FLOWER | Democratizing Generalist Robot Policies with Efficient VLA Flow Policies | 2025 | π |
| Ο0 | A Vision-Language-Action Flow Model for General Robot Control | 2024 | π Β· π |
Explicit chain-of-thought / embodied reasoning, or a slow System-2 planner paired with a fast System-1 controller.
| Model | Title | Year | Links |
|---|---|---|---|
| ACoT-VLA | Action Chain-of-Thought for VLA Models | 2026 | π Β· π» |
| Gemini Robotics 1.5 | Embodied Reasoning & Motion Transfer | 2025 | π |
| ThinkAct | VLA Reasoning via Reinforced Visual Latent Planning | 2025 | π |
| OpenHelix | A Short Survey & Open-Source Dual-System VLA | 2025 | π |
| FiS-VLA | Fast-in-Slow: A Dual-System Foundation Model for Unified FastβSlow Reasoning | 2025 | π |
| WALL-OSS | Igniting VLMs toward the Embodied Space | 2025 | π Β· π» |
| CoT-VLA | Visual Chain-of-Thought Reasoning for VLA | 2025 | π |
Policies that reason over explicit 3D/4D structure (point clouds, occupancy, predicted future frames) rather than 2D images alone. (VoxPoser, a zero-shot 3D value-map planner, lives under Foundational Robot Policies.)
| Model | Title | Year | Links |
|---|---|---|---|
| 3D-VLA | A 3D Vision-Language-Action Generative World Model | 2024 | π |
Compression, caching, parallel decoding, and distillation to make VLAs small and fast enough for real-time / edge control (Guan et al., 2025).
| Model | Title | Year | Links |
|---|---|---|---|
| FASTER | Rethinking Real-Time Flow VLAs | 2026 | π |
| RTC | Real-Time Chunking: Running VLAs at Real-Time Speed | 2025 | π |
| NanoVLA | Routing-Decoupled VLA for Nano-Sized Generalist Policies | 2025 | π |
| VLA-Adapter | A Tiny-Scale VLA Paradigm | 2025 | π |
| OpenVLA-OFT | Fine-Tuning VLAs: Optimizing Speed and Success | 2025 | π Β· π |
| TinyVLA | Fast, Data-Efficient VLA Models | 2024 | π Β· π |
Reinforcement learning (often on top of flow-/diffusion-based VLAs) to improve over imitation-only training.
| Model | Title | Year | Links |
|---|---|---|---|
| Ο_RL | Online RL Fine-Tuning for Flow-based VLAs | 2025 | π |
| VLA-RFT | RL Fine-Tuning with Verified Rewards in World Simulators | 2025 | π |
| SimpleVLA-RL | Scaling VLA Training via Reinforcement Learning | 2025 | π |
| ConRFT | A Reinforced Fine-Tuning Method for VLA via Consistency Policy | 2025 | π |
Organized by what the model predicts and how it is built, following the embodied-world-model taxonomy of Li et al., 2025 and the WAM split popularized by awesome-vla-wam.
General-purpose models of environment dynamics β spanning classical latent world models for model-based RL (World Models, DreamerV3) and modern large-scale video / foundation world models β used for planning, neural simulation, or as backbones for WAMs.
| Model | Title | Year | Links |
|---|---|---|---|
| Cosmos-Predict2.5 | World Simulation with Video Foundation Models for Physical AI | 2025 | π Β· π» |
| Cosmos-Reason1 | From Physical Common Sense to Embodied Reasoning | 2025 | π |
| Cosmos | World Foundation Model Platform for Physical AI | 2025 | π Β· π |
| V-JEPA 2 | Self-Supervised Video Models Enable Understanding, Prediction & Planning | 2025 | π |
| iVideoGPT | Interactive VideoGPTs are Scalable World Models | 2024 | π |
| Genie | Generative Interactive Environments | 2024 | π |
| DreamerV3 | Mastering Diverse Domains through World Models | 2023 | π Β· π» |
| UniSim | Learning Interactive Real-World Simulators | 2023 | π |
| World Models | Recurrent latent world model + controller (origin of the term) | 2018 | π |
A (text-/image-conditioned) video generator imagines future frames; actions are recovered via an inverse-dynamics / action head.
| Model | Title | Year | Links |
|---|---|---|---|
| DreamZero | World Action Models are Zero-shot Policies | 2026 | π Β· π |
| DiT4DiT | Jointly Modeling Video Dynamics and Actions | 2026 | π |
| Cosmos Policy | Fine-Tuning Video Models for Visuomotor Control & Planning | 2026 | π Β· π |
| Video2Act | A Dual-System Video Diffusion Policy | 2025 | π |
| GR-2 | A Generative Video-Language-Action Model with Web-Scale Knowledge | 2024 | π |
| GR-1 | Large-Scale Video Generative Pre-training for Visual Robot Manipulation | 2023 | π |
A pretrained VLM is turned into a world model (e.g., predicting goal images / object-centric futures) that then drives action.
| Model | Title | Year | Links |
|---|---|---|---|
| DreamVLA | A VLA Model Dreamed with Comprehensive World Knowledge | 2025 | π |
| Goal-VLA | Image-Generative VLMs as Object-Centric World Models for VLA | 2025 | π |
Single architectures that jointly learn to act and to predict world dynamics, blurring the VLA/WAM boundary.
| Model | Title | Year | Links |
|---|---|---|---|
| RynnVLA-002 | A Unified Vision-Language-Action and World Model | 2025 | π Β· π» |
| WholeBodyVLA | Unified Latent VLA for Whole-Body Loco-Manipulation | 2025 | π Β· π» |
| WorldVLA | Towards an Autoregressive Action World Model | 2025 | π Β· π» |
Self-supervised latent predictive models (non-reconstructive joint-embedding / JEPA). The JEPA foundations (I-JEPA) learn to predict in representation space; the action-conditioned variant (V-JEPA 2-AC) turns that prior into a world model for planning.
| Model | Title | Year | Links |
|---|---|---|---|
| V-JEPA 2-AC | Action-Conditioned Latent World Model for Zero-Shot Planning | 2025 | π |
| I-JEPA | Image-based Joint-Embedding Predictive Architecture (representation foundation) | 2023 | π |
| Model | Title | Year | Links |
|---|---|---|---|
| GAIA-2 | A Controllable Multi-View Generative World Model for Autonomous Driving | 2025 | π |
| Navigation World Models | Conditional Diffusion Transformer for Navigation | 2024 | π |
| GAIA-1 | A Generative World Model for Autonomous Driving | 2023 | π |
Building blocks shared across VLA and WAM policies β how continuous actions become learnable targets.
| Method | Title | Year | Links |
|---|---|---|---|
| FAST | Efficient (DCT-based) Action Tokenization for VLAs | 2025 | π Β· π |
| BeT | Behavior Transformers: Cloning k Modes with One Stone | 2022 | π |
Heads that emit continuous action chunks β by denoising diffusion (Diffusion Policy) or by chunked sequence prediction with a CVAE (ACT). Flow-matching heads (Ο0, SmolVLA, β¦) are listed with their models under Flow-Matching VLA.
| Method | Title | Year | Links |
|---|---|---|---|
| Diffusion Policy | Visuomotor Policy Learning via Action Diffusion | 2023 | π Β· π |
| ACT / ALOHA | Action Chunking with Transformers | 2023 | π Β· π |
Non-VLA policies and planners that remain standard baselines in the experimental tables of the papers above. (Diffusion Policy, ACT, and BeT are described under Action Representations.)
| Method | Title | Year | Links |
|---|---|---|---|
| CrossFormer | Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion & Flight | 2024 | π Β· π» |
| RoboFlamingo | Vision-Language Foundation Models as Effective Robot Imitators | 2023 | π Β· π» |
| VoxPoser | Composable 3D Value Maps for Robotic Manipulation (zero-shot LLM + 3D planner) | 2023 | π Β· π |
| RT-1 | Robotics Transformer for Real-World Control at Scale | 2022 | π |
| Name | Description | Scale | Links |
|---|---|---|---|
| Open X-Embodiment | Cross-embodiment aggregation behind the RT-X models | 1M+ traj Β· 22 embodiments | π Β· π |
| AgiBot World | Large-scale real-world manipulation (Colosseo) | 1M+ traj Β· 217 tasks | π Β· π |
| EgoScale | Scaling dexterous manipulation with diverse egocentric human data | egocentric Β· 2026 | π |
| DexCanvas | Human demos β robot learning for dexterous manipulation | dexterous | π |
| Galaxea Open-World | Mobile-bimanual dataset paired with the G0 dual-system VLA | 500 hrs Β· 150 tasks | π Β· π» |
| DROID | In-the-wild Franka manipulation across 3 continents | 76K traj Β· 564 scenes | π Β· π |
| RoboMIND | Multi-embodiment teleop incl. labeled failures | 107K traj Β· 479 tasks | π Β· π |
| BridgeData V2 | WidowX manipulation w/ language + goal images | 60K traj Β· 24 envs | π Β· π |
| RH20T | Contact-rich skills w/ paired human demos | 110K+ seq Β· 147 tasks | π Β· π |
| Ego-Exo4D | Simultaneous ego + exo video of skilled activity | 1,286 hrs | π Β· π |
| Ego4D | Massive egocentric daily-life video | 3,670 hrs | π Β· π |
| Name | Description | Links |
|---|---|---|
| LIBERO | Lifelong robot-learning, 130 manipulation tasks (de-facto VLA eval) | π Β· π» |
| CALVIN | Long-horizon language-conditioned manipulation | π Β· π» |
| SimplerEnv | Real-to-sim evaluation for manipulation policies | π Β· π |
| RoboCasa | Large-scale kitchen simulation (100 tasks) | π Β· π |
| VLABench | World-knowledge & long-horizon language tasks | π Β· π |
| ManiSkill3 | GPU-parallel manipulation (30K+ FPS) | π Β· π |
| THE COLOSSEUM | Robustness under 14 environmental perturbations | π Β· π |
| RoboArena | Distributed crowd-sourced real-world policy eval | π Β· π |
| RoboChallenge | Large-scale real-robot evaluation of embodied policies | π |
| RobotArena β | Scalable robot benchmarking via real-to-sim translation | π Β· π |
| WorldArena | Perception & functional-utility benchmark for embodied world models | π |
| Meta-World | 50 tabletop tasks for multi-task / meta-RL | π Β· π» |
| RLBench | 100 hand-designed manipulation tasks | π Β· π» |
| Name | Description | Links |
|---|---|---|
| Isaac Sim / Isaac Lab | GPU-native robotics sim + RL/IL framework (Omniverse/USD) | π |
| MuJoCo / MJX | Standard rigid-body engine + JAX/XLA parallel variant | π |
| Genesis | Generative, multi-solver physics platform (up to ~43M FPS) | π |
| ManiSkill | GPU-parallel manipulation simulator on SAPIEN | π |
| SAPIEN | Part-level articulated-object simulator (PartNet-Mobility) | π Β· π |
| Habitat | Photorealistic indoor navigation & rearrangement | π |
| ThreeDWorld | Multimodal Unity3D sim (vision + audio + physics) | π Β· π |
| Newton | Open, differentiable GPU physics engine (NVIDIA + DeepMind + Disney) | π |
| Name | Description | Links |
|---|---|---|
| LeRobot | End-to-end PyTorch robot-learning library + datasets + low-cost HW | π Β· π» |
| openpi | Open models & training/inference for Ο0, Ο0-FAST, Ο0.5 | π» |
| Isaac GR00T | Open humanoid foundation-model framework + checkpoints | π» |
| OpenVLA | Training / LoRA fine-tuning for the 7B OpenVLA model | π» |
| Octo | JAX/Flax generalist transformer policy on OXE | π» |
| robomimic / robosuite | Learning-from-demonstration framework + MuJoCo manipulation sim | π» |
| HIL-SERL | Human-in-the-loop, sample-efficient real-world RL | π» |
A broader, continuously-mined index of recent arXiv work that complements the curated highlights above β 182 additional papers, newest first. Last updated: 2026-09-30. Auto-generated from
data/*.jsonbyscripts/expand_papers.py; papers already highlighted above are omitted here to avoid duplication.
| Paper | Authors | Date | Links |
|---|---|---|---|
| Do World Action Models Generalize Better than VLAs? A Robustness Study | Zhanguang Zhang, Zhiyuan Li et al. | 2026-03-23 | |
| Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models | Riccardo Andrea Izzo, Gianluca Bardaro et al. | 2026-03-05 | |
| Chain of World: World Model Thinking in Latent Motion | Fuxiang Yang, Donglin Di et al. | 2026-03-03 | |
| FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment | Han Zhao, Jingbo Wang et al. | 2026-02-19 | |
| VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation | Changhua Xu, Jie Lu et al. | 2026-02-07 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation | Jade Choghari, Pepijn Kooijmans et al. | 2026-09-29 | |
| Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models | Ruixing Jin, Zicheng Zhu et al. | 2026-03-24 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| Beyond Token Importance: Preserving Spatial Scaffolds for Efficient Vision-Language-Action Inference | Jiayu Chen, Shuyong Gao et al. | 2026-09-29 | |
| AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations | Rui Huang, Yanlin Mu et al. | 2026-09-29 | |
| LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior | Xinkai Wang, Chenyi Wang et al. | 2026-03-26 | |
| 3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models | Bin Yu, Shijie Lian et al. | 2026-03-25 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning | Ruixiao Xu, Wong Lik Hang Kenny et al. | 2026-09-29 | |
| StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks | Ziyi Yin, Sangmin Woo et al. | 2026-09-28 | |
| VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation | Zhide Zhong, Haodong Yan et al. | 2026-03-27 | |
| On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning | Changyu Liu, Yiyang Liu et al. | 2026-01-11 | |
| VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning | Si-Cheng Wang, Tian-Yu Xiang et al. | 2025-09-30 | |
| The arc-shaped radio source at the center of NGC 6334A: Is it a colliding wind region of two young massive stars or the bow shock of a runaway star? | Vanessa Yanza, Sergio A. Dzib et al. | 2025-02-24 | |
| VLA 22 GHz Imaging of Massive Star Formation in Local Wolf-Rayet Galaxies | Nicholas G. Ferraro, Jean L. Turner et al. | 2024-11-09 |
| Paper | Authors | Date | Links |
|---|---|---|---|
| V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving | Junwei You, Weizhe Tang et al. | 2026-09-29 | |
| EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning | Yichao Liang, Amber Li et al. | 2026-09-28 | |
| Enhancing Policy Learning with World-Action Model | Yuci Han, Alper Yilmaz | 2026-03-30 | |
| Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving | Linbo Wang, Yupeng Zheng et al. | 2026-03-25 | |
| NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation | Tianshuai Hu, Zeying Gong et al. | 2026-03-16 | |
| AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation | Ge Yuan, Qiyuan Qiao et al. | 2026-02-23 |
The curated tables above highlight landmark and representative work. For the exhaustive, auto-maintained index and the baseline methods extracted from experimental tables, see:
| Family | Key baselines |
|---|---|
| VLA | RT-1, RT-2, OpenVLA, Octo, Ο0, Ο0.5, X-VLA, UniVLA, SmolVLA |
| Policy | Diffusion Policy, ACT, BeT, RoboFlamingo, CrossFormer |
| World Model | DreamerV3, I-JEPA, V-JEPA 2, Genie, Cosmos, GR-1/GR-2 |
Contributions are very welcome! To add or fix a paper:
To run the discovery pipeline locally:
pip install -r requirements.txt
python scripts/arxiv_scraper.py --max-results 50 --days-back 30 # writes data/papers.json
python scripts/update_readme.py # refreshes the π auto section
python scripts/expand_papers.py # refreshes the ποΈ extended index
python scripts/build_site.py # rebuilds the GitHub Pages site (index.html)
Released under the MIT License.
Inspired by awesome-vla-wam, awesome-physical-ai, and awesome-vla-study. Taxonomy grounded in the surveys listed above.
Python
100.0%