A curated list of benchmarks for evaluating Vision-Language-Action (VLA) models
3
18 commits
updated Mar 20, 2026
If you want to add or update an entry, start with CONTRIBUTING.md. If this list is useful, consider giving the repository a star ⭐.
A curated list of benchmarks for evaluating Vision-Language-Action (VLA) models and closely related embodied agents across robotics, autonomous driving, GUI agents, interactive environments, and multimodal perception. This repository is benchmark-first and includes datasets, simulation platforms, and evaluation resources only when they directly support benchmark usage, comparison, or reproducibility.
This list does not aim to track model-only papers, generic multimodal benchmarks with no action or embodied relevance, or infrastructure projects that do not directly support benchmark use.
Benchmarks for embodied agents operating in physical or simulated robotic environments.
Benchmarks for language-conditioned manipulation with robot arms, mobile manipulators, or bimanual systems.
| Name | Highlights | References |
|---|---|---|
| Open X-Embodiment (OXE) | [ICRA 2024] Cross-embodiment dataset with 1M+ robot episodes from 22 institutions and 22 robot types | arXiv · GitHub · Website |
| DROID | [RSS 2024] 76k demonstrations across diverse real-world environments for scalable VLA training | arXiv · GitHub · Website |
| ManiSkill3 | [NeurIPS 2024] GPU-parallelized simulation with richer tasks and improved evaluation protocols | arXiv · GitHub · Website |
| RoboCasa | [RSS 2024] 100+ kitchen manipulation tasks in photorealistic MuJoCo-based environments | arXiv · GitHub · Website |
| SimplerEnv | [CoRL 2024] Simulation evaluation framework designed to correlate with real-world VLA performance | arXiv · GitHub · Website |
| Mobile ALOHA | [CoRL 2024] Extends ALOHA to a mobile whole-body platform for real-world household tasks | arXiv · GitHub · Website |
| RoboAgent | [ICRA 2024] Generalizable manipulation via semantic augmentations evaluated on 12 real-world tasks | arXiv · GitHub · Website |
| Octo | [RSS 2024] Generalist robot policy trained on OXE data, evaluated across BridgeV2 and Franka | arXiv · GitHub · HF · Website |
| GR-1 | [ICLR 2024] ByteDance's generalist manipulation policy evaluated on CALVIN and real-world tasks | arXiv · GitHub |
| OpenVLA | [CoRL 2024] Open-source 7B VLA model evaluated on BridgeV2 and OXE tasks | arXiv · GitHub · HF · Website |
| LIBERO | [NeurIPS 2023] Knowledge transfer across 130 language-conditioned tasks in 4 structured suites | arXiv · GitHub · Website |
| VIMA-BENCH | [ICML 2023] 17 task types specified via interleaved text-and-image prompts for multimodal manipulation | arXiv · GitHub · Website |
| FurnitureBench | [RSS 2023] Real and simulated furniture assembly requiring long-horizon dexterous manipulation | arXiv · GitHub · Website |
| BridgeData V2 | [CoRL 2023] Large-scale tabletop manipulation with language annotations across diverse kitchen scenes | arXiv · GitHub · Website |
| ManiSkill2 | [ICLR 2023] 20 task families with rigid/soft-body physics for generalizable manipulation skill evaluation | arXiv · GitHub · Website |
| ALOHA / ACT | [RSS 2023] Low-cost bimanual teleoperation with imitation learning demonstrations | arXiv · GitHub · Website |
| RT-1 | [RSS 2023] Google's Robotics Transformer trained on 700+ real-world manipulation tasks | arXiv · Website |
| RT-2 | [CoRL 2023] Combines internet-scale VLM pretraining with robot control for emergent generalization | arXiv · Website |
| CALVIN | [RA-L 2022] Long-horizon manipulation requiring chaining up to 5 language-conditioned tasks in a row | arXiv · GitHub · Website |
| RoboMimic | [CoRL 2021] Standardized demonstration datasets and algorithm baselines for robot learning | arXiv · GitHub · Website |
| RLBench | [RA-L 2020] 100+ unique language-conditioned tasks in CoppeliaSim simulator | arXiv · GitHub |
| MetaWorld | [CoRL 2020] 50 distinct tabletop manipulation tasks for multi-task and meta-RL | arXiv · GitHub · Website |
| Franka Kitchen | [CoRL 2019] Multi-task kitchen manipulation with a Franka Panda robot across 7 compound subtasks | arXiv · GitHub |
| SpatialBench | Evaluates spatial reasoning capabilities of VLMs in manipulation contexts | GitHub |
| π0 (pi-zero) | Physical Intelligence's large robot foundation model for diverse dexterous tasks | arXiv · Website |
| VLABench | [ICCV 2025] Large-scale benchmark for language-conditioned manipulation with long-horizon reasoning across 100 tasks and multiple dimensions | arXiv · GitHub · HF · Website |
| RoboVerse | [RSS 2025] Unified platform, dataset, and benchmark for scalable and generalizable robot learning across 5000+ tasks | arXiv · GitHub · Website |
| MIKASA-Robo | [ICLR 2026] Benchmark for robotic tabletop manipulation with 32 memory-intensive tasks covering object, spatial, sequential, and capacity memory | arXiv · GitHub · Website |
| RoboVLMs | Comprehensive evaluation of VLMs as robot policy foundation models across manipulation tasks | GitHub · Website |
| RoboTwin | [CVPR 2025] Dual-arm robot benchmark with generative digital twins for automated data collection across 50 diverse bimanual manipulation tasks | arXiv · GitHub · Website |
| BiGym | Demo-driven mobile bi-manual manipulation benchmark with 40 novel household tasks spanning easy to extremely difficult difficulty levels | arXiv · GitHub · Website |
Benchmarks for dexterous hands, whole-body control, and locomotion-oriented embodied control.
| Name | Highlights | References |
|---|---|---|
| HumanoidBench | [NeurIPS 2024] 27 diverse whole-body tasks for humanoid robot control evaluation | arXiv · GitHub · Website |
| DexArt | [CVPR 2023] 4 complex dexterous tasks for articulated object manipulation with a multi-fingered hand | arXiv · GitHub · Website |
| DexDeform | [ICLR 2023] 5 dexterous deformable object manipulation tasks for learning contact-rich policies | arXiv · GitHub |
| LEAP Hand | [CoRL 2023] Low-cost dexterous hand platform with benchmark tasks for in-hand manipulation | arXiv · Website |
| MyoSuite | [NeurIPS 2022] Musculoskeletal simulation for physiologically accurate hand and arm control | arXiv · GitHub · Website |
| IsaacGym Benchmark | [NeurIPS 2021] GPU-accelerated dexterous benchmarks including in-hand cube reorientation and pen spinning | arXiv · Website |
| D4RL | [NeurIPS 2020] Offline RL benchmark covering locomotion, manipulation, and navigation with standardized datasets | arXiv · GitHub |
| DeepMind Control Suite | [JMLR 2020] Continuous control tasks for locomotion and manipulation used as standard RL/VLA baselines | arXiv · GitHub |
| Adroit | [RSS 2018] 24-DoF anthropomorphic hand benchmark for pen, door, hammer, and relocate tasks | arXiv · GitHub |
Benchmarks for navigation in 3D environments from language, visual, or object-goal instructions.
| Name | Highlights | References |
|---|---|---|
| GOAT-Bench | [CVPR 2024] Multi-modal lifelong navigation benchmark with sequential object goals specified by category, language, or image in open-vocabulary settings | GitHub · Website |
| OpenEQA | [CVPR 2024] Embodied Question Answering in the era of foundation models with 1600+ questions spanning spatial, object, and functional understanding | arXiv · GitHub · Website |
| ScanQA | [CVPR 2022] Spatial question answering grounded in 3D indoor point cloud environments | arXiv · GitHub |
| SOON | [CVPR 2021] Scenario-oriented object navigation requiring reasoning about object-scene relations | arXiv · Website |
| ImageNav | [ICCV 2021] Navigation to a goal specified by an image rather than a category label | GitHub |
| VLN-BERT / HAMT | [CVPR 2021] BERT-based and history-aware transformer approaches evaluated on R2R and REVERIE | arXiv · GitHub |
| MultiON | [CVPR 2021] Multi-object navigation requiring sequential finding of multiple target objects | arXiv · Website |
| RxR (Room-Across-Rooms) | [EMNLP 2020] Multilingual navigation in 3 languages with denser annotations than R2R | arXiv · GitHub |
| REVERIE | [CVPR 2020] Combines long-horizon navigation with object grounding via referring expressions | arXiv · GitHub · Website |
| ObjectNav (Habitat) | [NeurIPS 2020] Navigate to a specified object category in novel photorealistic environments | arXiv · Website |
| VLN-CE | [EMNLP 2020] Continuous-environment R2R in Habitat without discrete viewpoint teleportation | arXiv · GitHub |
| CVDN | [CVPR 2019] Cooperative multi-turn dialogue navigation where a helper guides an agent to a target | arXiv · GitHub |
| TOUCHDOWN | [CVPR 2019] Street-view navigation in real-world NYC following natural language descriptions | arXiv · GitHub |
| R2R (Room-to-Room) | [CVPR 2018] Foundational VLN benchmark for instruction-following navigation in Matterport3D | arXiv · GitHub · Website |
| EQA (Embodied Question Answering) | [CVPR 2018] Agents navigate and answer visual questions about indoor environments | arXiv · Website |
| MP3D (Matterport3D) | [CVPR 2017] Large-scale RGB-D indoor dataset widely used as a navigation substrate | Website |
Benchmarks for household and open-world agents that combine navigation, manipulation, and task reasoning.
| Name | Highlights | References |
|---|---|---|
| Habitat 3.0 | [ICLR 2024] Social navigation, collaboration, and rearrangement tasks in embodied AI simulation | arXiv · GitHub · Website |
| EmbodiedBench | [ICML 2025] Comprehensive benchmark evaluating MLLMs as embodied agents across 1128 tasks spanning high-level (household, habitat) and low-level (manipulation, navigation) with 6 capability dimensions | arXiv · GitHub · Website |
| ARNOLD | [ICCV 2023] Language-conditioned articulated object state change tasks in 3D scenes | arXiv · GitHub · Website |
| TEACh | [AAAI 2022] Task completion in AI2-THOR via multi-turn dialogue history | arXiv · GitHub |
| ProcTHOR | [NeurIPS 2022] Procedurally generated AI2-THOR houses for training and evaluating generalist embodied agents | arXiv · GitHub · Website |
| ALFWorld | [ICLR 2021] Aligns TextWorld with AI2-THOR for language-conditioned household task completion | arXiv · GitHub · Website |
| ThreeDWorld (TDW) | [NeurIPS 2021] Multi-modal physical simulation for transport, manipulation, and multi-agent tasks | arXiv · GitHub · Website |
| Watch-And-Help | [ICLR 2021] Social intelligence benchmark for agent cooperation in VirtualHome environments | arXiv · GitHub |
| AI2-THOR | [CVPR 2018] Interactive 3D household simulation with physics for pick-and-place, slicing, cooking, and more | arXiv · GitHub · Website |
| VirtualHome | [CVPR 2018] Multi-step household simulation supporting natural language program execution | arXiv · GitHub · Website |
| BEHAVIOR-1K | 1000 household activities grounded in real human needs, simulated in OmniGibson | GitHub · Website |
| LANMP | Language-conditioned navigation and manipulation on a real mobile manipulator | arXiv · GitHub |
Benchmarks for socially aware, collaborative, or communicative robot behavior.
| Name | Highlights | References |
|---|---|---|
| Habitat 3.0 Social Nav | [ICLR 2024] Social navigation requiring robots to move around humans and follow social norms | arXiv · Website |
| ConcepFusion / LangNav | [RSS 2023] Open-vocabulary 3D feature fields enabling natural language queries for navigation | arXiv · GitHub |
| TDW-Social | [NeurIPS 2021] Multi-agent social simulation requiring cooperation and communication in TDW | arXiv · GitHub · Website |
| SEAN | [RA-L 2020] Navigation benchmark requiring socially compliant behavior around pedestrians | arXiv · Website |
| RoboTHOR Challenge | ObjectNav challenge bridging simulation and real robot deployment in home-like environments | arXiv · Website |
| ProCo | Proactive cooperation benchmark requiring initiative and dialogue from embodied agents | GitHub |
Benchmarks for driving agents that ground perception, language, and decision-making in traffic scenes.
| Name | Highlights | References |
|---|---|---|
| DriveLM | [ECCV 2024] Graph-structured vision-language QA benchmark for end-to-end driving on nuScenes | arXiv · GitHub · HF |
| LingoQA | [CVPR 2024] Video QA benchmark for driving scene understanding grounded in real-world footage | arXiv · GitHub |
| Rank2Tell | [WACV 2024] Dataset for ranking and describing important objects in driving scenes | arXiv · GitHub |
| MAPLM | [WACV 2024] Map-based language model benchmark for understanding HD maps in autonomous driving | arXiv · GitHub |
| NuScenes-QA | [AAAI 2024] 460k QA pairs for visual question answering built on nuScenes | arXiv · GitHub |
| TOD3Cap | [ECCV 2024] 3D dense captioning for autonomous driving with 2.3M natural language descriptions | arXiv · GitHub |
| DriveVLM | [CoRL 2024] Chain-of-thought VLM reasoning integrated with motion planning for autonomous driving | arXiv · Website |
| MMAD | [NeurIPS 2024] Massive multimodal AD benchmark with 18 perception and reasoning subtasks | arXiv · GitHub |
| nuScenes | [CVPR 2020] Large-scale AD dataset with 3D bounding boxes, HD maps, and rich multi-sensor data | GitHub · Website |
| Waymo Open Dataset | [CVPR 2020] High-quality driving dataset with LiDAR and camera data for perception and prediction | Website |
| BDD-X | [ECCV 2018] Explainable driving dataset with textual descriptions and justifications for actions | arXiv · GitHub |
| CARLA Challenge | [CoRL 2017] Open urban driving simulation benchmark for route completion and safety compliance | GitHub · Website |
| nuPlan | Closed-loop motion planning benchmark with expert demonstrations and reactive simulation | arXiv · GitHub · Website |
| DriveBench | Comprehensive benchmark evaluating VLMs across multiple driving perception and QA tasks | Website |
| VLAAD | Vision-Language-Action benchmark for autonomous driving with natural language instructions | GitHub |
| LaMPilot | [CVPR 2024] Open benchmark dataset for autonomous driving evaluated via language model programs | arXiv · GitHub |
| STSBench | Spatio-temporal scenario benchmark for evaluating MLLMs in interactive autonomous driving scenarios | GitHub |
| Drive4C | Closed-loop benchmark assessing what foundation models need to enable language-guided autonomous driving | GitHub |
| NAVSIM | [NeurIPS 2024] Non-reactive simulation benchmark for data-driven autonomous driving evaluation with 1000+ hours of real-world driving data | arXiv · GitHub |
| Bench2Drive | Closed-loop multi-ability benchmark for end-to-end autonomous driving across 220 scenarios in CARLA covering perception, prediction, and planning | arXiv · GitHub |
Benchmarks for agents that perceive graphical interfaces and act through keyboard and mouse inputs.
| Name | Highlights | References |
|---|---|---|
| AndroidWorld | [ICLR 2025] 116 programmatic tasks across 20 real Android apps | arXiv · GitHub · Website |
| GUI-Odyssey | [ICCV 2025] Cross-app Android navigation requiring multi-step actions across multiple apps | arXiv · GitHub |
| WebArena | [ICLR 2024] 812 realistic tasks across e-commerce, social, coding, and content websites | arXiv · GitHub · Website |
| VisualWebArena | [ACL 2024] Extends WebArena with visually grounded tasks requiring image-based reasoning | arXiv · GitHub · Website |
| OSWorld | [NeurIPS 2024] 369 open-ended computer tasks in real OS environments (Ubuntu, Windows, macOS) | arXiv · GitHub · HF · Website |
| ScreenSpot | [ACL 2024] GUI grounding benchmark for localizing UI elements from text instructions | arXiv · GitHub |
| AssistGUI | [CVPR 2024] Desktop productivity benchmark (Word, Excel, PowerPoint) with video-based evaluation | arXiv · GitHub · Website |
| AgentBench | [ICLR 2024] LLM-as-agent benchmark with 8 environments: OS, database, web, games, and more | arXiv · GitHub · Website |
| WorkArena | [ICML 2024] 33 enterprise software tasks on a real ServiceNow instance | arXiv · GitHub · Website |
| SWE-bench | [ICLR 2024] Requires agents to resolve real GitHub issues as a software engineering challenge | arXiv · GitHub · Website |
| τ-bench (tau-bench) | [NeurIPS 2024] Tool-agent-user interaction benchmark evaluating agentic task completion with tool use | arXiv · GitHub |
| Mind2Web | [NeurIPS 2023] 2000+ tasks from 137 real-world websites for generalist web agent evaluation | GitHub · Website |
| AITW (Android in the Wild) | [NeurIPS 2023] Large-scale human demonstration dataset on Android devices across diverse tasks | arXiv · GitHub |
| WebSRC | [EMNLP 2021] Structural reading comprehension for understanding web page layouts and content | arXiv · GitHub · Website |
| Screen2Words | [UIST 2021] 112k human-annotated mobile UI screen summaries for screen captioning | arXiv · GitHub |
| MiniWob++ | [ICLR 2018] 100+ browser mini-tasks (clicking, typing, form filling) for web interaction | GitHub · Website |
| WindowsAgentArena | 154 Windows OS tasks spanning productivity, web, and system applications | arXiv · GitHub · Website |
| ScreenSpot-Pro | GUI grounding benchmark for professional high-resolution computer use across complex desktop applications | arXiv · GitHub |
| Spider2-V | [NeurIPS 2024] 494 tasks testing multimodal agents at automating real-world data science and engineering workflows | arXiv · GitHub · Website |
| GUI-World | Comprehensive GUI benchmark for evaluating multimodal agents across mobile, web, and desktop interfaces with sequential tasks and QA scenarios | arXiv · GitHub · Website |
| AndroidLab | [ICLR 2025] Systematic environment and benchmark for training and evaluating Android autonomous agents across 138 tasks in 9 apps | arXiv · GitHub · Website |
| OmniACT | Dataset and benchmark for multimodal autonomous agents performing computer tasks via UI element grounding across desktop and web interfaces | arXiv · GitHub · Website |
| AgentTrek | Automated GUI agent trajectory synthesis benchmark derived from web tutorials with 15K+ step-level annotations for agent training and evaluation | arXiv · GitHub · Website |
Benchmarks for language-guided sequential decision-making in games and interactive simulated worlds.
| Name | Highlights | References |
|---|---|---|
| Voyager / DEPS | [NeurIPS 2023] LLM-powered lifelong learning Minecraft agent evaluation framework | arXiv · GitHub · Website |
| MineRL / MineDojo | [NeurIPS 2022] 3000+ open-ended Minecraft tasks leveraging internet-scale video knowledge | arXiv · GitHub · Website |
| ScienceWorld | [EMNLP 2022] 30 science experiment tasks requiring procedural multi-step reasoning and action | arXiv · GitHub · Website |
| CrafterBench | [TMLR 2022] Open-world survival game measuring 22 achievements requiring long-horizon planning | arXiv · GitHub |
| BEHAVIOR (iGibson) | [NeurIPS 2021] 100 household activities in iGibson simulation requiring physical commonsense reasoning | GitHub · Website |
| NetHack Learning Environment (NLE) | [NeurIPS 2020] Complex roguelike game for long-horizon decision making and language-conditioned play | arXiv · GitHub |
| Atari-HEAD | [AAAI 2019] Human eye-tracking data for Atari games enabling human-like visual attention evaluation | arXiv · GitHub |
| BabyAI | [ICLR 2019] Gridworld benchmark with 19 procedurally generated instruction difficulty levels | arXiv · GitHub |
| GROOT | Open-ended Minecraft agent benchmark with skill learning and generalization | GitHub |
| MiniGrid / MiniWorld | Lightweight 2D/3D grid environments for language-conditioned navigation and reasoning | GitHub |
| TextWorld | Framework for generating and playing text-based adventure games to train language agents | arXiv · GitHub |
Benchmarks for multimodal perception, temporal understanding, and grounding capabilities that support VLA systems.
Benchmarks for general multimodal reasoning, instruction following, and cross-modal grounding.
| Name | Highlights | References |
|---|---|---|
| MMMU | [CVPR 2024] 11.5k expert-level questions across 57 subjects for massive multidisciplinary evaluation | arXiv · GitHub · HF · Website |
| MMBench | [ECCV 2024] Systematic 20-dimension VLM evaluation with 3000 questions | arXiv · GitHub · Website |
| SEED-Bench | [CVPR 2024] 19K questions across 12 dimensions for multi-granularity generative MLLM evaluation | arXiv · GitHub |
| HallusionBench | [CVPR 2024] Hallucination detection benchmark for visual understanding in VLMs | arXiv · GitHub |
| MMStar | [NeurIPS 2024] 1500 carefully curated samples minimizing data leakage and requiring genuine VL reasoning | arXiv · GitHub · Website |
| CV-Bench | [EMNLP 2024] Compositional reasoning and spatial understanding benchmark for VLMs | GitHub |
| BLINK | [ECCV 2024] Visual perception benchmark targeting human intuitive abilities that challenge VLMs | arXiv · GitHub · Website |
| WildVision | [NeurIPS 2024] Real-world VLM evaluation via human preference collected from live user interactions | HF |
| VLMEvalKit | [ACM MM 2024] Open-source toolkit supporting 100+ VLMs across 50+ benchmarks | arXiv · GitHub |
| POPE | [EMNLP 2023] Polling-based object probing for evaluating object hallucination in VLMs | arXiv · GitHub |
| MME | 14-subtask benchmark covering MLLM perception and cognition capabilities | arXiv · GitHub |
| TouchStone | GPT-4 scored VLM evaluation across 5 ability dimensions | arXiv · GitHub |
| RealWorldQA | Real-world spatial understanding benchmark from xAI testing physical world reasoning | Website |
| MMTE | Multimodal task execution benchmark for following complex visual instructions | GitHub |
| MMIU | [ICLR 2025] Multimodal Multi-image Understanding benchmark with 7 image-relationship types, 52 tasks, and 11K questions for evaluating LVLMs | arXiv · GitHub · Website |
| MathVista | [ICLR 2024] Mathematical reasoning in visual contexts benchmark requiring compositional reasoning across 28 math tasks and 5 visual types | arXiv · GitHub · Website |
| MMMU-Pro | [NeurIPS 2024] More robust MMMU extension with college-level expert questions requiring genuine multimodal understanding, resistant to language-only shortcuts | arXiv · GitHub · Website |
| CharXiv | [NeurIPS 2024] Chart understanding benchmark exposing significant gaps in VLMs on 2323 scientific paper figures with descriptive and reasoning tasks | arXiv · GitHub · Website |
| OlympiadBench | [ACL 2024] Olympiad-level bilingual multimodal science benchmark with 8952 problems in mathematics and physics for frontier VLM evaluation | arXiv · GitHub |
| MEGA-Bench | [NeurIPS 2024] Scaling multimodal evaluation to 500+ diverse real-world tasks covering perception, reasoning, and generation across many domains | arXiv · GitHub · Website |
| MM-Vet v2 | Harder and more comprehensive integrated capabilities benchmark for large multimodal models using GPT-4 scoring across 16 skills | arXiv · GitHub |
Benchmarks for temporal video understanding, action reasoning, and long-context event interpretation.
| Name | Highlights | References |
|---|---|---|
| TemporalBench | [ICLR 2025] Fine-grained temporal video understanding benchmark for VLMs with 2M QA pairs | arXiv · GitHub |
| Video-MME | [CVPR 2025] First comprehensive evaluation benchmark of MLLMs in video analysis covering short/medium/long videos across 30 domains | arXiv · GitHub · Website |
| LongVideoBench | [NeurIPS 2024] Long-context interleaved video-language understanding benchmark with 6678 questions on videos ranging from minutes to 1 hour | arXiv · GitHub · Website |
| MLVU | [NeurIPS 2024] Multi-task long video understanding benchmark with 9 task types spanning holistic, single-detail, and multi-detail evaluation across 10–180 minute videos | arXiv · GitHub |
| MVBench | [CVPR 2024] 20 challenging temporal reasoning tasks for multi-task video understanding | arXiv · GitHub |
| EgoTaskQA | [NeurIPS 2022] Egocentric video QA with task goal inference, procedural reasoning, and state tracking | arXiv · GitHub |
| STAR (Situated Reasoning) | [NeurIPS 2021] Situated temporal action reasoning with 4 question types grounded in real videos | arXiv · Website |
| NExT-QA | [CVPR 2021] Video QA benchmark emphasizing causal and temporal reasoning about actions | arXiv · GitHub · Website |
| COIN | [CVPR 2019] 11,827 instructional videos across 180 tasks and 778 procedural steps | arXiv · GitHub · Website |
| CrossTask | [CVPR 2019] Procedural activity understanding across 83 instructional video tasks | arXiv · GitHub |
| HowTo100M | [ICCV 2019] Narrated instructional videos for learning action-language grounding at scale | arXiv · Website |
| Something-Something v2 | [ICCV 2017] Temporal reasoning requiring understanding of object interactions and motion direction | arXiv · Website |
| Charades | [ECCV 2016] Indoor activity dataset requiring temporal localization and compositional reasoning | arXiv · Website |
| ActivityNet | [CVPR 2015] Activity recognition, localization, and dense video captioning at large scale | arXiv · Website |
| Kinetics-700 | Large-scale action recognition dataset with 700 human action classes | arXiv · Website |
| Video-Bench | Video-exclusive benchmark for evaluating video understanding capabilities in MLLMs | arXiv · GitHub |
| VideoVista | Diverse video understanding benchmark spanning 14 task types and 35 categories | arXiv · GitHub |
| DREAM-1K | Procedural video understanding benchmark for evaluating robot task planning | arXiv · GitHub |
Benchmarks for egocentric perception and activity understanding from first-person video.
| Name | Highlights | References |
|---|---|---|
| Ego-Exo4D | [CVPR 2024] Paired egocentric and exocentric video dataset for skill assessment and correspondence | arXiv · Website |
| Ego4D | [CVPR 2022] 3670h egocentric video with benchmarks for episodic memory, forecasting, and hand-object interaction | arXiv · GitHub · Website |
| EPIC-Kitchens-100 | [IJCV 2022] 100 hours of unscripted egocentric cooking actions with fine-grained annotations | arXiv · GitHub · Website |
| Assembly101 | [CVPR 2022] Egocentric procedural dataset for assembling and disassembling 101 toy vehicles | arXiv · GitHub · Website |
| HOI4D | [CVPR 2022] Egocentric 4D dataset for category-level human-object interaction manipulation | arXiv · GitHub · Website |
| EgoProceL | [ECCV 2022] Egocentric procedural learning benchmark for keystep recognition from instructional videos | arXiv · GitHub |
| LEMMA | [ECCV 2020] Multi-person, multi-task dataset for compositional action understanding | arXiv · Website |
| EGTEA Gaze+ | [ECCV 2018] First-person cooking dataset with gaze and hand masks for fine-grained recognition | arXiv · Website |
| EPIC-Kitchens Challenges | Annual challenges on action recognition, detection, anticipation, and retrieval | Website |
| EgoPlan-Bench | [IJCV 2024] Benchmarking MLLMs for human-level task planning from egocentric video with real-world household scenarios | arXiv · GitHub |
Benchmarks for grounding language and reasoning in 3D scenes for embodied perception.
| Name | Highlights | References |
|---|---|---|
| EmbodiedScan | [CVPR 2024] Holistic 3D perception benchmark with RGB-D input from egocentric views | arXiv · GitHub · Website |
| LEO | [ICML 2024] Embodied generalist benchmark requiring 3D scene understanding for grounded dialogue and planning | arXiv · GitHub · Website |
| SQA3D (Situated QA) | [ICLR 2023] Situated reasoning benchmark where agents answer questions from a 3D first-person perspective | arXiv · GitHub · Website |
| 3D-VisTA | [ICCV 2023] Pre-trained transformer for 3D VL tasks: grounding, QA, and dense captioning | arXiv · GitHub · Website |
| MultiScan | [NeurIPS 2022] Multi-scan indoor reconstruction with articulated object annotations for VLA research | arXiv · GitHub · Website |
| ScanEnts3D | [ECCV 2022] Links 3D scene entities to text mentions for fine-grained object grounding | arXiv · GitHub |
| ScanRefer | [ECCV 2020] Localizes objects in 3D point clouds using free-form language descriptions | arXiv · GitHub · Website |
| CLEVR3D | 3D extension of CLEVR for spatial reasoning questions in synthetic environments | GitHub |
| Nu-Scenes QA | QA benchmark grounded in outdoor 3D LiDAR scenes for driving reasoning | arXiv · GitHub |
| MMScan | [NeurIPS 2024] Multi-modal 3D scene dataset with 1.4M hierarchical grounded language annotations on 109k objects and benchmarks for visual grounding and QA | arXiv · GitHub · Website |
| SceneVerse | [ECCV 2024] Large-scale 3D vision-language dataset with 68K indoor scenes and 2.5M language-scene pairs for grounded scene understanding and spatial reasoning | arXiv · GitHub · Website |
Benchmarks for visual grounding, referring expressions, and spatial language understanding.
| Name | Highlights | References |
|---|---|---|
| SeedBench-2-Plus | [NeurIPS 2024] Extended SEED-Bench focusing on charts, maps, web pages, and document comprehension | arXiv · GitHub |
| ARO (Attribution, Relation, Order) | [ICLR 2023] Reveals VLMs' limited sensitivity to word order and relational structure | arXiv · GitHub |
| VSR (Visual Spatial Reasoning) | [NAACL 2022] True/false benchmark testing spatial relationships between objects in images | arXiv · GitHub |
| Winoground | [CVPR 2022] Tests visio-linguistic compositional reasoning via image-caption matching with swapped words | arXiv · HF |
| FLUTE | [EMNLP 2022] Figurative language understanding for VLMs involving metaphors and similes | arXiv · GitHub |
| GQA | [CVPR 2019] Compositional QA benchmark requiring multi-step spatial and relational reasoning | arXiv · Website |
| Visual Genome | [IJCV 2017] Densely annotated images with region descriptions, QA, and scene graphs for grounded understanding | arXiv · GitHub · Website |
| CLEVR | [CVPR 2017] Compositional language and elementary visual reasoning diagnostic benchmark | arXiv · GitHub · Website |
| RefCOCO / RefCOCO+ / RefCOCOg | [ECCV 2016] Referring expression comprehension benchmarks for localizing objects from natural language | arXiv · GitHub |
| SpatialBot | Spatial understanding benchmark evaluating depth-aware reasoning in VLMs | GitHub |
Supporting resources used to run, compare, and reproduce VLA benchmark results.
Simulation platforms and task environments used to build or execute VLA benchmarks.
| Name | Highlights | References |
|---|---|---|
| IsaacGym | [NeurIPS 2021] Legacy GPU physics simulation for massively parallel RL training on dexterous tasks | arXiv · Website |
| Robosuite | [IROS 2021] Modular robot learning framework built on MuJoCo with standardized task suites | arXiv · GitHub · Website |
| SAPIEN | [CVPR 2020] Simulation platform for articulated object manipulation with PhysX physics | arXiv · GitHub · Website |
| Habitat-Sim | [ICCV 2019] High-performance 3D simulator for embodied AI with photorealistic rendering and physics | arXiv · GitHub · Website |
| CARLA | [CoRL 2017] Open urban driving simulator with rich sensor modalities for autonomous driving research | arXiv · GitHub · Website |
| MuJoCo | [IROS 2012] Fast and accurate physics simulation engine widely used for locomotion and manipulation | arXiv · GitHub · Website |
| IsaacSim / IsaacLab | NVIDIA's GPU-accelerated simulation platform for robot learning and VLA development | GitHub · Website |
| Genesis | Generative simulation engine with ultra-fast physics and photorealistic rendering | arXiv · GitHub · Website |
| PyBullet / Panda-Gym | Open-source physics simulation with Panda robot gym environments for tabletop manipulation | GitHub · Website |
| OmniGibson | NVIDIA Omniverse-based simulation for BEHAVIOR-1K with realistic physics and rendering | arXiv · GitHub · Website |
| AI2-THOR (Simulator) | Photorealistic interactive indoor simulation for embodied household task research | arXiv · GitHub · Website |
| CoppeliaSim (V-REP) | Versatile robot simulation platform underlying RLBench and other benchmarks | Website |
| SUMO | Microscopic traffic simulation supporting multi-modal transportation research | GitHub · Website |
| XR-EgoBench | Extended reality egocentric benchmark platform for AR/VR embodied agent evaluation | GitHub |
Datasets and model hubs that provide training assets, benchmark data, or evaluation-facing artifacts for VLA work.
| Name | Highlights | References |
|---|---|---|
| HuggingFace LeRobot Datasets | Standardized real-robot demonstration datasets for VLA training and benchmarking | GitHub · HF |
| HuggingFace Open-VLA | HuggingFace hub for OpenVLA model weights, training data, and evaluation scripts | GitHub · HF |
| HuggingFace Robot Benchmarks | Curated collection of robot learning benchmarks and evaluation protocols | HF |
| Open X-Embodiment Dataset | HuggingFace mirror of the OXE cross-embodiment dataset aggregating 1M+ robot episodes | arXiv · HF |
| RoboSet | Large-scale robot manipulation dataset with 100k+ demonstrations across 12 task categories | arXiv · Website |
| DROID Dataset | 76k robot manipulation demonstrations on HuggingFace for diverse real-world VLA training | arXiv · HF |
| BridgeData V2 (HF) | HuggingFace version of BridgeData V2 for easy access and VLA pretraining | arXiv · HF |
| EgoMimic Dataset | Paired egocentric human video and robot demonstration dataset for VLA transfer learning | arXiv · Website |
| RH20T | Large-scale robotic dataset with 110k demonstrations across 20 tasks | arXiv · GitHub · Website |
| Ego4D HuggingFace | Ego4D benchmark data accessible via HuggingFace for convenient evaluation | arXiv · HF |
Leaderboards and evaluation services that host benchmark results, competitions, or public comparisons.
| Name | Highlights | References |
|---|---|---|
| EvalAI | Cloud-based challenge platform hosting embodied AI, VQA, navigation, and VLA competitions | Website |
| PaperWithCode Embodied AI | Live leaderboards tracking SOTA across navigation, manipulation, and QA benchmarks | Website |
| HuggingFace Open VLA Leaderboard | Community leaderboard for real-world robot VLA performance | HF |
| Open-Compass (OpenVLM Leaderboard) | Comprehensive VLM leaderboard comparing models across 50+ benchmarks | Website |
| LIBERO Leaderboard | Official leaderboard for the LIBERO manipulation benchmark suite | Website |
| CARLA Autonomous Driving Leaderboard | Official competition leaderboard for autonomous driving agents in CARLA | Website |
| WebArena Leaderboard | Community tracking sheet for web agent performance on WebArena | Website |
| OSWorld Leaderboard | Live leaderboard for OS computer-use agents on the OSWorld benchmark | Website |
| Ego4D Challenges (EvalAI) | Annual Ego4D challenge tracks on episodic memory, forecasting, AV, hands, and narrations | Website |
| BEHAVIOR-1K Leaderboard | Evaluation results for household activity agents in the BEHAVIOR benchmark | Website |
| AgentBench Leaderboard | Rankings for LLM-as-agent performance across OS, web, game, and database tasks | Website |
| EmbodiedBench Leaderboard | CVPR 2026 challenge leaderboard for MLLM-based embodied agents across high-level and low-level tasks | Website |
See CONTRIBUTING.md for repository scope, placement rules, row format, sorting expectations, and the pull request checklist.
This list is released under the CC0 1.0 Universal public domain dedication. You are free to copy, modify, and distribute this work, even for commercial purposes, without asking permission.
13 commits
5 commits
A curated list of benchmarks for evaluating Vision-Language-Action (VLA) models
3
18 commits
updated Mar 20, 2026
If you want to add or update an entry, start with CONTRIBUTING.md. If this list is useful, consider giving the repository a star ⭐.
A curated list of benchmarks for evaluating Vision-Language-Action (VLA) models and closely related embodied agents across robotics, autonomous driving, GUI agents, interactive environments, and multimodal perception. This repository is benchmark-first and includes datasets, simulation platforms, and evaluation resources only when they directly support benchmark usage, comparison, or reproducibility.
This list does not aim to track model-only papers, generic multimodal benchmarks with no action or embodied relevance, or infrastructure projects that do not directly support benchmark use.
Benchmarks for embodied agents operating in physical or simulated robotic environments.
Benchmarks for language-conditioned manipulation with robot arms, mobile manipulators, or bimanual systems.
| Name | Highlights | References |
|---|---|---|
| Open X-Embodiment (OXE) | [ICRA 2024] Cross-embodiment dataset with 1M+ robot episodes from 22 institutions and 22 robot types | arXiv · GitHub · Website |
| DROID | [RSS 2024] 76k demonstrations across diverse real-world environments for scalable VLA training | arXiv · GitHub · Website |
| ManiSkill3 | [NeurIPS 2024] GPU-parallelized simulation with richer tasks and improved evaluation protocols | arXiv · GitHub · Website |
| RoboCasa | [RSS 2024] 100+ kitchen manipulation tasks in photorealistic MuJoCo-based environments | arXiv · GitHub · Website |
| SimplerEnv | [CoRL 2024] Simulation evaluation framework designed to correlate with real-world VLA performance | arXiv · GitHub · Website |
| Mobile ALOHA | [CoRL 2024] Extends ALOHA to a mobile whole-body platform for real-world household tasks | arXiv · GitHub · Website |
| RoboAgent | [ICRA 2024] Generalizable manipulation via semantic augmentations evaluated on 12 real-world tasks | arXiv · GitHub · Website |
| Octo | [RSS 2024] Generalist robot policy trained on OXE data, evaluated across BridgeV2 and Franka | arXiv · GitHub · HF · Website |
| GR-1 | [ICLR 2024] ByteDance's generalist manipulation policy evaluated on CALVIN and real-world tasks | arXiv · GitHub |
| OpenVLA | [CoRL 2024] Open-source 7B VLA model evaluated on BridgeV2 and OXE tasks | arXiv · GitHub · HF · Website |
| LIBERO | [NeurIPS 2023] Knowledge transfer across 130 language-conditioned tasks in 4 structured suites | arXiv · GitHub · Website |
| VIMA-BENCH | [ICML 2023] 17 task types specified via interleaved text-and-image prompts for multimodal manipulation | arXiv · GitHub · Website |
| FurnitureBench | [RSS 2023] Real and simulated furniture assembly requiring long-horizon dexterous manipulation | arXiv · GitHub · Website |
| BridgeData V2 | [CoRL 2023] Large-scale tabletop manipulation with language annotations across diverse kitchen scenes | arXiv · GitHub · Website |
| ManiSkill2 | [ICLR 2023] 20 task families with rigid/soft-body physics for generalizable manipulation skill evaluation | arXiv · GitHub · Website |
| ALOHA / ACT | [RSS 2023] Low-cost bimanual teleoperation with imitation learning demonstrations | arXiv · GitHub · Website |
| RT-1 | [RSS 2023] Google's Robotics Transformer trained on 700+ real-world manipulation tasks | arXiv · Website |
| RT-2 | [CoRL 2023] Combines internet-scale VLM pretraining with robot control for emergent generalization | arXiv · Website |
| CALVIN | [RA-L 2022] Long-horizon manipulation requiring chaining up to 5 language-conditioned tasks in a row | arXiv · GitHub · Website |
| RoboMimic | [CoRL 2021] Standardized demonstration datasets and algorithm baselines for robot learning | arXiv · GitHub · Website |
| RLBench | [RA-L 2020] 100+ unique language-conditioned tasks in CoppeliaSim simulator | arXiv · GitHub |
| MetaWorld | [CoRL 2020] 50 distinct tabletop manipulation tasks for multi-task and meta-RL | arXiv · GitHub · Website |
| Franka Kitchen | [CoRL 2019] Multi-task kitchen manipulation with a Franka Panda robot across 7 compound subtasks | arXiv · GitHub |
| SpatialBench | Evaluates spatial reasoning capabilities of VLMs in manipulation contexts | GitHub |
| π0 (pi-zero) | Physical Intelligence's large robot foundation model for diverse dexterous tasks | arXiv · Website |
| VLABench | [ICCV 2025] Large-scale benchmark for language-conditioned manipulation with long-horizon reasoning across 100 tasks and multiple dimensions | arXiv · GitHub · HF · Website |
| RoboVerse | [RSS 2025] Unified platform, dataset, and benchmark for scalable and generalizable robot learning across 5000+ tasks | arXiv · GitHub · Website |
| MIKASA-Robo | [ICLR 2026] Benchmark for robotic tabletop manipulation with 32 memory-intensive tasks covering object, spatial, sequential, and capacity memory | arXiv · GitHub · Website |
| RoboVLMs | Comprehensive evaluation of VLMs as robot policy foundation models across manipulation tasks | GitHub · Website |
| RoboTwin | [CVPR 2025] Dual-arm robot benchmark with generative digital twins for automated data collection across 50 diverse bimanual manipulation tasks | arXiv · GitHub · Website |
| BiGym | Demo-driven mobile bi-manual manipulation benchmark with 40 novel household tasks spanning easy to extremely difficult difficulty levels | arXiv · GitHub · Website |
Benchmarks for dexterous hands, whole-body control, and locomotion-oriented embodied control.
| Name | Highlights | References |
|---|---|---|
| HumanoidBench | [NeurIPS 2024] 27 diverse whole-body tasks for humanoid robot control evaluation | arXiv · GitHub · Website |
| DexArt | [CVPR 2023] 4 complex dexterous tasks for articulated object manipulation with a multi-fingered hand | arXiv · GitHub · Website |
| DexDeform | [ICLR 2023] 5 dexterous deformable object manipulation tasks for learning contact-rich policies | arXiv · GitHub |
| LEAP Hand | [CoRL 2023] Low-cost dexterous hand platform with benchmark tasks for in-hand manipulation | arXiv · Website |
| MyoSuite | [NeurIPS 2022] Musculoskeletal simulation for physiologically accurate hand and arm control | arXiv · GitHub · Website |
| IsaacGym Benchmark | [NeurIPS 2021] GPU-accelerated dexterous benchmarks including in-hand cube reorientation and pen spinning | arXiv · Website |
| D4RL | [NeurIPS 2020] Offline RL benchmark covering locomotion, manipulation, and navigation with standardized datasets | arXiv · GitHub |
| DeepMind Control Suite | [JMLR 2020] Continuous control tasks for locomotion and manipulation used as standard RL/VLA baselines | arXiv · GitHub |
| Adroit | [RSS 2018] 24-DoF anthropomorphic hand benchmark for pen, door, hammer, and relocate tasks | arXiv · GitHub |
Benchmarks for navigation in 3D environments from language, visual, or object-goal instructions.
| Name | Highlights | References |
|---|---|---|
| GOAT-Bench | [CVPR 2024] Multi-modal lifelong navigation benchmark with sequential object goals specified by category, language, or image in open-vocabulary settings | GitHub · Website |
| OpenEQA | [CVPR 2024] Embodied Question Answering in the era of foundation models with 1600+ questions spanning spatial, object, and functional understanding | arXiv · GitHub · Website |
| ScanQA | [CVPR 2022] Spatial question answering grounded in 3D indoor point cloud environments | arXiv · GitHub |
| SOON | [CVPR 2021] Scenario-oriented object navigation requiring reasoning about object-scene relations | arXiv · Website |
| ImageNav | [ICCV 2021] Navigation to a goal specified by an image rather than a category label | GitHub |
| VLN-BERT / HAMT | [CVPR 2021] BERT-based and history-aware transformer approaches evaluated on R2R and REVERIE | arXiv · GitHub |
| MultiON | [CVPR 2021] Multi-object navigation requiring sequential finding of multiple target objects | arXiv · Website |
| RxR (Room-Across-Rooms) | [EMNLP 2020] Multilingual navigation in 3 languages with denser annotations than R2R | arXiv · GitHub |
| REVERIE | [CVPR 2020] Combines long-horizon navigation with object grounding via referring expressions | arXiv · GitHub · Website |
| ObjectNav (Habitat) | [NeurIPS 2020] Navigate to a specified object category in novel photorealistic environments | arXiv · Website |
| VLN-CE | [EMNLP 2020] Continuous-environment R2R in Habitat without discrete viewpoint teleportation | arXiv · GitHub |
| CVDN | [CVPR 2019] Cooperative multi-turn dialogue navigation where a helper guides an agent to a target | arXiv · GitHub |
| TOUCHDOWN | [CVPR 2019] Street-view navigation in real-world NYC following natural language descriptions | arXiv · GitHub |
| R2R (Room-to-Room) | [CVPR 2018] Foundational VLN benchmark for instruction-following navigation in Matterport3D | arXiv · GitHub · Website |
| EQA (Embodied Question Answering) | [CVPR 2018] Agents navigate and answer visual questions about indoor environments | arXiv · Website |
| MP3D (Matterport3D) | [CVPR 2017] Large-scale RGB-D indoor dataset widely used as a navigation substrate | Website |
Benchmarks for household and open-world agents that combine navigation, manipulation, and task reasoning.
| Name | Highlights | References |
|---|---|---|
| Habitat 3.0 | [ICLR 2024] Social navigation, collaboration, and rearrangement tasks in embodied AI simulation | arXiv · GitHub · Website |
| EmbodiedBench | [ICML 2025] Comprehensive benchmark evaluating MLLMs as embodied agents across 1128 tasks spanning high-level (household, habitat) and low-level (manipulation, navigation) with 6 capability dimensions | arXiv · GitHub · Website |
| ARNOLD | [ICCV 2023] Language-conditioned articulated object state change tasks in 3D scenes | arXiv · GitHub · Website |
| TEACh | [AAAI 2022] Task completion in AI2-THOR via multi-turn dialogue history | arXiv · GitHub |
| ProcTHOR | [NeurIPS 2022] Procedurally generated AI2-THOR houses for training and evaluating generalist embodied agents | arXiv · GitHub · Website |
| ALFWorld | [ICLR 2021] Aligns TextWorld with AI2-THOR for language-conditioned household task completion | arXiv · GitHub · Website |
| ThreeDWorld (TDW) | [NeurIPS 2021] Multi-modal physical simulation for transport, manipulation, and multi-agent tasks | arXiv · GitHub · Website |
| Watch-And-Help | [ICLR 2021] Social intelligence benchmark for agent cooperation in VirtualHome environments | arXiv · GitHub |
| AI2-THOR | [CVPR 2018] Interactive 3D household simulation with physics for pick-and-place, slicing, cooking, and more | arXiv · GitHub · Website |
| VirtualHome | [CVPR 2018] Multi-step household simulation supporting natural language program execution | arXiv · GitHub · Website |
| BEHAVIOR-1K | 1000 household activities grounded in real human needs, simulated in OmniGibson | GitHub · Website |
| LANMP | Language-conditioned navigation and manipulation on a real mobile manipulator | arXiv · GitHub |
Benchmarks for socially aware, collaborative, or communicative robot behavior.
| Name | Highlights | References |
|---|---|---|
| Habitat 3.0 Social Nav | [ICLR 2024] Social navigation requiring robots to move around humans and follow social norms | arXiv · Website |
| ConcepFusion / LangNav | [RSS 2023] Open-vocabulary 3D feature fields enabling natural language queries for navigation | arXiv · GitHub |
| TDW-Social | [NeurIPS 2021] Multi-agent social simulation requiring cooperation and communication in TDW | arXiv · GitHub · Website |
| SEAN | [RA-L 2020] Navigation benchmark requiring socially compliant behavior around pedestrians | arXiv · Website |
| RoboTHOR Challenge | ObjectNav challenge bridging simulation and real robot deployment in home-like environments | arXiv · Website |
| ProCo | Proactive cooperation benchmark requiring initiative and dialogue from embodied agents | GitHub |
Benchmarks for driving agents that ground perception, language, and decision-making in traffic scenes.
| Name | Highlights | References |
|---|---|---|
| DriveLM | [ECCV 2024] Graph-structured vision-language QA benchmark for end-to-end driving on nuScenes | arXiv · GitHub · HF |
| LingoQA | [CVPR 2024] Video QA benchmark for driving scene understanding grounded in real-world footage | arXiv · GitHub |
| Rank2Tell | [WACV 2024] Dataset for ranking and describing important objects in driving scenes | arXiv · GitHub |
| MAPLM | [WACV 2024] Map-based language model benchmark for understanding HD maps in autonomous driving | arXiv · GitHub |
| NuScenes-QA | [AAAI 2024] 460k QA pairs for visual question answering built on nuScenes | arXiv · GitHub |
| TOD3Cap | [ECCV 2024] 3D dense captioning for autonomous driving with 2.3M natural language descriptions | arXiv · GitHub |
| DriveVLM | [CoRL 2024] Chain-of-thought VLM reasoning integrated with motion planning for autonomous driving | arXiv · Website |
| MMAD | [NeurIPS 2024] Massive multimodal AD benchmark with 18 perception and reasoning subtasks | arXiv · GitHub |
| nuScenes | [CVPR 2020] Large-scale AD dataset with 3D bounding boxes, HD maps, and rich multi-sensor data | GitHub · Website |
| Waymo Open Dataset | [CVPR 2020] High-quality driving dataset with LiDAR and camera data for perception and prediction | Website |
| BDD-X | [ECCV 2018] Explainable driving dataset with textual descriptions and justifications for actions | arXiv · GitHub |
| CARLA Challenge | [CoRL 2017] Open urban driving simulation benchmark for route completion and safety compliance | GitHub · Website |
| nuPlan | Closed-loop motion planning benchmark with expert demonstrations and reactive simulation | arXiv · GitHub · Website |
| DriveBench | Comprehensive benchmark evaluating VLMs across multiple driving perception and QA tasks | Website |
| VLAAD | Vision-Language-Action benchmark for autonomous driving with natural language instructions | GitHub |
| LaMPilot | [CVPR 2024] Open benchmark dataset for autonomous driving evaluated via language model programs | arXiv · GitHub |
| STSBench | Spatio-temporal scenario benchmark for evaluating MLLMs in interactive autonomous driving scenarios | GitHub |
| Drive4C | Closed-loop benchmark assessing what foundation models need to enable language-guided autonomous driving | GitHub |
| NAVSIM | [NeurIPS 2024] Non-reactive simulation benchmark for data-driven autonomous driving evaluation with 1000+ hours of real-world driving data | arXiv · GitHub |
| Bench2Drive | Closed-loop multi-ability benchmark for end-to-end autonomous driving across 220 scenarios in CARLA covering perception, prediction, and planning | arXiv · GitHub |
Benchmarks for agents that perceive graphical interfaces and act through keyboard and mouse inputs.
| Name | Highlights | References |
|---|---|---|
| AndroidWorld | [ICLR 2025] 116 programmatic tasks across 20 real Android apps | arXiv · GitHub · Website |
| GUI-Odyssey | [ICCV 2025] Cross-app Android navigation requiring multi-step actions across multiple apps | arXiv · GitHub |
| WebArena | [ICLR 2024] 812 realistic tasks across e-commerce, social, coding, and content websites | arXiv · GitHub · Website |
| VisualWebArena | [ACL 2024] Extends WebArena with visually grounded tasks requiring image-based reasoning | arXiv · GitHub · Website |
| OSWorld | [NeurIPS 2024] 369 open-ended computer tasks in real OS environments (Ubuntu, Windows, macOS) | arXiv · GitHub · HF · Website |
| ScreenSpot | [ACL 2024] GUI grounding benchmark for localizing UI elements from text instructions | arXiv · GitHub |
| AssistGUI | [CVPR 2024] Desktop productivity benchmark (Word, Excel, PowerPoint) with video-based evaluation | arXiv · GitHub · Website |
| AgentBench | [ICLR 2024] LLM-as-agent benchmark with 8 environments: OS, database, web, games, and more | arXiv · GitHub · Website |
| WorkArena | [ICML 2024] 33 enterprise software tasks on a real ServiceNow instance | arXiv · GitHub · Website |
| SWE-bench | [ICLR 2024] Requires agents to resolve real GitHub issues as a software engineering challenge | arXiv · GitHub · Website |
| τ-bench (tau-bench) | [NeurIPS 2024] Tool-agent-user interaction benchmark evaluating agentic task completion with tool use | arXiv · GitHub |
| Mind2Web | [NeurIPS 2023] 2000+ tasks from 137 real-world websites for generalist web agent evaluation | GitHub · Website |
| AITW (Android in the Wild) | [NeurIPS 2023] Large-scale human demonstration dataset on Android devices across diverse tasks | arXiv · GitHub |
| WebSRC | [EMNLP 2021] Structural reading comprehension for understanding web page layouts and content | arXiv · GitHub · Website |
| Screen2Words | [UIST 2021] 112k human-annotated mobile UI screen summaries for screen captioning | arXiv · GitHub |
| MiniWob++ | [ICLR 2018] 100+ browser mini-tasks (clicking, typing, form filling) for web interaction | GitHub · Website |
| WindowsAgentArena | 154 Windows OS tasks spanning productivity, web, and system applications | arXiv · GitHub · Website |
| ScreenSpot-Pro | GUI grounding benchmark for professional high-resolution computer use across complex desktop applications | arXiv · GitHub |
| Spider2-V | [NeurIPS 2024] 494 tasks testing multimodal agents at automating real-world data science and engineering workflows | arXiv · GitHub · Website |
| GUI-World | Comprehensive GUI benchmark for evaluating multimodal agents across mobile, web, and desktop interfaces with sequential tasks and QA scenarios | arXiv · GitHub · Website |
| AndroidLab | [ICLR 2025] Systematic environment and benchmark for training and evaluating Android autonomous agents across 138 tasks in 9 apps | arXiv · GitHub · Website |
| OmniACT | Dataset and benchmark for multimodal autonomous agents performing computer tasks via UI element grounding across desktop and web interfaces | arXiv · GitHub · Website |
| AgentTrek | Automated GUI agent trajectory synthesis benchmark derived from web tutorials with 15K+ step-level annotations for agent training and evaluation | arXiv · GitHub · Website |
Benchmarks for language-guided sequential decision-making in games and interactive simulated worlds.
| Name | Highlights | References |
|---|---|---|
| Voyager / DEPS | [NeurIPS 2023] LLM-powered lifelong learning Minecraft agent evaluation framework | arXiv · GitHub · Website |
| MineRL / MineDojo | [NeurIPS 2022] 3000+ open-ended Minecraft tasks leveraging internet-scale video knowledge | arXiv · GitHub · Website |
| ScienceWorld | [EMNLP 2022] 30 science experiment tasks requiring procedural multi-step reasoning and action | arXiv · GitHub · Website |
| CrafterBench | [TMLR 2022] Open-world survival game measuring 22 achievements requiring long-horizon planning | arXiv · GitHub |
| BEHAVIOR (iGibson) | [NeurIPS 2021] 100 household activities in iGibson simulation requiring physical commonsense reasoning | GitHub · Website |
| NetHack Learning Environment (NLE) | [NeurIPS 2020] Complex roguelike game for long-horizon decision making and language-conditioned play | arXiv · GitHub |
| Atari-HEAD | [AAAI 2019] Human eye-tracking data for Atari games enabling human-like visual attention evaluation | arXiv · GitHub |
| BabyAI | [ICLR 2019] Gridworld benchmark with 19 procedurally generated instruction difficulty levels | arXiv · GitHub |
| GROOT | Open-ended Minecraft agent benchmark with skill learning and generalization | GitHub |
| MiniGrid / MiniWorld | Lightweight 2D/3D grid environments for language-conditioned navigation and reasoning | GitHub |
| TextWorld | Framework for generating and playing text-based adventure games to train language agents | arXiv · GitHub |
Benchmarks for multimodal perception, temporal understanding, and grounding capabilities that support VLA systems.
Benchmarks for general multimodal reasoning, instruction following, and cross-modal grounding.
| Name | Highlights | References |
|---|---|---|
| MMMU | [CVPR 2024] 11.5k expert-level questions across 57 subjects for massive multidisciplinary evaluation | arXiv · GitHub · HF · Website |
| MMBench | [ECCV 2024] Systematic 20-dimension VLM evaluation with 3000 questions | arXiv · GitHub · Website |
| SEED-Bench | [CVPR 2024] 19K questions across 12 dimensions for multi-granularity generative MLLM evaluation | arXiv · GitHub |
| HallusionBench | [CVPR 2024] Hallucination detection benchmark for visual understanding in VLMs | arXiv · GitHub |
| MMStar | [NeurIPS 2024] 1500 carefully curated samples minimizing data leakage and requiring genuine VL reasoning | arXiv · GitHub · Website |
| CV-Bench | [EMNLP 2024] Compositional reasoning and spatial understanding benchmark for VLMs | GitHub |
| BLINK | [ECCV 2024] Visual perception benchmark targeting human intuitive abilities that challenge VLMs | arXiv · GitHub · Website |
| WildVision | [NeurIPS 2024] Real-world VLM evaluation via human preference collected from live user interactions | HF |
| VLMEvalKit | [ACM MM 2024] Open-source toolkit supporting 100+ VLMs across 50+ benchmarks | arXiv · GitHub |
| POPE | [EMNLP 2023] Polling-based object probing for evaluating object hallucination in VLMs | arXiv · GitHub |
| MME | 14-subtask benchmark covering MLLM perception and cognition capabilities | arXiv · GitHub |
| TouchStone | GPT-4 scored VLM evaluation across 5 ability dimensions | arXiv · GitHub |
| RealWorldQA | Real-world spatial understanding benchmark from xAI testing physical world reasoning | Website |
| MMTE | Multimodal task execution benchmark for following complex visual instructions | GitHub |
| MMIU | [ICLR 2025] Multimodal Multi-image Understanding benchmark with 7 image-relationship types, 52 tasks, and 11K questions for evaluating LVLMs | arXiv · GitHub · Website |
| MathVista | [ICLR 2024] Mathematical reasoning in visual contexts benchmark requiring compositional reasoning across 28 math tasks and 5 visual types | arXiv · GitHub · Website |
| MMMU-Pro | [NeurIPS 2024] More robust MMMU extension with college-level expert questions requiring genuine multimodal understanding, resistant to language-only shortcuts | arXiv · GitHub · Website |
| CharXiv | [NeurIPS 2024] Chart understanding benchmark exposing significant gaps in VLMs on 2323 scientific paper figures with descriptive and reasoning tasks | arXiv · GitHub · Website |
| OlympiadBench | [ACL 2024] Olympiad-level bilingual multimodal science benchmark with 8952 problems in mathematics and physics for frontier VLM evaluation | arXiv · GitHub |
| MEGA-Bench | [NeurIPS 2024] Scaling multimodal evaluation to 500+ diverse real-world tasks covering perception, reasoning, and generation across many domains | arXiv · GitHub · Website |
| MM-Vet v2 | Harder and more comprehensive integrated capabilities benchmark for large multimodal models using GPT-4 scoring across 16 skills | arXiv · GitHub |
Benchmarks for temporal video understanding, action reasoning, and long-context event interpretation.
| Name | Highlights | References |
|---|---|---|
| TemporalBench | [ICLR 2025] Fine-grained temporal video understanding benchmark for VLMs with 2M QA pairs | arXiv · GitHub |
| Video-MME | [CVPR 2025] First comprehensive evaluation benchmark of MLLMs in video analysis covering short/medium/long videos across 30 domains | arXiv · GitHub · Website |
| LongVideoBench | [NeurIPS 2024] Long-context interleaved video-language understanding benchmark with 6678 questions on videos ranging from minutes to 1 hour | arXiv · GitHub · Website |
| MLVU | [NeurIPS 2024] Multi-task long video understanding benchmark with 9 task types spanning holistic, single-detail, and multi-detail evaluation across 10–180 minute videos | arXiv · GitHub |
| MVBench | [CVPR 2024] 20 challenging temporal reasoning tasks for multi-task video understanding | arXiv · GitHub |
| EgoTaskQA | [NeurIPS 2022] Egocentric video QA with task goal inference, procedural reasoning, and state tracking | arXiv · GitHub |
| STAR (Situated Reasoning) | [NeurIPS 2021] Situated temporal action reasoning with 4 question types grounded in real videos | arXiv · Website |
| NExT-QA | [CVPR 2021] Video QA benchmark emphasizing causal and temporal reasoning about actions | arXiv · GitHub · Website |
| COIN | [CVPR 2019] 11,827 instructional videos across 180 tasks and 778 procedural steps | arXiv · GitHub · Website |
| CrossTask | [CVPR 2019] Procedural activity understanding across 83 instructional video tasks | arXiv · GitHub |
| HowTo100M | [ICCV 2019] Narrated instructional videos for learning action-language grounding at scale | arXiv · Website |
| Something-Something v2 | [ICCV 2017] Temporal reasoning requiring understanding of object interactions and motion direction | arXiv · Website |
| Charades | [ECCV 2016] Indoor activity dataset requiring temporal localization and compositional reasoning | arXiv · Website |
| ActivityNet | [CVPR 2015] Activity recognition, localization, and dense video captioning at large scale | arXiv · Website |
| Kinetics-700 | Large-scale action recognition dataset with 700 human action classes | arXiv · Website |
| Video-Bench | Video-exclusive benchmark for evaluating video understanding capabilities in MLLMs | arXiv · GitHub |
| VideoVista | Diverse video understanding benchmark spanning 14 task types and 35 categories | arXiv · GitHub |
| DREAM-1K | Procedural video understanding benchmark for evaluating robot task planning | arXiv · GitHub |
Benchmarks for egocentric perception and activity understanding from first-person video.
| Name | Highlights | References |
|---|---|---|
| Ego-Exo4D | [CVPR 2024] Paired egocentric and exocentric video dataset for skill assessment and correspondence | arXiv · Website |
| Ego4D | [CVPR 2022] 3670h egocentric video with benchmarks for episodic memory, forecasting, and hand-object interaction | arXiv · GitHub · Website |
| EPIC-Kitchens-100 | [IJCV 2022] 100 hours of unscripted egocentric cooking actions with fine-grained annotations | arXiv · GitHub · Website |
| Assembly101 | [CVPR 2022] Egocentric procedural dataset for assembling and disassembling 101 toy vehicles | arXiv · GitHub · Website |
| HOI4D | [CVPR 2022] Egocentric 4D dataset for category-level human-object interaction manipulation | arXiv · GitHub · Website |
| EgoProceL | [ECCV 2022] Egocentric procedural learning benchmark for keystep recognition from instructional videos | arXiv · GitHub |
| LEMMA | [ECCV 2020] Multi-person, multi-task dataset for compositional action understanding | arXiv · Website |
| EGTEA Gaze+ | [ECCV 2018] First-person cooking dataset with gaze and hand masks for fine-grained recognition | arXiv · Website |
| EPIC-Kitchens Challenges | Annual challenges on action recognition, detection, anticipation, and retrieval | Website |
| EgoPlan-Bench | [IJCV 2024] Benchmarking MLLMs for human-level task planning from egocentric video with real-world household scenarios | arXiv · GitHub |
Benchmarks for grounding language and reasoning in 3D scenes for embodied perception.
| Name | Highlights | References |
|---|---|---|
| EmbodiedScan | [CVPR 2024] Holistic 3D perception benchmark with RGB-D input from egocentric views | arXiv · GitHub · Website |
| LEO | [ICML 2024] Embodied generalist benchmark requiring 3D scene understanding for grounded dialogue and planning | arXiv · GitHub · Website |
| SQA3D (Situated QA) | [ICLR 2023] Situated reasoning benchmark where agents answer questions from a 3D first-person perspective | arXiv · GitHub · Website |
| 3D-VisTA | [ICCV 2023] Pre-trained transformer for 3D VL tasks: grounding, QA, and dense captioning | arXiv · GitHub · Website |
| MultiScan | [NeurIPS 2022] Multi-scan indoor reconstruction with articulated object annotations for VLA research | arXiv · GitHub · Website |
| ScanEnts3D | [ECCV 2022] Links 3D scene entities to text mentions for fine-grained object grounding | arXiv · GitHub |
| ScanRefer | [ECCV 2020] Localizes objects in 3D point clouds using free-form language descriptions | arXiv · GitHub · Website |
| CLEVR3D | 3D extension of CLEVR for spatial reasoning questions in synthetic environments | GitHub |
| Nu-Scenes QA | QA benchmark grounded in outdoor 3D LiDAR scenes for driving reasoning | arXiv · GitHub |
| MMScan | [NeurIPS 2024] Multi-modal 3D scene dataset with 1.4M hierarchical grounded language annotations on 109k objects and benchmarks for visual grounding and QA | arXiv · GitHub · Website |
| SceneVerse | [ECCV 2024] Large-scale 3D vision-language dataset with 68K indoor scenes and 2.5M language-scene pairs for grounded scene understanding and spatial reasoning | arXiv · GitHub · Website |
Benchmarks for visual grounding, referring expressions, and spatial language understanding.
| Name | Highlights | References |
|---|---|---|
| SeedBench-2-Plus | [NeurIPS 2024] Extended SEED-Bench focusing on charts, maps, web pages, and document comprehension | arXiv · GitHub |
| ARO (Attribution, Relation, Order) | [ICLR 2023] Reveals VLMs' limited sensitivity to word order and relational structure | arXiv · GitHub |
| VSR (Visual Spatial Reasoning) | [NAACL 2022] True/false benchmark testing spatial relationships between objects in images | arXiv · GitHub |
| Winoground | [CVPR 2022] Tests visio-linguistic compositional reasoning via image-caption matching with swapped words | arXiv · HF |
| FLUTE | [EMNLP 2022] Figurative language understanding for VLMs involving metaphors and similes | arXiv · GitHub |
| GQA | [CVPR 2019] Compositional QA benchmark requiring multi-step spatial and relational reasoning | arXiv · Website |
| Visual Genome | [IJCV 2017] Densely annotated images with region descriptions, QA, and scene graphs for grounded understanding | arXiv · GitHub · Website |
| CLEVR | [CVPR 2017] Compositional language and elementary visual reasoning diagnostic benchmark | arXiv · GitHub · Website |
| RefCOCO / RefCOCO+ / RefCOCOg | [ECCV 2016] Referring expression comprehension benchmarks for localizing objects from natural language | arXiv · GitHub |
| SpatialBot | Spatial understanding benchmark evaluating depth-aware reasoning in VLMs | GitHub |
Supporting resources used to run, compare, and reproduce VLA benchmark results.
Simulation platforms and task environments used to build or execute VLA benchmarks.
| Name | Highlights | References |
|---|---|---|
| IsaacGym | [NeurIPS 2021] Legacy GPU physics simulation for massively parallel RL training on dexterous tasks | arXiv · Website |
| Robosuite | [IROS 2021] Modular robot learning framework built on MuJoCo with standardized task suites | arXiv · GitHub · Website |
| SAPIEN | [CVPR 2020] Simulation platform for articulated object manipulation with PhysX physics | arXiv · GitHub · Website |
| Habitat-Sim | [ICCV 2019] High-performance 3D simulator for embodied AI with photorealistic rendering and physics | arXiv · GitHub · Website |
| CARLA | [CoRL 2017] Open urban driving simulator with rich sensor modalities for autonomous driving research | arXiv · GitHub · Website |
| MuJoCo | [IROS 2012] Fast and accurate physics simulation engine widely used for locomotion and manipulation | arXiv · GitHub · Website |
| IsaacSim / IsaacLab | NVIDIA's GPU-accelerated simulation platform for robot learning and VLA development | GitHub · Website |
| Genesis | Generative simulation engine with ultra-fast physics and photorealistic rendering | arXiv · GitHub · Website |
| PyBullet / Panda-Gym | Open-source physics simulation with Panda robot gym environments for tabletop manipulation | GitHub · Website |
| OmniGibson | NVIDIA Omniverse-based simulation for BEHAVIOR-1K with realistic physics and rendering | arXiv · GitHub · Website |
| AI2-THOR (Simulator) | Photorealistic interactive indoor simulation for embodied household task research | arXiv · GitHub · Website |
| CoppeliaSim (V-REP) | Versatile robot simulation platform underlying RLBench and other benchmarks | Website |
| SUMO | Microscopic traffic simulation supporting multi-modal transportation research | GitHub · Website |
| XR-EgoBench | Extended reality egocentric benchmark platform for AR/VR embodied agent evaluation | GitHub |
Datasets and model hubs that provide training assets, benchmark data, or evaluation-facing artifacts for VLA work.
| Name | Highlights | References |
|---|---|---|
| HuggingFace LeRobot Datasets | Standardized real-robot demonstration datasets for VLA training and benchmarking | GitHub · HF |
| HuggingFace Open-VLA | HuggingFace hub for OpenVLA model weights, training data, and evaluation scripts | GitHub · HF |
| HuggingFace Robot Benchmarks | Curated collection of robot learning benchmarks and evaluation protocols | HF |
| Open X-Embodiment Dataset | HuggingFace mirror of the OXE cross-embodiment dataset aggregating 1M+ robot episodes | arXiv · HF |
| RoboSet | Large-scale robot manipulation dataset with 100k+ demonstrations across 12 task categories | arXiv · Website |
| DROID Dataset | 76k robot manipulation demonstrations on HuggingFace for diverse real-world VLA training | arXiv · HF |
| BridgeData V2 (HF) | HuggingFace version of BridgeData V2 for easy access and VLA pretraining | arXiv · HF |
| EgoMimic Dataset | Paired egocentric human video and robot demonstration dataset for VLA transfer learning | arXiv · Website |
| RH20T | Large-scale robotic dataset with 110k demonstrations across 20 tasks | arXiv · GitHub · Website |
| Ego4D HuggingFace | Ego4D benchmark data accessible via HuggingFace for convenient evaluation | arXiv · HF |
Leaderboards and evaluation services that host benchmark results, competitions, or public comparisons.
| Name | Highlights | References |
|---|---|---|
| EvalAI | Cloud-based challenge platform hosting embodied AI, VQA, navigation, and VLA competitions | Website |
| PaperWithCode Embodied AI | Live leaderboards tracking SOTA across navigation, manipulation, and QA benchmarks | Website |
| HuggingFace Open VLA Leaderboard | Community leaderboard for real-world robot VLA performance | HF |
| Open-Compass (OpenVLM Leaderboard) | Comprehensive VLM leaderboard comparing models across 50+ benchmarks | Website |
| LIBERO Leaderboard | Official leaderboard for the LIBERO manipulation benchmark suite | Website |
| CARLA Autonomous Driving Leaderboard | Official competition leaderboard for autonomous driving agents in CARLA | Website |
| WebArena Leaderboard | Community tracking sheet for web agent performance on WebArena | Website |
| OSWorld Leaderboard | Live leaderboard for OS computer-use agents on the OSWorld benchmark | Website |
| Ego4D Challenges (EvalAI) | Annual Ego4D challenge tracks on episodic memory, forecasting, AV, hands, and narrations | Website |
| BEHAVIOR-1K Leaderboard | Evaluation results for household activity agents in the BEHAVIOR benchmark | Website |
| AgentBench Leaderboard | Rankings for LLM-as-agent performance across OS, web, game, and database tasks | Website |
| EmbodiedBench Leaderboard | CVPR 2026 challenge leaderboard for MLLM-based embodied agents across high-level and low-level tasks | Website |
See CONTRIBUTING.md for repository scope, placement rules, row format, sorting expectations, and the pull request checklist.
This list is released under the CC0 1.0 Universal public domain dedication. You are free to copy, modify, and distribute this work, even for commercial purposes, without asking permission.
13 commits
5 commits