STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
5
9 commits
1 linked in READMEs
updated May 16, 2025
We leveraged our STSBench framework to construct a benchmark from the nuScenes dataset for current expert driving models about their spatio-temporal reasoning capabilities in traffic scenes. In particular, we automatically gathered and manually verified scenarios from all 150 scenes of the validation set, considering only annotated key frames. In contrast to prior benchmarks, focusing primarily on ego-vehicle actions that mainly occur in the front-view, STSnu evaluates spatio-temporal reasoning across a broader set of interactions and multiple views. This includes reasoning about other agents and their interactions with the ego-vehicle or with one another. To support this, we define four distinct scenario categories:
Ego-vehicle scenarios. The first category includes all actions related exclusively to the ego-vehicle, such as acceleration/deceleration, left/right turn, or lane change. Important for control decisions and collision prevention, driving models must be aware of the ego-vehicle status and behavior. Although these scenarios are part of existing benchmarks in different forms and relatively straightforward to detect, they provide valuable negatives for scenarios with ego-agent interactions.
Agent scenarios. Similar to ego-vehicle scenarios, agent scenarios involve a single object. However, this category additionally contains vulnerable road users such as pedestrians and cyclists. Pedestrians, contrary to vehicles, perform actions such as walking, running, or crossing. Awareness of other traffic participants and their actions is crucial when it comes to risk assessment, planning the next ego action, or analyzing the situation in a dynamic environment. In contrast to ego-vehicle actions, other road users may be occluded or far away and, therefore, pose a particular challenge.
Ego-to-agent scenarios. The third category of scenarios describes ego-related agent actions. Directly influencing the driving behavior of each other, this category is similarly important to the ego-vehicle scenarios with respect to the immediate control decisions. Ego-agent scenarios contain maneuvers such as overtaking, passing, following, or leading. The scenarios focus on agents in the immediate vicinity of the ego-vehicle and direct interactions.
Agent-to-agent scenarios. The most challenging group of scenarios concerns interactions between two agents, not considering the ego-vehicle. These scenarios describe the spatio-temporal relationship between objects. For instance, a vehicle that overtakes another vehicle in motion or pedestrians moving alongside each other. The latter is a perfect example of interactions that do not actively influence the driving behavior of the expert model. However, we argue that a holistic understanding of the scene should not be restricted to the immediate surroundings of the ego-vehicle.
STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
5
9 commits
1 linked in READMEs
updated May 16, 2025
We leveraged our STSBench framework to construct a benchmark from the nuScenes dataset for current expert driving models about their spatio-temporal reasoning capabilities in traffic scenes. In particular, we automatically gathered and manually verified scenarios from all 150 scenes of the validation set, considering only annotated key frames. In contrast to prior benchmarks, focusing primarily on ego-vehicle actions that mainly occur in the front-view, STSnu evaluates spatio-temporal reasoning across a broader set of interactions and multiple views. This includes reasoning about other agents and their interactions with the ego-vehicle or with one another. To support this, we define four distinct scenario categories:
Ego-vehicle scenarios. The first category includes all actions related exclusively to the ego-vehicle, such as acceleration/deceleration, left/right turn, or lane change. Important for control decisions and collision prevention, driving models must be aware of the ego-vehicle status and behavior. Although these scenarios are part of existing benchmarks in different forms and relatively straightforward to detect, they provide valuable negatives for scenarios with ego-agent interactions.
Agent scenarios. Similar to ego-vehicle scenarios, agent scenarios involve a single object. However, this category additionally contains vulnerable road users such as pedestrians and cyclists. Pedestrians, contrary to vehicles, perform actions such as walking, running, or crossing. Awareness of other traffic participants and their actions is crucial when it comes to risk assessment, planning the next ego action, or analyzing the situation in a dynamic environment. In contrast to ego-vehicle actions, other road users may be occluded or far away and, therefore, pose a particular challenge.
Ego-to-agent scenarios. The third category of scenarios describes ego-related agent actions. Directly influencing the driving behavior of each other, this category is similarly important to the ego-vehicle scenarios with respect to the immediate control decisions. Ego-agent scenarios contain maneuvers such as overtaking, passing, following, or leading. The scenarios focus on agents in the immediate vicinity of the ego-vehicle and direct interactions.
Agent-to-agent scenarios. The most challenging group of scenarios concerns interactions between two agents, not considering the ego-vehicle. These scenarios describe the spatio-temporal relationship between objects. For instance, a vehicle that overtakes another vehicle in motion or pedestrians moving alongside each other. The latter is a perfect example of interactions that do not actively influence the driving behavior of the expert model. However, we argue that a holistic understanding of the scene should not be restricted to the immediate surroundings of the ego-vehicle.