FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
7
10 commits
1 linked in READMEs
updated Jan 22, 2026
Predicting the future requires listening as well as seeing.
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed to evaluate omni-modal future forecasting from audio–visual environments. To succeed, models must perform cross-modal causal and temporal reasoning while effectively leveraging internal knowledge to predict future events.
The dataset consists of 1,034 high-quality multiple-choice QA pairs over 919 videos.
from datasets import load_dataset
# Load the benchmark evaluation set
dataset_test = load_dataset("OpenMOSS-Team/FutureOmni", split="test")
print(dataset_test[0])
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
7
10 commits
1 linked in READMEs
updated Jan 22, 2026
Predicting the future requires listening as well as seeing.
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed to evaluate omni-modal future forecasting from audio–visual environments. To succeed, models must perform cross-modal causal and temporal reasoning while effectively leveraging internal knowledge to predict future events.
The dataset consists of 1,034 high-quality multiple-choice QA pairs over 919 videos.
from datasets import load_dataset
# Load the benchmark evaluation set
dataset_test = load_dataset("OpenMOSS-Team/FutureOmni", split="test")
print(dataset_test[0])