WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
See the codeJack Hong1, Shilin Yan1†, Jiayin Cai1, Xiaolong Jiang1, Yao Hu1, Weidi Xie2‡
1Xiaohongshu Inc. 2Shanghai Jiao Tong University
We welcome researchers and developers from the community to submit their models for evaluation and inclusion on the WorldSense leaderboard. To streamline the process, please send an email to jaaackhong@gmail.com and CC tattoo.ysl@gmail.com.
we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features:
Based on our WorldSense, we extensively evaluate various state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding real-world scenarios (48% best accuracy). We hope our WorldSense can provide a platform for evaluating the ability in constructing and understanding coherent contexts from omni-modality.
Please download our WorldSense from here.
📍 Evaluation: Thanks for the reproduction of our evaluation through VLMEvalkit. Please refer to VLMEvalkit for details.
This dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
You are free to use, copy, modify, distribute, and build upon this dataset, including for commercial purposes, provided that appropriate credit is given to the original source.
If you find WorldSense helpful for your research, please consider citing our work. Thanks!
@article{hong2025worldsenseevaluatingrealworldomnimodal,
title={WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs},
author={Jack Hong and Shilin Yan and Jiayin Cai and Xiaolong Jiang and Yao Hu and Weidi Xie},
year={2025},
eprint={2502.04326},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2502.04326},
}
JavaScript
58.8%
HTML
31.8%
CSS
9.3%
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
See the codeJack Hong1, Shilin Yan1†, Jiayin Cai1, Xiaolong Jiang1, Yao Hu1, Weidi Xie2‡
1Xiaohongshu Inc. 2Shanghai Jiao Tong University
We welcome researchers and developers from the community to submit their models for evaluation and inclusion on the WorldSense leaderboard. To streamline the process, please send an email to jaaackhong@gmail.com and CC tattoo.ysl@gmail.com.
we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features:
Based on our WorldSense, we extensively evaluate various state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding real-world scenarios (48% best accuracy). We hope our WorldSense can provide a platform for evaluating the ability in constructing and understanding coherent contexts from omni-modality.
Please download our WorldSense from here.
📍 Evaluation: Thanks for the reproduction of our evaluation through VLMEvalkit. Please refer to VLMEvalkit for details.
This dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
You are free to use, copy, modify, distribute, and build upon this dataset, including for commercial purposes, provided that appropriate credit is given to the original source.
If you find WorldSense helpful for your research, please consider citing our work. Thanks!
@article{hong2025worldsenseevaluatingrealworldomnimodal,
title={WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs},
author={Jack Hong and Shilin Yan and Jiayin Cai and Xiaolong Jiang and Yao Hu and Weidi Xie},
year={2025},
eprint={2502.04326},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2502.04326},
}
JavaScript
58.8%
HTML
31.8%
CSS
9.3%