Jack Hong1, Shilin Yan1†, Jiayin Cai1, Xiaolong Jiang1, Yao Hu1, Weidi Xie2‡
5
14 commits
2 linked in READMEs
updated Jul 12, 2026
Jack Hong1, Shilin Yan1†, Jiayin Cai1, Xiaolong Jiang1, Yao Hu1, Weidi Xie2‡
1Xiaohongshu Inc. 2Shanghai Jiao Tong University
2025.02.07 🌟 We release WorldSense, the first benchmark for real-world omnimodal understanding of MLLMs.we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features:
Based on our WorldSense, we extensively evaluate various state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding real-world scenarios (48% best accuracy). We hope our WorldSense can provide a platform for evaluating the ability in constructing and understanding coherent contexts from omni-modality.
Please download our WorldSense from here.
📍 Evaluation: Thanks for the reproduction of our evaluation through VLMEvalkit. Please refer to VLMEvalkit for details.
📍 Leaderboard:
If you want to add your model to our leaderboard, please contact jaaackhong@gmail.com.
If you find WorldSense helpful for your research, please consider citing our work. Thanks!
@article{hong2025worldsenseevaluatingrealworldomnimodal,
title={WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs},
author={Jack Hong and Shilin Yan and Jiayin Cai and Xiaolong Jiang and Yao Hu and Weidi Xie},
year={2025},
eprint={2502.04326},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2502.04326},
}
14 commits
Jack Hong1, Shilin Yan1†, Jiayin Cai1, Xiaolong Jiang1, Yao Hu1, Weidi Xie2‡
5
14 commits
2 linked in READMEs
updated Jul 12, 2026
Jack Hong1, Shilin Yan1†, Jiayin Cai1, Xiaolong Jiang1, Yao Hu1, Weidi Xie2‡
1Xiaohongshu Inc. 2Shanghai Jiao Tong University
2025.02.07 🌟 We release WorldSense, the first benchmark for real-world omnimodal understanding of MLLMs.we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features:
Based on our WorldSense, we extensively evaluate various state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding real-world scenarios (48% best accuracy). We hope our WorldSense can provide a platform for evaluating the ability in constructing and understanding coherent contexts from omni-modality.
Please download our WorldSense from here.
📍 Evaluation: Thanks for the reproduction of our evaluation through VLMEvalkit. Please refer to VLMEvalkit for details.
📍 Leaderboard:
If you want to add your model to our leaderboard, please contact jaaackhong@gmail.com.
If you find WorldSense helpful for your research, please consider citing our work. Thanks!
@article{hong2025worldsenseevaluatingrealworldomnimodal,
title={WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs},
author={Jack Hong and Shilin Yan and Jiayin Cai and Xiaolong Jiang and Yao Hu and Weidi Xie},
year={2025},
eprint={2502.04326},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2502.04326},
}
14 commits