cpystan/MSMU

Dataset

2

stars

18

commits

1

linked in READMEs

Jun 3, 2026

updated

README

MSMU (Massive Spatial Measuring and Understanding Dataset for Spatial Intelligence)

🌐 Homepage | 🤗 Dataset | 📖 arXiv | GitHub

Dataset Details

Dataset Description

We introduce MSMU and MSMU-Bench: a new benchmark designed to enhance and evaluate multimodal models on spatial measuring and understanding. MSMU is featured as metric-accurate spatial annotations which are sourced from high-precision 3D scenes. It contains , 25K images, 700K QA pairs, and 2.5M numerical values, covering a wide range of quantitative spatial tasks (Existence, Counting, Scale Estimation, Grounding, Relative Position, Absolute Distance, Scale Comparison, and Reference Object Estimation).

🎯 We have released a full set of MSMU and MSMU-Bench.

Dataset Creation

We categorize the spatial tasks in MSMU into 8 types, the distribution of which is illustrated in Figure below (left). The QA distribution of MSMU-Bench is also shown in Figure below (right) which provides a detailed breakdown of these eight categories.

🏆 Mini-Leaderboard

We show a mini-leaderboard here. It shows the results of each sub-category and the overall performance.

Results

ModelExistenceObject
Counting
Scale
Est.
GroundingRelative
Position
Absolute
Distance
Scale
Comparison
Ref. Object
Est.
Average
Large Language Models (LLMs): Text only
GPT-4-Turbo12.765.2113.5112.6424.847.5036.7912.0415.66
Qwen2.54.250.000.7813.790.620.0016.041.574.63
DeepSeek-V30.005.241.546.9010.560.0025.475.247.39
Vision-Language Models (VLMs): Image + Text
GPT-4o44.6841.673.8627.5967.0820.0054.722.0932.28
Gemini-238.3043.7523.9419.5454.6612.5069.8118.8535.17
Qwen2.5-VL-72B59.5735.421.5413.7957.762.5066.049.9530.82
Qwen2.5-VL-32B29.7941.6710.8118.3960.252.5046.2310.9927.59
Qwen2.5-VL-7B12.764.170.001.151.240.005.660.523.19
Intern-VL3-78B47.6242.716.4726.3256.9413.3364.1016.4633.63
Intern-VL3-8B36.1741.674.6318.3960.252.5049.068.3828.54
LLaVA-1.5-7B1.5436.465.0220.6942.865.0038.680.5219.45
Depth-encoded VLMs: Image + Depth + Text
SpatialBot10.6446.8815.8328.7466.465.0050.948.9029.17
SpatialRGPT10.6436.4620.0817.2460.2515.0062.269.9528.98
Ours87.2347.9251.3542.5375.1640.0055.6646.0756.31

Citation

BibTeX:

@inproceedings{chen2025sdvlm,
      title={SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models}, 
      author={Pingyi Chen and Yujing Lou and Shen Cao and Jinhui Guo and Lubin Fan and Yue Wu and Lin Yang and Lizhuang Ma and Jieping Ye},
      booktitle={NeurIPS},
      year={2025},
}

Contributors

cpystan

18 commits

cpystan/MSMU

Dataset

2

stars

18

commits

1

linked in READMEs

Jun 3, 2026

updated

README

MSMU (Massive Spatial Measuring and Understanding Dataset for Spatial Intelligence)

🌐 Homepage | 🤗 Dataset | 📖 arXiv | GitHub

Dataset Details

Dataset Description

We introduce MSMU and MSMU-Bench: a new benchmark designed to enhance and evaluate multimodal models on spatial measuring and understanding. MSMU is featured as metric-accurate spatial annotations which are sourced from high-precision 3D scenes. It contains , 25K images, 700K QA pairs, and 2.5M numerical values, covering a wide range of quantitative spatial tasks (Existence, Counting, Scale Estimation, Grounding, Relative Position, Absolute Distance, Scale Comparison, and Reference Object Estimation).

🎯 We have released a full set of MSMU and MSMU-Bench.

Dataset Creation

We categorize the spatial tasks in MSMU into 8 types, the distribution of which is illustrated in Figure below (left). The QA distribution of MSMU-Bench is also shown in Figure below (right) which provides a detailed breakdown of these eight categories.

🏆 Mini-Leaderboard

We show a mini-leaderboard here. It shows the results of each sub-category and the overall performance.

Results

ModelExistenceObject
Counting
Scale
Est.
GroundingRelative
Position
Absolute
Distance
Scale
Comparison
Ref. Object
Est.
Average
Large Language Models (LLMs): Text only
GPT-4-Turbo12.765.2113.5112.6424.847.5036.7912.0415.66
Qwen2.54.250.000.7813.790.620.0016.041.574.63
DeepSeek-V30.005.241.546.9010.560.0025.475.247.39
Vision-Language Models (VLMs): Image + Text
GPT-4o44.6841.673.8627.5967.0820.0054.722.0932.28
Gemini-238.3043.7523.9419.5454.6612.5069.8118.8535.17
Qwen2.5-VL-72B59.5735.421.5413.7957.762.5066.049.9530.82
Qwen2.5-VL-32B29.7941.6710.8118.3960.252.5046.2310.9927.59
Qwen2.5-VL-7B12.764.170.001.151.240.005.660.523.19
Intern-VL3-78B47.6242.716.4726.3256.9413.3364.1016.4633.63
Intern-VL3-8B36.1741.674.6318.3960.252.5049.068.3828.54
LLaVA-1.5-7B1.5436.465.0220.6942.865.0038.680.5219.45
Depth-encoded VLMs: Image + Depth + Text
SpatialBot10.6446.8815.8328.7466.465.0050.948.9029.17
SpatialRGPT10.6436.4620.0817.2460.2515.0062.269.9528.98
Ours87.2347.9251.3542.5375.1640.0055.6646.0756.31

Citation

BibTeX:

@inproceedings{chen2025sdvlm,
      title={SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models}, 
      author={Pingyi Chen and Yujing Lou and Shen Cao and Jinhui Guo and Lubin Fan and Yue Wu and Lin Yang and Lizhuang Ma and Jieping Ye},
      booktitle={NeurIPS},
      year={2025},
}

Contributors

cpystan

18 commits