EN | δΈζ
π This is the official full-scale training dataset of the SenseNova-SI series. SenseNova-SI-8M contains ~8.16 million carefully curated training samples spanning ~2.72 million unique images, organized under a rigorous taxonomy of spatial capabilities. It is the dataset used to train the recommended released model SenseNova-SI-1.1-InternVL3-8B and serves as the canonical training corpus for spatial intelligence research with SenseNova-SI.
Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation models to cultivate spatial intelligence within the SenseNova-SI family, built upon established multimodal foundations including visual understanding models (i.e., Qwen3-VL and InternVL3) and unified understanding and generation models (i.e., Bagel). We take a principled approach to constructing high-performing and robust spatial intelligence by systematically curating SenseNova-SI-8M: eight million diverse data samples under a rigorous taxonomy of spatial capabilities. SenseNova-SI demonstrates unprecedented performance across a broad range of spatial intelligence benchmarks, while maintaining strong general multimodal understanding. More importantly, we analyze the impact of data scaling, discuss early signs of emergent generalization capabilities enabled by diverse data training, analyze the risk of overfitting and language shortcuts, present a preliminary study on spatial chain-of-thought reasoning, and validate the potential downstream application. SenseNova-SI is an ongoing project, and this report will be updated continuously. All newly trained multimodal foundation models are publicly released to facilitate further research in this direction. In the future, SenseNova-SI will be integrated with larger-scale in-house models.
SenseNova-SI-8M is the official, full-scale training dataset of the SenseNova-SI series. The recommended model SenseNova-SI-1.1-InternVL3-8B is trained on this dataset and is our primary recommended release for spatial intelligence research and downstream use.
The previously released SenseNova-SI-800K subset is provided only as a reference for scaling-law analysis; the model trained on it is not the recommended release.
| Model | SI Dataset | Avg | VSI | MMSI | MindCube-Tiny | ViewSpatial | SITE | BLINK | 3DSRBench | EmbSpatial |
|---|---|---|---|---|---|---|---|---|---|---|
| InternVL3-8B | - | 45.7 | 42.1 | 28.0 | 41.5 | 38.7 | 41.1 | 53.5 | 44.3 | 76.3 |
| VST-7B-SFT | VST-P-4.1M | 50.8 | 55.5 | 32.5 | 39.7 | 50.5 | 39.7 | 61.9 | 53.1 | 73.7 |
| Cambrian-S-7B | VSI-590K | 45.1 | 62.9 | 27.1 | 37.9 | 41.3 | 36.1 | 37.9 | 45.0 | 72.8 |
| *SenseNova-SI-1.1-InternVL3-8B-800K | SenseNova-SI-800K (scaling-law reference) | - | 60.9 | 36.4 | 56.9 | 52.5 | 47.7 | - | - | - |
| SenseNova-SI-1.1-InternVL3-8B (recommended) | SenseNova-SI-8M (this dataset) | 61.5 | 68.8 | 43.3 | 85.7 | 54.7 | 47.7 | 63.9 | 55.5 | 72.0 |
| Aspect | SenseNova-SI-800K | SenseNova-SI-8M (this release) |
|---|---|---|
| Role | Scaling-law reference subset | Official full training set |
| # Training samples | ~832K | ~8.16M (β 800K Γ 10) |
| # Unique images | ~410K | ~2.72M |
| Total image storage | ~370 GB | ~1.1 TB |
| Image archive | 93 Γ 4 GB independent zips | 53 independent zips (~21 GB each, last ~7 GB) |
| Annotation file | SenseNova-SI-800K.jsonl | SenseNova-SI-8M.jsonl |
| Viewer preview | 1,000-sample parquet | SenseNova-SI-8M_1000samples.parquet |
| Recommended trained model | (reference only) | SenseNova-SI-1.1-InternVL3-8B |
All zip archives are independent (not split volumes) β each can be extracted on its own and together they reconstruct the full images/ tree. One-click extraction scripts are included for both Linux/macOS and Windows (see Download & Extract Images).
The data is stored in SenseNova-SI-8M.jsonl using the JSONL (JSON Lines) format, where each line represents an independent data entry. Each entry is a dictionary organized in the following format, containing three main fields: id, conversations, and image.
The id serves as a unique identifier for each data sample.
The image field is a list of strings specifying image paths, all given as paths relative to the root data directory.
The conversations field is a list of dialogue turns, where each turn is a dictionary with two key-value pairs: from, indicating the speaker identity (e.g., human or gpt), and value, indicating the textual content. Within value, the <image> placeholder marks where images are inserted, and the number of <image> placeholders matches the number of images listed in the image field.
{
"id": 0,
"conversations": [
{"from": "human", "value": "<image>\nuser input <image>\nuser input"},
{"from": "gpt", "value": "assistant output"},
{"from": "human", "value": "<image>\nuser input"},
{"from": "gpt", "value": "assistant output"}
],
"image": ["path/to/image1.jpg", "path/to/image2.jpg", "path/to/image3.jpg"]
}
The image data is packaged into 53 independent zip files (images_part_001.zip through images_part_053.zip; ~21 GB each except the last (~7 GB); total ~1.1 TB). Each zip can be extracted on its own β they are not split volumes, so you don't need all parts to extract any one of them. Every zip preserves the full images/ directory structure and extracting them all to the same destination reconstructs the complete image tree.
Two one-click extraction scripts are included in the repo root:
Linux / macOS / Git Bash:
bash extract_all.sh # extract to the script's parent directory
bash extract_all.sh /path/to/dir # extract to a specified directory
Windows PowerShell:
.\extract_all.ps1 # extract to the script's parent directory
.\extract_all.ps1 -Dest D:\data # extract to a specified directory
You can also extract manually with any zip tool (e.g. unzip, 7-Zip, WinRAR) β each zip is a standard archive.
After training, you can use EASI to evaluate your model on mainstream spatial intelligence benchmarks.
EASI supports over 20 spatial intelligence models and more than 10 spatial benchmarks, offering Docker for one-click spatial intelligence evaluation.
@InProceedings{sensenova-si,
title = {Scaling Spatial Intelligence with Multimodal Foundation Models},
author = {Cai, Zhongang and Wang, Ruisi and Gu, Chenyang and Pu, Fanyi and Xu, Junxiang and Wang, Yubo and Yin, Wanqi and Yang, Zhitao and Wei, Chen and Sun, Qingping and Zhou, Tongxi and Li, Jiaqi and Pang, Hui En and Qian, Oscar and Wei, Yukun and Lin, Zhiqian and Shi, Xuanke and Deng, Kewang and Han, Xiaoyang and Chen, Zukai and Fan, Xiangyu and Deng, Hanming and Lu, Lewei and Pan, Liang and Li, Bo and Liu, Ziwei and Wang, Quan and Lin, Dahua and Yang, Lei},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
EN | δΈζ
π This is the official full-scale training dataset of the SenseNova-SI series. SenseNova-SI-8M contains ~8.16 million carefully curated training samples spanning ~2.72 million unique images, organized under a rigorous taxonomy of spatial capabilities. It is the dataset used to train the recommended released model SenseNova-SI-1.1-InternVL3-8B and serves as the canonical training corpus for spatial intelligence research with SenseNova-SI.
Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation models to cultivate spatial intelligence within the SenseNova-SI family, built upon established multimodal foundations including visual understanding models (i.e., Qwen3-VL and InternVL3) and unified understanding and generation models (i.e., Bagel). We take a principled approach to constructing high-performing and robust spatial intelligence by systematically curating SenseNova-SI-8M: eight million diverse data samples under a rigorous taxonomy of spatial capabilities. SenseNova-SI demonstrates unprecedented performance across a broad range of spatial intelligence benchmarks, while maintaining strong general multimodal understanding. More importantly, we analyze the impact of data scaling, discuss early signs of emergent generalization capabilities enabled by diverse data training, analyze the risk of overfitting and language shortcuts, present a preliminary study on spatial chain-of-thought reasoning, and validate the potential downstream application. SenseNova-SI is an ongoing project, and this report will be updated continuously. All newly trained multimodal foundation models are publicly released to facilitate further research in this direction. In the future, SenseNova-SI will be integrated with larger-scale in-house models.
SenseNova-SI-8M is the official, full-scale training dataset of the SenseNova-SI series. The recommended model SenseNova-SI-1.1-InternVL3-8B is trained on this dataset and is our primary recommended release for spatial intelligence research and downstream use.
The previously released SenseNova-SI-800K subset is provided only as a reference for scaling-law analysis; the model trained on it is not the recommended release.
| Model | SI Dataset | Avg | VSI | MMSI | MindCube-Tiny | ViewSpatial | SITE | BLINK | 3DSRBench | EmbSpatial |
|---|---|---|---|---|---|---|---|---|---|---|
| InternVL3-8B | - | 45.7 | 42.1 | 28.0 | 41.5 | 38.7 | 41.1 | 53.5 | 44.3 | 76.3 |
| VST-7B-SFT | VST-P-4.1M | 50.8 | 55.5 | 32.5 | 39.7 | 50.5 | 39.7 | 61.9 | 53.1 | 73.7 |
| Cambrian-S-7B | VSI-590K | 45.1 | 62.9 | 27.1 | 37.9 | 41.3 | 36.1 | 37.9 | 45.0 | 72.8 |
| *SenseNova-SI-1.1-InternVL3-8B-800K | SenseNova-SI-800K (scaling-law reference) | - | 60.9 | 36.4 | 56.9 | 52.5 | 47.7 | - | - | - |
| SenseNova-SI-1.1-InternVL3-8B (recommended) | SenseNova-SI-8M (this dataset) | 61.5 | 68.8 | 43.3 | 85.7 | 54.7 | 47.7 | 63.9 | 55.5 | 72.0 |
| Aspect | SenseNova-SI-800K | SenseNova-SI-8M (this release) |
|---|---|---|
| Role | Scaling-law reference subset | Official full training set |
| # Training samples | ~832K | ~8.16M (β 800K Γ 10) |
| # Unique images | ~410K | ~2.72M |
| Total image storage | ~370 GB | ~1.1 TB |
| Image archive | 93 Γ 4 GB independent zips | 53 independent zips (~21 GB each, last ~7 GB) |
| Annotation file | SenseNova-SI-800K.jsonl | SenseNova-SI-8M.jsonl |
| Viewer preview | 1,000-sample parquet | SenseNova-SI-8M_1000samples.parquet |
| Recommended trained model | (reference only) | SenseNova-SI-1.1-InternVL3-8B |
All zip archives are independent (not split volumes) β each can be extracted on its own and together they reconstruct the full images/ tree. One-click extraction scripts are included for both Linux/macOS and Windows (see Download & Extract Images).
The data is stored in SenseNova-SI-8M.jsonl using the JSONL (JSON Lines) format, where each line represents an independent data entry. Each entry is a dictionary organized in the following format, containing three main fields: id, conversations, and image.
The id serves as a unique identifier for each data sample.
The image field is a list of strings specifying image paths, all given as paths relative to the root data directory.
The conversations field is a list of dialogue turns, where each turn is a dictionary with two key-value pairs: from, indicating the speaker identity (e.g., human or gpt), and value, indicating the textual content. Within value, the <image> placeholder marks where images are inserted, and the number of <image> placeholders matches the number of images listed in the image field.
{
"id": 0,
"conversations": [
{"from": "human", "value": "<image>\nuser input <image>\nuser input"},
{"from": "gpt", "value": "assistant output"},
{"from": "human", "value": "<image>\nuser input"},
{"from": "gpt", "value": "assistant output"}
],
"image": ["path/to/image1.jpg", "path/to/image2.jpg", "path/to/image3.jpg"]
}
The image data is packaged into 53 independent zip files (images_part_001.zip through images_part_053.zip; ~21 GB each except the last (~7 GB); total ~1.1 TB). Each zip can be extracted on its own β they are not split volumes, so you don't need all parts to extract any one of them. Every zip preserves the full images/ directory structure and extracting them all to the same destination reconstructs the complete image tree.
Two one-click extraction scripts are included in the repo root:
Linux / macOS / Git Bash:
bash extract_all.sh # extract to the script's parent directory
bash extract_all.sh /path/to/dir # extract to a specified directory
Windows PowerShell:
.\extract_all.ps1 # extract to the script's parent directory
.\extract_all.ps1 -Dest D:\data # extract to a specified directory
You can also extract manually with any zip tool (e.g. unzip, 7-Zip, WinRAR) β each zip is a standard archive.
After training, you can use EASI to evaluate your model on mainstream spatial intelligence benchmarks.
EASI supports over 20 spatial intelligence models and more than 10 spatial benchmarks, offering Docker for one-click spatial intelligence evaluation.
@InProceedings{sensenova-si,
title = {Scaling Spatial Intelligence with Multimodal Foundation Models},
author = {Cai, Zhongang and Wang, Ruisi and Gu, Chenyang and Pu, Fanyi and Xu, Junxiang and Wang, Yubo and Yin, Wanqi and Yang, Zhitao and Wei, Chen and Sun, Qingping and Zhou, Tongxi and Li, Jiaqi and Pang, Hui En and Qian, Oscar and Wei, Yukun and Lin, Zhiqian and Shi, Xuanke and Deng, Kewang and Han, Xiaoyang and Chen, Zukai and Fan, Xiangyu and Deng, Hanming and Lu, Lewei and Pan, Liang and Li, Bo and Liu, Ziwei and Wang, Quan and Lin, Dahua and Yang, Lei},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}