Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
See the codeUsing VLMEvalKit backend (default):
python scripts/submissions/run_easi_eval.py \
--model sensenova/SenseNova-SI-1.5-InternVL3-8B \
--nproc 4
Using lmms-eval backend:
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model internvl2 \
--model-args "pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B" \
--nproc 4
Under the hood, EASI wraps VLMEvalKit and lmms-eval with a unified CLI. See the respective repos for advanced usage and adding custom models.
EASI is a unified evaluation suite for Spatial Intelligence. It benchmarks state-of-the-art proprietary and open-source multimodal LLMs across a growing set of spatial benchmarks.
Full details are available at 👉 Supported Models & Benchmarks. EASI also provides transparent 👉 Benchmark Verification against official scores.
🌟 [2026-04-27] EASI v0.2.2 is released. Major updates include:
For the full release history and detailed changelog, please see 👉 Changelog.
The setup script installs both evaluation backends (VLMEvalKit and lmms-eval) with pinned dependencies:
git clone --recursive https://github.com/EvolvingLMMs-Lab/EASI.git
cd EASI
bash scripts/setup.sh
source .venv/bin/activate
This creates a Python 3.11 virtual environment with both backends, flash-attn, and all required dependencies. See scripts/setup.sh for details.
bash dockerfiles/EASI/build_runtime_docker.sh
docker run --gpus all -it --rm \
-v /path/to/your/data:/mnt/data \
--name easi-runtime \
VLMEvalKit_EASI:latest \
/bin/bash
EASI provides a unified evaluation script that supports both VLMEvalKit and lmms-eval backends. The script handles dataset preparation, evaluation, result collection, and optional leaderboard submission.
VLMEvalKit backend (default):
# Run EASI-8 core benchmarks on 4 GPUs
python scripts/submissions/run_easi_eval.py \
--model sensenova/SenseNova-SI-1.5-InternVL3-8B \
--nproc 4
lmms-eval backend:
# Run EASI-8 core benchmarks on 4 GPUs
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model internvl2 \
--model-args "pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B" \
--nproc 4
With automated submission:
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model internvl2 \
--model-args "pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B" \
--nproc 4 \
--submit \
--submission-configs '{
"modelName": "sensenova/SenseNova-SI-1.5-InternVL3-8B",
"modelType": "instruction",
"precision": "bfloat16"
}'
More options:
# Run specific benchmarks only
python scripts/submissions/run_easi_eval.py \
--model Qwen/Qwen2.5-VL-7B-Instruct \
--benchmarks vsi_bench,blink,sitebench
# Include extra benchmarks (MMSI-Video, OmniSpatial, SPAR-Bench, VSI-Debiased)
python scripts/submissions/run_easi_eval.py \
--model Qwen/Qwen2.5-VL-7B-Instruct \
--nproc 8 --include-extra
# Force re-evaluation (ignore previous results)
python scripts/submissions/run_easi_eval.py \
--model Qwen/Qwen2.5-VL-7B-Instruct \
--nproc 8 --rerun
# lmms-eval OOM recovery: complete failed benchmarks in single-GPU mode
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model qwen3_vl \
--model-args "pretrained=Qwen/Qwen3-VL-8B-Instruct,attn_implementation=flash_attention_2" \
--no-accelerate
Full CLI options and submission config details at 👉 Submission Guide.
For advanced usage or custom model integration, you can also call the backends directly:
VLMEvalKit:
cd VLMEvalKit/
python run.py --data MindCubeBench_tiny_raw_qa \
--model SenseNova-SI-1.5-InternVL3-8B \
--verbose --reuse --judge extract_matching
lmms-eval:
CUDA_VISIBLE_DEVICES=0,1,2,3 accelerate launch \
--num_processes=4 -m lmms_eval \
--model internvl2 \
--model_args=pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B \
--tasks vsibench_multiimage \
--batch_size 1 --log_samples --output_path ./logs/
For more details, refer to the VLMEvalKit documentation and lmms-eval documentation.
vlmeval/config.py. Verify inference with vlmutil check {MODEL_NAME}.qwen2_5_vl, llava, internvl2, etc.). See the lmms-eval models directory.You can submit your evaluation results at 👉 EASI Leaderboard Submission.
Full details and file format examples are available at 👉 Submission Guide.
EASI is an open and evolving evaluation suite. We warmly welcome community contributions, including:
If you are interested in contributing, or have questions about integration, please contact us at 📧 easi-lmms-lab@outlook.com
@article{easi2025,
title={Holistic Evaluation of Multimodal LLMs on Spatial Intelligence},
author={Cai, Zhongang and Wang, Yubo and Sun, Qingping and Wang, Ruisi and Gu, Chenyang and Yin, Wanqi and Lin, Zhiqian and Yang, Zhitao and Wei, Chen and Shi, Xuanke and Deng, Kewang and Han, Xiaoyang and Chen, Zukai and Li, Jiaqi and Fan, Xiangyu and Deng, Hanming and Lu, Lewei and Li, Bo and Liu, Ziwei and Wang, Quan and Lin, Dahua and Yang, Lei},
journal={arXiv preprint arXiv:2508.13142},
year={2025}
}
Python
95.6%
Dockerfile
2.2%
Shell
2.2%
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
See the codeUsing VLMEvalKit backend (default):
python scripts/submissions/run_easi_eval.py \
--model sensenova/SenseNova-SI-1.5-InternVL3-8B \
--nproc 4
Using lmms-eval backend:
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model internvl2 \
--model-args "pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B" \
--nproc 4
Under the hood, EASI wraps VLMEvalKit and lmms-eval with a unified CLI. See the respective repos for advanced usage and adding custom models.
EASI is a unified evaluation suite for Spatial Intelligence. It benchmarks state-of-the-art proprietary and open-source multimodal LLMs across a growing set of spatial benchmarks.
Full details are available at 👉 Supported Models & Benchmarks. EASI also provides transparent 👉 Benchmark Verification against official scores.
🌟 [2026-04-27] EASI v0.2.2 is released. Major updates include:
For the full release history and detailed changelog, please see 👉 Changelog.
The setup script installs both evaluation backends (VLMEvalKit and lmms-eval) with pinned dependencies:
git clone --recursive https://github.com/EvolvingLMMs-Lab/EASI.git
cd EASI
bash scripts/setup.sh
source .venv/bin/activate
This creates a Python 3.11 virtual environment with both backends, flash-attn, and all required dependencies. See scripts/setup.sh for details.
bash dockerfiles/EASI/build_runtime_docker.sh
docker run --gpus all -it --rm \
-v /path/to/your/data:/mnt/data \
--name easi-runtime \
VLMEvalKit_EASI:latest \
/bin/bash
EASI provides a unified evaluation script that supports both VLMEvalKit and lmms-eval backends. The script handles dataset preparation, evaluation, result collection, and optional leaderboard submission.
VLMEvalKit backend (default):
# Run EASI-8 core benchmarks on 4 GPUs
python scripts/submissions/run_easi_eval.py \
--model sensenova/SenseNova-SI-1.5-InternVL3-8B \
--nproc 4
lmms-eval backend:
# Run EASI-8 core benchmarks on 4 GPUs
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model internvl2 \
--model-args "pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B" \
--nproc 4
With automated submission:
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model internvl2 \
--model-args "pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B" \
--nproc 4 \
--submit \
--submission-configs '{
"modelName": "sensenova/SenseNova-SI-1.5-InternVL3-8B",
"modelType": "instruction",
"precision": "bfloat16"
}'
More options:
# Run specific benchmarks only
python scripts/submissions/run_easi_eval.py \
--model Qwen/Qwen2.5-VL-7B-Instruct \
--benchmarks vsi_bench,blink,sitebench
# Include extra benchmarks (MMSI-Video, OmniSpatial, SPAR-Bench, VSI-Debiased)
python scripts/submissions/run_easi_eval.py \
--model Qwen/Qwen2.5-VL-7B-Instruct \
--nproc 8 --include-extra
# Force re-evaluation (ignore previous results)
python scripts/submissions/run_easi_eval.py \
--model Qwen/Qwen2.5-VL-7B-Instruct \
--nproc 8 --rerun
# lmms-eval OOM recovery: complete failed benchmarks in single-GPU mode
python scripts/submissions/run_easi_eval.py \
--backend lmms-eval \
--model qwen3_vl \
--model-args "pretrained=Qwen/Qwen3-VL-8B-Instruct,attn_implementation=flash_attention_2" \
--no-accelerate
Full CLI options and submission config details at 👉 Submission Guide.
For advanced usage or custom model integration, you can also call the backends directly:
VLMEvalKit:
cd VLMEvalKit/
python run.py --data MindCubeBench_tiny_raw_qa \
--model SenseNova-SI-1.5-InternVL3-8B \
--verbose --reuse --judge extract_matching
lmms-eval:
CUDA_VISIBLE_DEVICES=0,1,2,3 accelerate launch \
--num_processes=4 -m lmms_eval \
--model internvl2 \
--model_args=pretrained=sensenova/SenseNova-SI-1.5-InternVL3-8B \
--tasks vsibench_multiimage \
--batch_size 1 --log_samples --output_path ./logs/
For more details, refer to the VLMEvalKit documentation and lmms-eval documentation.
vlmeval/config.py. Verify inference with vlmutil check {MODEL_NAME}.qwen2_5_vl, llava, internvl2, etc.). See the lmms-eval models directory.You can submit your evaluation results at 👉 EASI Leaderboard Submission.
Full details and file format examples are available at 👉 Submission Guide.
EASI is an open and evolving evaluation suite. We warmly welcome community contributions, including:
If you are interested in contributing, or have questions about integration, please contact us at 📧 easi-lmms-lab@outlook.com
@article{easi2025,
title={Holistic Evaluation of Multimodal LLMs on Spatial Intelligence},
author={Cai, Zhongang and Wang, Yubo and Sun, Qingping and Wang, Ruisi and Gu, Chenyang and Yin, Wanqi and Lin, Zhiqian and Yang, Zhitao and Wei, Chen and Shi, Xuanke and Deng, Kewang and Han, Xiaoyang and Chen, Zukai and Li, Jiaqi and Fan, Xiangyu and Deng, Hanming and Lu, Lewei and Li, Bo and Liu, Ziwei and Wang, Quan and Lin, Dahua and Yang, Lei},
journal={arXiv preprint arXiv:2508.13142},
year={2025}
}
Python
95.6%
Dockerfile
2.2%
Shell
2.2%