Inst-IT/Inst-It-Bench

Dataset

Inst-It Bench

1

12 commits

6 linked in READMEs

updated Mar 3, 2025

See the code

README

Inst-It Bench

Homepage | Code | Paper | arXiv

Inst-It Bench is a fine-grained multimodal benchmark for evaluating LMMs at the instance-Level, which is introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning.

  • Size: 1,000 image QAs and 1,000 video QAs
  • Splits: Image split and Video split
  • Evaluation Formats: Open-Ended and Multiple-Choice

Introduction

Existing multimodal benchmarks primarily focus on global understanding, failing to provide more in-depth insights into the instance-level comprehension capability of models. Specifically, Inst-IT Bench includes two parts: image-split and video-split, and is able to evaluate the models' ability in understanding instances in both images and videos. The image-split contains 1,036 QA pairs for 338 images, while the video-split contains 1,001 QA pairs for 206 videos. Each QA pair is available in both open-ended and multiple-choices formats. The followings are some examples from the video-split:


Click here to unfold more data examples:



Evaluate your model on Inst-IT Bench

If you want to evaluate your model on our Inst-IT Bench, please refer to our GitHub code for more instructions.

We conducted an extensive evaluation of Inst-IT Bench

We conduct extensive evaluations on our benchmark, including state-of-the-art open-source image models, video models, and cutting-edge proprietary models. The results that even state-of-the-art models struggle with fine-grained, instance-level understanding.

#IT indicates the number of training samples used during the instruction-tuning stage. N/A indicates that the number is unknown.

ModelLLM#ITOpen-Ended Q&AMulti-Choice Q&AOpen-Ended Q&AMulti-Choice Q&A
Random Guess-N/A-25.0-25.0
GPT-4o-N/A74.184.865.581.0
Gemini-1.5-pro-N/A69.979.761.476.7
Gemini-1.5-flash-N/A65.379.557.975.8
LLaVA-1.5Vicuna-7B665K41.632.1--
ViP-LLaVAVicuna-7B~1.2M42.129.2--
SoM-LLaVAVicuna-7B695K45.140.0--
LLaVA-NextVicuna-7B765K46.042.4--
LLaVA-NeXT-VideoVicuna-7B860K46.539.525.824.8
ShareGPT4VideoLlama3-8B~1.0M43.248.727.816.1
MiniCPM-V 2.6Qwen2-7B~7.0M57.666.840.045.2
LLaVA-OV (SI)Qwen2-7B~7.2M60.361.831.436.4
LLaVA-OVQwen2-7B~8.8M48.071.733.245.6
LLaVA-VideoQwen2-7B~7.4M45.167.034.153.2
InternVL2InternLM2.5-7BN/A58.666.539.845.5
Qwen2-VL-InstructQwen2-7BN/A48.364.938.259.4
Qwen2-VL-InstructQwen2-72BN/A55.574.745.574.6
LLaVA-Next-Inst-ITVicuna-7B920K68.663.049.342.1
LLaVA-Next-Inst-ITQwen2-7B920K67.975.345.753.3

Contact

Feel free to contact us if you have any questions or suggestions

Citation

If you find our work helpful, please consider citing our paper ✒️ and like our dataset ❤️ :

  @article{peng2024inst,
    title={Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning},
    author={Peng, Wujian and Meng, Lingchen and Chen, Yitong and Xie, Yiweng and Liu, Yang and Gui, Tao and Xu, Hang and Qiu, Xipeng and Wu, Zuxuan and Jiang, Yu-Gang},
    journal={arXiv preprint arXiv:2412.03565},
    year={2024}
  }
image
multimodal-instance-understanding
video

Contributors

wjpoom

12 commits

Inst-IT/Inst-It-Bench

Dataset

Inst-It Bench

1

12 commits

6 linked in READMEs

updated Mar 3, 2025

See the code

README

Inst-It Bench

Homepage | Code | Paper | arXiv

Inst-It Bench is a fine-grained multimodal benchmark for evaluating LMMs at the instance-Level, which is introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning.

  • Size: 1,000 image QAs and 1,000 video QAs
  • Splits: Image split and Video split
  • Evaluation Formats: Open-Ended and Multiple-Choice

Introduction

Existing multimodal benchmarks primarily focus on global understanding, failing to provide more in-depth insights into the instance-level comprehension capability of models. Specifically, Inst-IT Bench includes two parts: image-split and video-split, and is able to evaluate the models' ability in understanding instances in both images and videos. The image-split contains 1,036 QA pairs for 338 images, while the video-split contains 1,001 QA pairs for 206 videos. Each QA pair is available in both open-ended and multiple-choices formats. The followings are some examples from the video-split:


Click here to unfold more data examples:



Evaluate your model on Inst-IT Bench

If you want to evaluate your model on our Inst-IT Bench, please refer to our GitHub code for more instructions.

We conducted an extensive evaluation of Inst-IT Bench

We conduct extensive evaluations on our benchmark, including state-of-the-art open-source image models, video models, and cutting-edge proprietary models. The results that even state-of-the-art models struggle with fine-grained, instance-level understanding.

#IT indicates the number of training samples used during the instruction-tuning stage. N/A indicates that the number is unknown.

ModelLLM#ITOpen-Ended Q&AMulti-Choice Q&AOpen-Ended Q&AMulti-Choice Q&A
Random Guess-N/A-25.0-25.0
GPT-4o-N/A74.184.865.581.0
Gemini-1.5-pro-N/A69.979.761.476.7
Gemini-1.5-flash-N/A65.379.557.975.8
LLaVA-1.5Vicuna-7B665K41.632.1--
ViP-LLaVAVicuna-7B~1.2M42.129.2--
SoM-LLaVAVicuna-7B695K45.140.0--
LLaVA-NextVicuna-7B765K46.042.4--
LLaVA-NeXT-VideoVicuna-7B860K46.539.525.824.8
ShareGPT4VideoLlama3-8B~1.0M43.248.727.816.1
MiniCPM-V 2.6Qwen2-7B~7.0M57.666.840.045.2
LLaVA-OV (SI)Qwen2-7B~7.2M60.361.831.436.4
LLaVA-OVQwen2-7B~8.8M48.071.733.245.6
LLaVA-VideoQwen2-7B~7.4M45.167.034.153.2
InternVL2InternLM2.5-7BN/A58.666.539.845.5
Qwen2-VL-InstructQwen2-7BN/A48.364.938.259.4
Qwen2-VL-InstructQwen2-72BN/A55.574.745.574.6
LLaVA-Next-Inst-ITVicuna-7B920K68.663.049.342.1
LLaVA-Next-Inst-ITQwen2-7B920K67.975.345.753.3

Contact

Feel free to contact us if you have any questions or suggestions

Citation

If you find our work helpful, please consider citing our paper ✒️ and like our dataset ❤️ :

  @article{peng2024inst,
    title={Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning},
    author={Peng, Wujian and Meng, Lingchen and Chen, Yitong and Xie, Yiweng and Liu, Yang and Gui, Tao and Xu, Hang and Qiu, Xipeng and Wu, Zuxuan and Jiang, Yu-Gang},
    journal={arXiv preprint arXiv:2412.03565},
    year={2024}
  }
image
multimodal-instance-understanding
video

Contributors

wjpoom

12 commits