TIGER-Lab/Mantis-Eval

Dataset

6

stars

64

commits

2

linked in READMEs

Nov 15, 2024

updated

README

Overview

This is a newly curated dataset to evaluate multimodal language models' capability to reason over multiple images. More details are shown in https://tiger-ai-lab.github.io/Mantis/.

Statistics

This evaluation dataset contains 217 human-annotated challenging multi-image reasoning problems.

Leaderboard

We list the current results as follows:

ModelsSizeMantis-Eval
LLaVA OneVision72B77.60
LLaVA OneVision7B64.20
GPT-4V-62.67
Mantis-SigLIP8B59.45
Mantis-Idefics28B57.14
Mantis-CLIP8B55.76
VILA8B51.15
BLIP-213B49.77
Idefics28B48.85
InstructBLIP13B45.62
LLaVA-V1.67B45.62
CogVLM17B45.16
LLaVA OneVision0.5B39.60
Qwen-VL-Chat7B39.17
Emu2-Chat37B37.79
VideoLLaVA7B35.04
Mantis-Flamingo9B32.72
LLaVA-v1.57B31.34
Kosmos21.6B30.41
Idefics19B28.11
Fuyu8B27.19
OpenFlamingo9B12.44
Otter-Image9B14.29

Citation

If you are using this dataset, please cite our work with

@article{Jiang2024MANTISIM,
  title={MANTIS: Interleaved Multi-Image Instruction Tuning},
  author={Dongfu Jiang and Xuan He and Huaye Zeng and Cong Wei and Max W.F. Ku and Qian Liu and Wenhu Chen},
  journal={Transactions on Machine Learning Research},
  year={2024},
  volume={2024},
  url={https://openreview.net/forum?id=skLtdUVaJa}
}

Contributors

DJ
Dongfu Jiang

33 commits

DongfuJiang

21 commits

wenhu

10 commits

TIGER-Lab/Mantis-Eval

Dataset

6

stars

64

commits

2

linked in READMEs

Nov 15, 2024

updated

README

Overview

This is a newly curated dataset to evaluate multimodal language models' capability to reason over multiple images. More details are shown in https://tiger-ai-lab.github.io/Mantis/.

Statistics

This evaluation dataset contains 217 human-annotated challenging multi-image reasoning problems.

Leaderboard

We list the current results as follows:

ModelsSizeMantis-Eval
LLaVA OneVision72B77.60
LLaVA OneVision7B64.20
GPT-4V-62.67
Mantis-SigLIP8B59.45
Mantis-Idefics28B57.14
Mantis-CLIP8B55.76
VILA8B51.15
BLIP-213B49.77
Idefics28B48.85
InstructBLIP13B45.62
LLaVA-V1.67B45.62
CogVLM17B45.16
LLaVA OneVision0.5B39.60
Qwen-VL-Chat7B39.17
Emu2-Chat37B37.79
VideoLLaVA7B35.04
Mantis-Flamingo9B32.72
LLaVA-v1.57B31.34
Kosmos21.6B30.41
Idefics19B28.11
Fuyu8B27.19
OpenFlamingo9B12.44
Otter-Image9B14.29

Citation

If you are using this dataset, please cite our work with

@article{Jiang2024MANTISIM,
  title={MANTIS: Interleaved Multi-Image Instruction Tuning},
  author={Dongfu Jiang and Xuan He and Huaye Zeng and Cong Wei and Max W.F. Ku and Qian Liu and Wenhu Chen},
  journal={Transactions on Machine Learning Research},
  year={2024},
  volume={2024},
  url={https://openreview.net/forum?id=skLtdUVaJa}
}

Contributors

DJ
Dongfu Jiang

33 commits

DongfuJiang

21 commits

wenhu

10 commits