MM-MT-Bench is a multi-turn LLM-as-a-judge evaluation benchmark similar to the text MT-Bench for testing multimodal instruction-tuned models. While existing benchmarks like MMMU, MathVista, ChartQA and so on are focused on closed-ended questions with short responses, they do not evaluate model's ability to follow user instructions in multi-turn dialogues and answer open-ended questions in a zero-shot manner. MM MT-Bench is designed to overcome this limitation. The mistral-evals repository provides an example script on how to evaluate a model on this benchmark.
8 commits
MM-MT-Bench is a multi-turn LLM-as-a-judge evaluation benchmark similar to the text MT-Bench for testing multimodal instruction-tuned models. While existing benchmarks like MMMU, MathVista, ChartQA and so on are focused on closed-ended questions with short responses, they do not evaluate model's ability to follow user instructions in multi-turn dialogues and answer open-ended questions in a zero-shot manner. MM MT-Bench is designed to overcome this limitation. The mistral-evals repository provides an example script on how to evaluate a model on this benchmark.
8 commits