Arxiv Link, Project Page, GitHub Page
An example of test looks as follows:
{'file_name': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=2560x3747>,
'ID': '031_31_01_001',
'Language': 'Italian',
'Category': 'Lifestyle',
'Question_Type': 'Short Questions',
'English_Question': 'What type of clothing are the people in the image wearing?',
'English_Answer': 'The people in the image are wearing professional clothing.',
'Translated_Question': " Che tipo di abbigliamento indossano le persone nell'immagine?",
'Translated_Answer': " Le persone nell'immagine indossano abiti professionali.",
'Image_Url': 'https://assets.vogue.com/photos/650c97c9e5c5af360f4668ac/master/w_2560%2Cc_limit/GettyImages-1499571723.jpg'
}
Data Fields
The data fields are:
- 'file_name': ,
- 'ID': A unique ID in the language#_cat#_img# format.
- 'Language': A language from the 100 languages.
- 'Category': A category from our total 19 categories.
- 'Question_Type': One of four question types, MCQs, T/F, SVQAs, and LVQAs.
- 'English_Question': The original question in the English Language.
- 'English_Answer': The original answer in the English Language.
- 'Translated_Question': The translated and annotated question in the Native language.
- 'Translated_Answer': The translated and annotated answer in the Native language.
- 'Image_Url': The image URL that we have retrieved from the internet.
Data statistics of our ALM-bench showing the diversity of the scripts, global coverage, comprehensive categories, and various question types. Our dataset contains 22.7K high-quality
question-answers in total, covering 100 languages and 24 scripts. All the samples are manually verified by native speakers.

Comparison of various LMM benchmarks with a focus on multilingual and cultural understanding. The Domains indicate the range of aspects covered by the dataset for each language. Question Form is categorized as "Diverse" if the questions phrasing varies, and "Fixed" otherwise. Annotation Types are classified as "Manual" if questions were originally in the local language, "Manual+Auto" if questions were generated or translated using GPT-4/Google API and subsequently validated by human experts, and "Auto" if generated or translated automatically without human validation. Bias Correction reflects whether the dataset is balanced across cultures and countries, while Diversity indicates whether the dataset includes both Western and non-Western minority cultures. ‘-’ means information not available.
BibTeX:
@misc{vayani2024alm,
title={All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages},
author={Ashmal Vayani and Dinura Dissanayake and Hasindri Watawana and Noor Ahsan and Nevasini Sasikumar and Omkar Thawakar and Henok Biadglign Ademtew and Yahya Hmaiti and Amandeep Kumar and Kartik Kuckreja and Mykola Maslych and Wafa Al Ghallabi and Mihail Mihaylov and Chao Qin and Abdelrahman M Shaker and Mike Zhang and Mahardika Krisna Ihsani and Amiel Esplana and Monil Gokani and Shachar Mirkin and Harsh Singh and Ashay Srivastava and Endre Hamerlik and Fathinah Asma Izzati and Fadillah Adamsyah Maani and Sebastian Cavada and Jenny Chim and Rohit Gupta and Sanjay Manjunath and Kamila Zhumakhanova and Feno Heriniaina Rabevohitra and Azril Amirudin and Muhammad Ridzuan and Daniya Kareem and Ketan More and Kunyang Li and Pramesh Shakya and Muhammad Saad and Amirpouya Ghasemaghaei and Amirbek Djanibekov and Dilshod Azizov and Branislava Jankovic and Naman Bhatia and Alvaro Cabrera and Johan Obando-Ceron and Olympiah Otieno and Fabian Farestam and Muztoba Rabbani and Sanoojan Baliah and Santosh Sanjeev and Abduragim Shtanchaev and Maheen Fatima and Thao Nguyen and Amrin Kareem and Toluwani Aremu and Nathan Xavier and Amit Bhatkal and Hawau Toyin and Aman Chadha and Hisham Cholakkal and Rao Muhammad Anwer and Michael Felsberg and Jorma Laaksonen and Thamar Solorio and Monojit Choudhury and Ivan Laptev and Mubarak Shah and Salman Khan and Fahad Khan},
year={2024},
eprint={2411.16508},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2411.16508},
}
We release our work under CC BY-NC 4.0 License. The CC BY-NC 4.0 license allows others to share, remix, and adapt the work, as long as it's for non-commercial purposes and proper attribution is given to the original creator.
Arxiv Link, Project Page, GitHub Page
An example of test looks as follows:
{'file_name': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=2560x3747>,
'ID': '031_31_01_001',
'Language': 'Italian',
'Category': 'Lifestyle',
'Question_Type': 'Short Questions',
'English_Question': 'What type of clothing are the people in the image wearing?',
'English_Answer': 'The people in the image are wearing professional clothing.',
'Translated_Question': " Che tipo di abbigliamento indossano le persone nell'immagine?",
'Translated_Answer': " Le persone nell'immagine indossano abiti professionali.",
'Image_Url': 'https://assets.vogue.com/photos/650c97c9e5c5af360f4668ac/master/w_2560%2Cc_limit/GettyImages-1499571723.jpg'
}
Data Fields
The data fields are:
- 'file_name': ,
- 'ID': A unique ID in the language#_cat#_img# format.
- 'Language': A language from the 100 languages.
- 'Category': A category from our total 19 categories.
- 'Question_Type': One of four question types, MCQs, T/F, SVQAs, and LVQAs.
- 'English_Question': The original question in the English Language.
- 'English_Answer': The original answer in the English Language.
- 'Translated_Question': The translated and annotated question in the Native language.
- 'Translated_Answer': The translated and annotated answer in the Native language.
- 'Image_Url': The image URL that we have retrieved from the internet.
Data statistics of our ALM-bench showing the diversity of the scripts, global coverage, comprehensive categories, and various question types. Our dataset contains 22.7K high-quality
question-answers in total, covering 100 languages and 24 scripts. All the samples are manually verified by native speakers.

Comparison of various LMM benchmarks with a focus on multilingual and cultural understanding. The Domains indicate the range of aspects covered by the dataset for each language. Question Form is categorized as "Diverse" if the questions phrasing varies, and "Fixed" otherwise. Annotation Types are classified as "Manual" if questions were originally in the local language, "Manual+Auto" if questions were generated or translated using GPT-4/Google API and subsequently validated by human experts, and "Auto" if generated or translated automatically without human validation. Bias Correction reflects whether the dataset is balanced across cultures and countries, while Diversity indicates whether the dataset includes both Western and non-Western minority cultures. ‘-’ means information not available.
BibTeX:
@misc{vayani2024alm,
title={All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages},
author={Ashmal Vayani and Dinura Dissanayake and Hasindri Watawana and Noor Ahsan and Nevasini Sasikumar and Omkar Thawakar and Henok Biadglign Ademtew and Yahya Hmaiti and Amandeep Kumar and Kartik Kuckreja and Mykola Maslych and Wafa Al Ghallabi and Mihail Mihaylov and Chao Qin and Abdelrahman M Shaker and Mike Zhang and Mahardika Krisna Ihsani and Amiel Esplana and Monil Gokani and Shachar Mirkin and Harsh Singh and Ashay Srivastava and Endre Hamerlik and Fathinah Asma Izzati and Fadillah Adamsyah Maani and Sebastian Cavada and Jenny Chim and Rohit Gupta and Sanjay Manjunath and Kamila Zhumakhanova and Feno Heriniaina Rabevohitra and Azril Amirudin and Muhammad Ridzuan and Daniya Kareem and Ketan More and Kunyang Li and Pramesh Shakya and Muhammad Saad and Amirpouya Ghasemaghaei and Amirbek Djanibekov and Dilshod Azizov and Branislava Jankovic and Naman Bhatia and Alvaro Cabrera and Johan Obando-Ceron and Olympiah Otieno and Fabian Farestam and Muztoba Rabbani and Sanoojan Baliah and Santosh Sanjeev and Abduragim Shtanchaev and Maheen Fatima and Thao Nguyen and Amrin Kareem and Toluwani Aremu and Nathan Xavier and Amit Bhatkal and Hawau Toyin and Aman Chadha and Hisham Cholakkal and Rao Muhammad Anwer and Michael Felsberg and Jorma Laaksonen and Thamar Solorio and Monojit Choudhury and Ivan Laptev and Mubarak Shah and Salman Khan and Fahad Khan},
year={2024},
eprint={2411.16508},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2411.16508},
}
We release our work under CC BY-NC 4.0 License. The CC BY-NC 4.0 license allows others to share, remix, and adapt the work, as long as it's for non-commercial purposes and proper attribution is given to the original creator.