We present MMMG, a comprehensive and human-aligned benchmark for multimodal generation across 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio), with a focus on tasks that present significant challenges for generation models, while still enabling reliable automatic evaluation.
This huggingface page only contains the raw dataset of MMMG, for full evaluation suite, please refer to our github page: https://github.com/yaojh18/MMMG.
Please refer to our paper for detailed information: https://arxiv.org/abs/2505.17613v1.
The leaderboard is avaliable at: https://yaojh18.github.io/mmmg-leaderboard/
18 commits
We present MMMG, a comprehensive and human-aligned benchmark for multimodal generation across 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio), with a focus on tasks that present significant challenges for generation models, while still enabling reliable automatic evaluation.
This huggingface page only contains the raw dataset of MMMG, for full evaluation suite, please refer to our github page: https://github.com/yaojh18/MMMG.
Please refer to our paper for detailed information: https://arxiv.org/abs/2505.17613v1.
The leaderboard is avaliable at: https://yaojh18.github.io/mmmg-leaderboard/
18 commits