T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.
This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical consistency).
| Category | Models |
|---|---|
| Diffusion Models | SD-3-Medium, SD-3.5-Medium, SD-3.5-Large, FLUX.1-schnell, FLUX.1-dev, FLUX.1-Krea-dev, FLUX.2-dev, FLUX.2-klein-4B, FLUX.2-klein-9B, PixArt-$\alpha$, PixArt-$\Sigma$, HiDream-I1, Qwen-Image, Qwen-Image-2512, HunyuanImage-3.0, Z-Image-Turbo, Z-Image, LongCat-Image |
| Autogressive Models | Infinity-8B and GoT-R1-7B |
| Unified Models | BAGEL, BAGEL w/ Think, show-o2-1.5B, show-o2-7B, Janus-Pro-1B, Janus-Pro-7B, BLIP3o-4B, BLIP3o-8B, OmniGen2-7B |
| Closed-Source Models | Seedream 3.0, Seedream 4.0, Seedream 4.5, Gemini 2.0 Flash, Nano Banana, Nano Banana Pro, Nano Banana 2, Imagen 4, Imagen 4 Ultra, GPT-Image (GPT-4o), GPT-Image-1.5 |
If you find this dataset useful, please cite our paper:
@inproceedings{
li2026easier,
title={Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?},
author={Ouxiang Li and Yuan Wang and Xinting Hu and Huijuan Huang and Rui Chen and Jiarong Ou and Xin Tao and Pengfei Wan and Xiaojuan Qi and Fuli Feng},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=iqAFhWistW}
}
T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.
This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical consistency).
| Category | Models |
|---|---|
| Diffusion Models | SD-3-Medium, SD-3.5-Medium, SD-3.5-Large, FLUX.1-schnell, FLUX.1-dev, FLUX.1-Krea-dev, FLUX.2-dev, FLUX.2-klein-4B, FLUX.2-klein-9B, PixArt-$\alpha$, PixArt-$\Sigma$, HiDream-I1, Qwen-Image, Qwen-Image-2512, HunyuanImage-3.0, Z-Image-Turbo, Z-Image, LongCat-Image |
| Autogressive Models | Infinity-8B and GoT-R1-7B |
| Unified Models | BAGEL, BAGEL w/ Think, show-o2-1.5B, show-o2-7B, Janus-Pro-1B, Janus-Pro-7B, BLIP3o-4B, BLIP3o-8B, OmniGen2-7B |
| Closed-Source Models | Seedream 3.0, Seedream 4.0, Seedream 4.5, Gemini 2.0 Flash, Nano Banana, Nano Banana Pro, Nano Banana 2, Imagen 4, Imagen 4 Ultra, GPT-Image (GPT-4o), GPT-Image-1.5 |
If you find this dataset useful, please cite our paper:
@inproceedings{
li2026easier,
title={Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?},
author={Ouxiang Li and Yuan Wang and Xinting Hu and Huijuan Huang and Rui Chen and Jiarong Ou and Xin Tao and Pengfei Wan and Xiaojuan Qi and Fuli Feng},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=iqAFhWistW}
}