lioooox/T2I-CoReBench-Images

Dataset

T2I-CoReBench-Images

5

500 commits

1 linked in READMEs

updated Mar 9, 2026

See the code

README

T2I-CoReBench-Images

📖 Overview

T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.

This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical consistency).


📊 Dataset Contents

  • 1,080 prompts (aligned with T2I-CoReBench) and 4 images per prompt per model
  • 40 Evaluated T2I models included (see list below)
  • Total images: (1,080 Prompts × 4 Images × 40 Models) = 172,800 Images

📌 Models Included

CategoryModels
Diffusion ModelsSD-3-Medium, SD-3.5-Medium, SD-3.5-Large, FLUX.1-schnell, FLUX.1-dev, FLUX.1-Krea-dev, FLUX.2-dev, FLUX.2-klein-4B, FLUX.2-klein-9B, PixArt-$\alpha$, PixArt-$\Sigma$, HiDream-I1, Qwen-Image, Qwen-Image-2512, HunyuanImage-3.0, Z-Image-Turbo, Z-Image, LongCat-Image
Autogressive ModelsInfinity-8B and GoT-R1-7B
Unified ModelsBAGEL, BAGEL w/ Think, show-o2-1.5B, show-o2-7B, Janus-Pro-1B, Janus-Pro-7B, BLIP3o-4B, BLIP3o-8B, OmniGen2-7B
Closed-Source ModelsSeedream 3.0, Seedream 4.0, Seedream 4.5, Gemini 2.0 Flash, Nano Banana, Nano Banana Pro, Nano Banana 2, Imagen 4, Imagen 4 Ultra, GPT-Image (GPT-4o), GPT-Image-1.5

📜 Citation

If you find this dataset useful, please cite our paper:

@inproceedings{
  li2026easier,
  title={Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?},
  author={Ouxiang Li and Yuan Wang and Xinting Hu and Huijuan Huang and Rui Chen and Jiarong Ou and Xin Tao and Pengfei Wan and Xiaojuan Qi and Fuli Feng},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
  url={https://openreview.net/forum?id=iqAFhWistW}
}

Autoregressive Models
Benchmark
Closed-Source Models
Diffusion Models
Evaluation

Contributors

lioooox

496 commits

LI
lioox

4 commits

lioooox/T2I-CoReBench-Images

Dataset

T2I-CoReBench-Images

5

500 commits

1 linked in READMEs

updated Mar 9, 2026

See the code

README

T2I-CoReBench-Images

📖 Overview

T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.

This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical consistency).


📊 Dataset Contents

  • 1,080 prompts (aligned with T2I-CoReBench) and 4 images per prompt per model
  • 40 Evaluated T2I models included (see list below)
  • Total images: (1,080 Prompts × 4 Images × 40 Models) = 172,800 Images

📌 Models Included

CategoryModels
Diffusion ModelsSD-3-Medium, SD-3.5-Medium, SD-3.5-Large, FLUX.1-schnell, FLUX.1-dev, FLUX.1-Krea-dev, FLUX.2-dev, FLUX.2-klein-4B, FLUX.2-klein-9B, PixArt-$\alpha$, PixArt-$\Sigma$, HiDream-I1, Qwen-Image, Qwen-Image-2512, HunyuanImage-3.0, Z-Image-Turbo, Z-Image, LongCat-Image
Autogressive ModelsInfinity-8B and GoT-R1-7B
Unified ModelsBAGEL, BAGEL w/ Think, show-o2-1.5B, show-o2-7B, Janus-Pro-1B, Janus-Pro-7B, BLIP3o-4B, BLIP3o-8B, OmniGen2-7B
Closed-Source ModelsSeedream 3.0, Seedream 4.0, Seedream 4.5, Gemini 2.0 Flash, Nano Banana, Nano Banana Pro, Nano Banana 2, Imagen 4, Imagen 4 Ultra, GPT-Image (GPT-4o), GPT-Image-1.5

📜 Citation

If you find this dataset useful, please cite our paper:

@inproceedings{
  li2026easier,
  title={Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?},
  author={Ouxiang Li and Yuan Wang and Xinting Hu and Huijuan Huang and Rui Chen and Jiarong Ou and Xin Tao and Pengfei Wan and Xiaojuan Qi and Fuli Feng},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
  url={https://openreview.net/forum?id=iqAFhWistW}
}

Autoregressive Models
Benchmark
Closed-Source Models
Diffusion Models
Evaluation

Contributors

lioooox

496 commits

LI
lioox

4 commits