LeX-10K is a curated dataset of 10K high-resolution, visually diverse 1024×1024 images tailored for text-to-image generation with a focus on aesthetics, text fidelity, and stylistic richness.
We compare LeX-10K with two widely used datasets: AnyWord-3M and MARIO-10M.
As shown below, LeX-10K significantly outperforms both in terms of aesthetic quality, text readability, and visual diversity.

Figure: Visual comparison of samples from AnyWord-3M, MARIO-10M, and LeX-10K. LeX-10K exhibits better style variety, color harmony, and clarity in text rendering.
@article{zhao2025lexart,
title={LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis},
author={Zhao, Shitian and Wu, Qilong and Li, Xinyue and Zhang, Bo and Li, Ming and Qin, Qi and Liu, Dongyang and Zhang, Kaipeng and Li, Hongsheng and Qiao, Yu and Gao, Peng and Fu, Bin and Li, Zhen},
journal={arXiv preprint arXiv:2503.21749},
year={2025}
}
LeX-10K is a curated dataset of 10K high-resolution, visually diverse 1024×1024 images tailored for text-to-image generation with a focus on aesthetics, text fidelity, and stylistic richness.
We compare LeX-10K with two widely used datasets: AnyWord-3M and MARIO-10M.
As shown below, LeX-10K significantly outperforms both in terms of aesthetic quality, text readability, and visual diversity.

Figure: Visual comparison of samples from AnyWord-3M, MARIO-10M, and LeX-10K. LeX-10K exhibits better style variety, color harmony, and clarity in text rendering.
@article{zhao2025lexart,
title={LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis},
author={Zhao, Shitian and Wu, Qilong and Li, Xinyue and Zhang, Bo and Li, Ming and Qin, Qi and Liu, Dongyang and Zhang, Kaipeng and Li, Hongsheng and Qiao, Yu and Gao, Peng and Fu, Bin and Li, Zhen},
journal={arXiv preprint arXiv:2503.21749},
year={2025}
}