CSU-JPG/TextAtlas5M

Dataset

TextAtlas5M

39

37 commits

1 linked in READMEs

updated Oct 14, 2025

See the code

README

TextAtlas5M

This dataset is a training set for TextAtlas.

Paper: https://huggingface.co/papers/2502.07870

(All the data in this repo is uploaded :>)

Dataset subsets

Subsets in this dataset are CleanTextSynth, PPT2Details, PPT2Structured,LongWordsSubset-A,LongWordsSubset-M,Cover Book,Paper2Text,TextVisionBlend,StyledTextSynth and TextScenesHQ. The dataset features are as follows:

Dataset Features

  • image (img): The GT image.
  • annotation (string): The input prompt used to generate the text.
  • image_path (string): The image name.

CleanTextSynth

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "CleanTextSynth", split="train")

PPT2Details

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "PPT2Details", split="train")

PPT2Structured

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "PPT2Structured", split="train")

LongWordsSubset-A

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "LongWordsSubset-A", split="train")

LongWordsSubset-M

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "LongWordsSubset-M", split="train")

Cover Book

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "CoverBook", split="train")

Paper2Text

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "Paper2Text", split="train")

TextVisionBlend

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "TextVisionBlend", split="train")

StyledTextSynth

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "StyledTextSynth", split="train")

TextScenesHQ

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "TextScenesHQ", split="train")

Citation

If you found our work useful, please consider citing:

@article{wang2025textatlas5m,
  title={TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation},
  author={Wang, Alex Jinpeng and Mao, Dongxing and Zhang, Jiawei and Han, Weiming and Dong, Zhuobai and Li, Linjie and Lin, Yiqi and Yang, Zhengyuan and Qin, Libo and Zhang, Fuwei and others},
  journal={arXiv preprint arXiv:2502.07870},
  year={2025}
}

Contributors

neversa

36 commits

nielsr

1 commits

CSU-JPG/TextAtlas5M

Dataset

TextAtlas5M

39

37 commits

1 linked in READMEs

updated Oct 14, 2025

See the code

README

TextAtlas5M

This dataset is a training set for TextAtlas.

Paper: https://huggingface.co/papers/2502.07870

(All the data in this repo is uploaded :>)

Dataset subsets

Subsets in this dataset are CleanTextSynth, PPT2Details, PPT2Structured,LongWordsSubset-A,LongWordsSubset-M,Cover Book,Paper2Text,TextVisionBlend,StyledTextSynth and TextScenesHQ. The dataset features are as follows:

Dataset Features

  • image (img): The GT image.
  • annotation (string): The input prompt used to generate the text.
  • image_path (string): The image name.

CleanTextSynth

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "CleanTextSynth", split="train")

PPT2Details

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "PPT2Details", split="train")

PPT2Structured

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "PPT2Structured", split="train")

LongWordsSubset-A

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "LongWordsSubset-A", split="train")

LongWordsSubset-M

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "LongWordsSubset-M", split="train")

Cover Book

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "CoverBook", split="train")

Paper2Text

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "Paper2Text", split="train")

TextVisionBlend

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "TextVisionBlend", split="train")

StyledTextSynth

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "StyledTextSynth", split="train")

TextScenesHQ

To load the dataset

from datasets import load_dataset
ds = load_dataset("CSU-JPG/TextAtlas5M", "TextScenesHQ", split="train")

Citation

If you found our work useful, please consider citing:

@article{wang2025textatlas5m,
  title={TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation},
  author={Wang, Alex Jinpeng and Mao, Dongxing and Zhang, Jiawei and Han, Weiming and Dong, Zhuobai and Li, Linjie and Lin, Yiqi and Yang, Zhengyuan and Qin, Libo and Zhang, Fuwei and others},
  journal={arXiv preprint arXiv:2502.07870},
  year={2025}
}

Contributors

neversa

36 commits

nielsr

1 commits