YesianRohn/TextSSR

[ICCV2025] TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition

Python

92

19 commits

updated Mar 6, 2026

See the code

README

🎰TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition

Intro Image

$$ TextSSR ~ Capability ~ Showcase. $$

πŸ“’News

[2025.06.26] - This paper has been accepted by ICCV2025 πŸŽ‰!

[2025.02.13] - The benchmark and evaluation code are available!

[2024.12.05] - The training dataset and generative dataset(v1: 0.43m and v2: 3.55m) are released!

[2024.12.04] - We released the latest model and online demo, check on ModelScope.

[2024.12.03] - Our paper is available at here.

πŸ“TODOs

  • Upload the new revised version of the manuscript
  • Release the expanded large-scale synthetic dataset TextSSR-F, which contains 3.55M text instances.
  • Provide publicly checkpoints and gradio demo
  • Release TextSSR-benchmark dataset and evaluation code
  • Release processed AnyWord-lmdb dataset
  • Release our scene text synthesis dataset, TextSSR-F
  • Release training and inference code

πŸ’ŽVisualization

Intro Model

$$ Model ~ Architecture ~ Display. $$

Intro Framework

$$ Data ~ Synthesis ~ Pipeline. $$

Results

$$ Results ~ Presentation. $$

πŸ› Installation

Environment Settings

  1. Clone the TextSSR Repository:

    git clone https://github.com/YesianRohn/TextSSR.git
    cd TextSSR
    
  2. Create a New Environment for TextSSR:

    conda create -n textssr python=3.10
    conda activate textssr
    
  3. Install Required Dependencies:

    • Install PyTorch, TorchVision, Torchaudio, and the necessary CUDA version:
    conda install pytorch==2.1.0 torchvision==0.16.0 torchaudio==2.1.0 pytorch-cuda=11.8 -c pytorch -c nvidia
    
    • Install the rest of the dependencies listed in the requirements.txt file:
    pip install -r requirements.txt
    
    • Install our modified diffusers:
    cd diffusers
    pip install -e .
    cd ..
    

Checkpoints/Data Preparation

  1. Data Preparation:

    • You can use the Anyword-3M dataset provided by Anytext. However, you will need to modify the data loading code to use AnyWordDataset instead of AnyWordLmdbDataset.
    • If you have obtained our AnyWord-lmdb dataset, simply place it in the TextSSR folder.
  2. Font File Preparation:

    • You can either download the Alibaba PuHuiTi font from here, which should be named AlibabaPuHuiTi-3-85-Bold.ttf, or you can use your own custom font file.
    • Place your font file in the TextSSR folder.
  3. Model Preparation:

  • If you want to train the model from scratch, first download the SD2-1 model from Hugging Face.
    • Place the downloaded model in the model folder.
    • During the training process, you will obtain several model checkpoints. These should be placed sequentially in the model folder as follows:
      • vae_ft (trained VAE model)
      • step1 (trained CDM after step 1)
      • step2 (trained CDM after step 2)

After the preparations outlined above, you will have the following file structure:

TextSSR/
β”œβ”€β”€ model/
β”‚   β”œβ”€β”€ stable-diffusion-v2-1
β”‚   β”œβ”€β”€ vae_ft
β”‚       β”œβ”€β”€ checkpoint-x/
β”‚       	β”œβ”€β”€ vae/
β”‚       	└── ...
β”‚   β”œβ”€β”€ step1
β”‚       β”œβ”€β”€ checkpoint-x/
β”‚       	β”œβ”€β”€ unet/
β”‚       	└── ...
β”‚   β”œβ”€β”€ step2
β”‚       β”œβ”€β”€ checkpoint-x/
β”‚       	β”œβ”€β”€ unet/
β”‚       	└── ...
β”‚   └── AnyWord-lmdb/                      
β”‚       β”œβ”€β”€ step1_lmdb/
β”‚       β”œβ”€β”€ step2-lmdb/
β”œβ”€β”€ AlibabaPuHuiTi-3-85-Bold.ttf
β”œβ”€β”€ ...(the same as the GitHub code)

πŸš‚ Training

  1. Step 1: Fine-tune the VAE:

    accelerate launch --num_processes 8 train_vae.py --config configs/train_vae_cfg.py
    
  2. Step 2: First stage of CDM training:

    accelerate launch --num_processes 8 train_diff.py --config configs/train_diff_step1_cfg.py
    
  3. Step 3: Second stage of CDM training:

    accelerate launch --num_processes 8 train_diff.py --config configs/train_diff_step2_cfg.py
    

πŸ” Inference

  • Ensure the benchmark path is correctly set in infer.py.
  • Run the inference process with:
    python infer.py
    

This will start the inference and generate the results.

πŸ“ŠEvaluation

See here.

πŸ”—Citation

@InProceedings{Ye_2025_ICCV,
    author    = {Ye, Xingsong and Du, Yongkun and Tao, Yunbo and Chen, Zhineng},
    title     = {TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition},
    booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
    month     = {October},
    year      = {2025},
    pages     = {17464-17473}
}

🌟 Acknowledgements

Many thanks to these great projects for their contributions, which have influenced and supported our work in various ways: SynthText, TextOCR, DiffUTE, Textdiffuser & Textdiffuser-2, AnyText, UDiffText, SceneVTG, and SVTRv2.

Special thanks also go to the training frameworks: STR-Fewer-Labels and OpenOCR.

YesianRohn/TextSSR

[ICCV2025] TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition

Python

92

19 commits

updated Mar 6, 2026

See the code

README

🎰TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition

Intro Image

$$ TextSSR ~ Capability ~ Showcase. $$

πŸ“’News

[2025.06.26] - This paper has been accepted by ICCV2025 πŸŽ‰!

[2025.02.13] - The benchmark and evaluation code are available!

[2024.12.05] - The training dataset and generative dataset(v1: 0.43m and v2: 3.55m) are released!

[2024.12.04] - We released the latest model and online demo, check on ModelScope.

[2024.12.03] - Our paper is available at here.

πŸ“TODOs

  • Upload the new revised version of the manuscript
  • Release the expanded large-scale synthetic dataset TextSSR-F, which contains 3.55M text instances.
  • Provide publicly checkpoints and gradio demo
  • Release TextSSR-benchmark dataset and evaluation code
  • Release processed AnyWord-lmdb dataset
  • Release our scene text synthesis dataset, TextSSR-F
  • Release training and inference code

πŸ’ŽVisualization

Intro Model

$$ Model ~ Architecture ~ Display. $$

Intro Framework

$$ Data ~ Synthesis ~ Pipeline. $$

Results

$$ Results ~ Presentation. $$

πŸ› Installation

Environment Settings

  1. Clone the TextSSR Repository:

    git clone https://github.com/YesianRohn/TextSSR.git
    cd TextSSR
    
  2. Create a New Environment for TextSSR:

    conda create -n textssr python=3.10
    conda activate textssr
    
  3. Install Required Dependencies:

    • Install PyTorch, TorchVision, Torchaudio, and the necessary CUDA version:
    conda install pytorch==2.1.0 torchvision==0.16.0 torchaudio==2.1.0 pytorch-cuda=11.8 -c pytorch -c nvidia
    
    • Install the rest of the dependencies listed in the requirements.txt file:
    pip install -r requirements.txt
    
    • Install our modified diffusers:
    cd diffusers
    pip install -e .
    cd ..
    

Checkpoints/Data Preparation

  1. Data Preparation:

    • You can use the Anyword-3M dataset provided by Anytext. However, you will need to modify the data loading code to use AnyWordDataset instead of AnyWordLmdbDataset.
    • If you have obtained our AnyWord-lmdb dataset, simply place it in the TextSSR folder.
  2. Font File Preparation:

    • You can either download the Alibaba PuHuiTi font from here, which should be named AlibabaPuHuiTi-3-85-Bold.ttf, or you can use your own custom font file.
    • Place your font file in the TextSSR folder.
  3. Model Preparation:

  • If you want to train the model from scratch, first download the SD2-1 model from Hugging Face.
    • Place the downloaded model in the model folder.
    • During the training process, you will obtain several model checkpoints. These should be placed sequentially in the model folder as follows:
      • vae_ft (trained VAE model)
      • step1 (trained CDM after step 1)
      • step2 (trained CDM after step 2)

After the preparations outlined above, you will have the following file structure:

TextSSR/
β”œβ”€β”€ model/
β”‚   β”œβ”€β”€ stable-diffusion-v2-1
β”‚   β”œβ”€β”€ vae_ft
β”‚       β”œβ”€β”€ checkpoint-x/
β”‚       	β”œβ”€β”€ vae/
β”‚       	└── ...
β”‚   β”œβ”€β”€ step1
β”‚       β”œβ”€β”€ checkpoint-x/
β”‚       	β”œβ”€β”€ unet/
β”‚       	└── ...
β”‚   β”œβ”€β”€ step2
β”‚       β”œβ”€β”€ checkpoint-x/
β”‚       	β”œβ”€β”€ unet/
β”‚       	└── ...
β”‚   └── AnyWord-lmdb/                      
β”‚       β”œβ”€β”€ step1_lmdb/
β”‚       β”œβ”€β”€ step2-lmdb/
β”œβ”€β”€ AlibabaPuHuiTi-3-85-Bold.ttf
β”œβ”€β”€ ...(the same as the GitHub code)

πŸš‚ Training

  1. Step 1: Fine-tune the VAE:

    accelerate launch --num_processes 8 train_vae.py --config configs/train_vae_cfg.py
    
  2. Step 2: First stage of CDM training:

    accelerate launch --num_processes 8 train_diff.py --config configs/train_diff_step1_cfg.py
    
  3. Step 3: Second stage of CDM training:

    accelerate launch --num_processes 8 train_diff.py --config configs/train_diff_step2_cfg.py
    

πŸ” Inference

  • Ensure the benchmark path is correctly set in infer.py.
  • Run the inference process with:
    python infer.py
    

This will start the inference and generate the results.

πŸ“ŠEvaluation

See here.

πŸ”—Citation

@InProceedings{Ye_2025_ICCV,
    author    = {Ye, Xingsong and Du, Yongkun and Tao, Yunbo and Chen, Zhineng},
    title     = {TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition},
    booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
    month     = {October},
    year      = {2025},
    pages     = {17464-17473}
}

🌟 Acknowledgements

Many thanks to these great projects for their contributions, which have influenced and supported our work in various ways: SynthText, TextOCR, DiffUTE, Textdiffuser & Textdiffuser-2, AnyText, UDiffText, SceneVTG, and SVTRv2.

Special thanks also go to the training frameworks: STR-Fewer-Labels and OpenOCR.

Languages

Python

98.8%

Jupyter Notebook

1.1%