[GitHub]
We aim to establish a unified benchmark for training and evaluating models in scene text detection and recognition. Building on this benchmark, we introduce a general OCR system with accuracy and efficiency, OpenOCR. This repository also serves as the official codebase of the OCR team from the FVL Laboratory, Fudan University.
We sincerely welcome the researcher to recommend OCR or relevant algorithms and point out any potential factual errors or bugs. Upon receiving the suggestions, we will promptly evaluate and critically reproduce them. We look forward to collaborating with you to advance the development of OpenOCR and continuously contribute to the OCR community!
conda create -n openocr python==3.8
conda activate openocr
conda install pytorch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0 pytorch-cuda=11.8 -c pytorch -c nvidia
After installing dependencies, the following two installation methods are available. Either one can be chosen.
pip install openocr-python
Usage:
from openocr import OpenOCR
engine = OpenOCR()
img_path = '/path/img_path or /path/img_file'
result, elapse = engine(img_path)
# Server mode
# engine = OpenOCR(mode='server')
git clone https://github.com/Topdu/OpenOCR.git
cd OpenOCR
pip install -r requirements.txt
wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_det_repvit_ch.pth
wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_repsvtr_ch.pth
# Rec Server model
# wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_svtrv2_ch.pth
Usage:
# OpenOCR system: Det + Rec model
python tools/infer_e2e.py --img_path=/path/img_fold or /path/img_file
# Det model
python tools/infer_det.py --c https://github.com/Topdu/OpenOCR/tree/main/configs/det/dbnet/repvit_db.yml --o Global.infer_img=/path/img_fold or /path/img_file
# Rec model
python tools/infer_rec.py --c https://github.com/Topdu/OpenOCR/tree/main/configs/rec/svtrv2/repsvtr_ch.yml --o Global.infer_img=/path/img_fold or /path/img_file
pip install gradio==4.20.0
wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/OCR_e2e_img.tar
tar xf OCR_e2e_img.tar
# start demo
python demo_gradio.py
| Method | Venue | Training | Evaluation | Contributor |
|---|---|---|---|---|
| CRNN | TPAMI 2016 | ✅ | ✅ | |
| ASTER | TPAMI 2019 | ✅ | ✅ | pretto0 |
| NRTR | ICDAR 2019 | ✅ | ✅ | |
| SAR | AAAI 2019 | ✅ | ✅ | pretto0 |
| MORAN | PR 2019 | ✅ | ✅ | Debug |
| DAN | AAAI 2020 | ✅ | ✅ | |
| RobustScanner | ECCV 2020 | ✅ | ✅ | pretto0 |
| AutoSTR | ECCV 2020 | ✅ | ✅ | |
| SRN | CVPR 2020 | ✅ | ✅ | pretto0 |
| SEED | CVPR 2020 | ✅ | ✅ | |
| ABINet | CVPR 2021 | ✅ | ✅ | YesianRohn |
| VisionLAN | ICCV 2021 | ✅ | ✅ | YesianRohn |
| SVTR | IJCAI 2022 | ✅ | ✅ | |
| PARSeq | ECCV 2022 | ✅ | ✅ | |
| MATRN | ECCV 2022 | ✅ | ✅ | |
| MGP-STR | ECCV 2022 | ✅ | ✅ | |
| CPPD | 2023 | ✅ | ✅ | |
| LPV | IJCAI 2023 | ✅ | ✅ | |
| MAERec(Union14M) | ICCV 2023 | ✅ | ✅ | |
| LISTER | ICCV 2023 | ✅ | ✅ | |
| CDistNet | IJCV 2024 | ✅ | ✅ | YesianRohn |
| BUSNet | AAAI 2024 | ✅ | ✅ | |
| DCTC | AAAI 2024 | TODO | ||
| CAM | PR 2024 | ✅ | ✅ | |
| OTE | CVPR 2024 | ✅ | ✅ | |
| CFF | IJCAI 2024 | TODO | ||
| DPTR | ACM MM 2024 | TODO | ||
| VIPTR | ACM CIKM 2024 | TODO | ||
| IGTR | 2024 | ✅ | ✅ | |
| SMTR | 2024 | ✅ | ✅ | |
| FocalSVTR-CTC | 2024 | ✅ | ✅ | |
| SVTRv2 | 2024 | ✅ | ✅ | |
| ResNet+Trans-CTC | ✅ | ✅ | ||
| ViT-CTC | ✅ | ✅ |
Yiming Lei (pretto0) and Xingsong Ye (YesianRohn) from the FVL Laboratory, Fudan University, with guidance from Dr. Zhineng Chen, completed the majority work of the algorithm reproduction. Grateful for their outstanding contributions.
TODO
TODO
This codebase is built based on the PaddleOCR, PytorchOCR, and MMOCR. Thanks for their awesome work!
[GitHub]
We aim to establish a unified benchmark for training and evaluating models in scene text detection and recognition. Building on this benchmark, we introduce a general OCR system with accuracy and efficiency, OpenOCR. This repository also serves as the official codebase of the OCR team from the FVL Laboratory, Fudan University.
We sincerely welcome the researcher to recommend OCR or relevant algorithms and point out any potential factual errors or bugs. Upon receiving the suggestions, we will promptly evaluate and critically reproduce them. We look forward to collaborating with you to advance the development of OpenOCR and continuously contribute to the OCR community!
conda create -n openocr python==3.8
conda activate openocr
conda install pytorch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0 pytorch-cuda=11.8 -c pytorch -c nvidia
After installing dependencies, the following two installation methods are available. Either one can be chosen.
pip install openocr-python
Usage:
from openocr import OpenOCR
engine = OpenOCR()
img_path = '/path/img_path or /path/img_file'
result, elapse = engine(img_path)
# Server mode
# engine = OpenOCR(mode='server')
git clone https://github.com/Topdu/OpenOCR.git
cd OpenOCR
pip install -r requirements.txt
wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_det_repvit_ch.pth
wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_repsvtr_ch.pth
# Rec Server model
# wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_svtrv2_ch.pth
Usage:
# OpenOCR system: Det + Rec model
python tools/infer_e2e.py --img_path=/path/img_fold or /path/img_file
# Det model
python tools/infer_det.py --c https://github.com/Topdu/OpenOCR/tree/main/configs/det/dbnet/repvit_db.yml --o Global.infer_img=/path/img_fold or /path/img_file
# Rec model
python tools/infer_rec.py --c https://github.com/Topdu/OpenOCR/tree/main/configs/rec/svtrv2/repsvtr_ch.yml --o Global.infer_img=/path/img_fold or /path/img_file
pip install gradio==4.20.0
wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/OCR_e2e_img.tar
tar xf OCR_e2e_img.tar
# start demo
python demo_gradio.py
| Method | Venue | Training | Evaluation | Contributor |
|---|---|---|---|---|
| CRNN | TPAMI 2016 | ✅ | ✅ | |
| ASTER | TPAMI 2019 | ✅ | ✅ | pretto0 |
| NRTR | ICDAR 2019 | ✅ | ✅ | |
| SAR | AAAI 2019 | ✅ | ✅ | pretto0 |
| MORAN | PR 2019 | ✅ | ✅ | Debug |
| DAN | AAAI 2020 | ✅ | ✅ | |
| RobustScanner | ECCV 2020 | ✅ | ✅ | pretto0 |
| AutoSTR | ECCV 2020 | ✅ | ✅ | |
| SRN | CVPR 2020 | ✅ | ✅ | pretto0 |
| SEED | CVPR 2020 | ✅ | ✅ | |
| ABINet | CVPR 2021 | ✅ | ✅ | YesianRohn |
| VisionLAN | ICCV 2021 | ✅ | ✅ | YesianRohn |
| SVTR | IJCAI 2022 | ✅ | ✅ | |
| PARSeq | ECCV 2022 | ✅ | ✅ | |
| MATRN | ECCV 2022 | ✅ | ✅ | |
| MGP-STR | ECCV 2022 | ✅ | ✅ | |
| CPPD | 2023 | ✅ | ✅ | |
| LPV | IJCAI 2023 | ✅ | ✅ | |
| MAERec(Union14M) | ICCV 2023 | ✅ | ✅ | |
| LISTER | ICCV 2023 | ✅ | ✅ | |
| CDistNet | IJCV 2024 | ✅ | ✅ | YesianRohn |
| BUSNet | AAAI 2024 | ✅ | ✅ | |
| DCTC | AAAI 2024 | TODO | ||
| CAM | PR 2024 | ✅ | ✅ | |
| OTE | CVPR 2024 | ✅ | ✅ | |
| CFF | IJCAI 2024 | TODO | ||
| DPTR | ACM MM 2024 | TODO | ||
| VIPTR | ACM CIKM 2024 | TODO | ||
| IGTR | 2024 | ✅ | ✅ | |
| SMTR | 2024 | ✅ | ✅ | |
| FocalSVTR-CTC | 2024 | ✅ | ✅ | |
| SVTRv2 | 2024 | ✅ | ✅ | |
| ResNet+Trans-CTC | ✅ | ✅ | ||
| ViT-CTC | ✅ | ✅ |
Yiming Lei (pretto0) and Xingsong Ye (YesianRohn) from the FVL Laboratory, Fudan University, with guidance from Dr. Zhineng Chen, completed the majority work of the algorithm reproduction. Grateful for their outstanding contributions.
TODO
TODO
This codebase is built based on the PaddleOCR, PytorchOCR, and MMOCR. Thanks for their awesome work!