MSRA Text Detection 500 Database (MSRA-TD500)
0
4 commits
1 linked in READMEs
updated Apr 30, 2024
The MSRA Text Detection 500 Database (MSRA-TD500) is a publicly released benchmark designed to evaluate text detection algorithms. This dataset aims to track recent progresses in the field of text detection within natural images, particularly focusing on texts of arbitrary orientations.
MSRA-TD500 contains 500 natural images sourced from indoor (e.g., office and mall) and outdoor (e.g., street) scenes captured with a pocket camera. The images depict various elements such as:
Images resolutions range from 1296x864 to 1920x1280. This dataset challenges users with the diversity of texts and complexity of backgrounds, featuring texts in different languages (Chinese, English, or both), fonts, sizes, colors, and orientations. Backgrounds may include elements like vegetation and repeated patterns that can be difficult to distinguish from text.
Figure 1: Typical images from MSRA-TD500 showing texts labeled as difficult due to factors like blur or occlusion.
The dataset is split into two sets:
All images are fully annotated, with the primary unit of annotation being the text line. This differs from the ICDAR datasets, which use the word as the basic unit.
Ground truth generation involves locating and bounding each text line using a four-vertex polygon, followed by fitting a minimum area rectangle around the polygon.
Figure 2: Ground truth generation process.
The evaluation protocol, designed to accommodate texts of arbitrary orientations, uses minimum area rectangles for tighter fitting. Texts labeled as "difficult" include additional challenges like small size, occlusion, blur, or truncation. Detection misses of such texts are not penalized.
Each image has a corresponding ground truth file. Each line in the file provides details about one text line, marking "difficult" texts with a label.
# Ground Truth Format Example
Index; Text Coords; Difficulty
0; x1,y1,x2,y2,x3,y3,x4,y4; 0
Figure 3: Illustration of the ground truth file format.
C. Yao, X. Bai, W. Liu, Y. Ma, and Z. Tu. "Detecting Texts of Arbitrary Orientations in Natural Images." CVPR 2012.
MSRA Text Detection 500 Database (MSRA-TD500)
0
4 commits
1 linked in READMEs
updated Apr 30, 2024
The MSRA Text Detection 500 Database (MSRA-TD500) is a publicly released benchmark designed to evaluate text detection algorithms. This dataset aims to track recent progresses in the field of text detection within natural images, particularly focusing on texts of arbitrary orientations.
MSRA-TD500 contains 500 natural images sourced from indoor (e.g., office and mall) and outdoor (e.g., street) scenes captured with a pocket camera. The images depict various elements such as:
Images resolutions range from 1296x864 to 1920x1280. This dataset challenges users with the diversity of texts and complexity of backgrounds, featuring texts in different languages (Chinese, English, or both), fonts, sizes, colors, and orientations. Backgrounds may include elements like vegetation and repeated patterns that can be difficult to distinguish from text.
Figure 1: Typical images from MSRA-TD500 showing texts labeled as difficult due to factors like blur or occlusion.
The dataset is split into two sets:
All images are fully annotated, with the primary unit of annotation being the text line. This differs from the ICDAR datasets, which use the word as the basic unit.
Ground truth generation involves locating and bounding each text line using a four-vertex polygon, followed by fitting a minimum area rectangle around the polygon.
Figure 2: Ground truth generation process.
The evaluation protocol, designed to accommodate texts of arbitrary orientations, uses minimum area rectangles for tighter fitting. Texts labeled as "difficult" include additional challenges like small size, occlusion, blur, or truncation. Detection misses of such texts are not penalized.
Each image has a corresponding ground truth file. Each line in the file provides details about one text line, marking "difficult" texts with a label.
# Ground Truth Format Example
Index; Text Coords; Difficulty
0; x1,y1,x2,y2,x3,y3,x4,y4; 0
Figure 3: Illustration of the ground truth file format.
C. Yao, X. Bai, W. Liu, Y. Ma, and Z. Tu. "Detecting Texts of Arbitrary Orientations in Natural Images." CVPR 2012.