This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column.
DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank:
This dataset is structured differently from the other repository ds4sd/DocLayNet, as this one includes the content (PDF cells) of the detections, and abandons the COCO format.
image: page PIL image.bboxes: a list of layout bounding boxes.category_id: a list of class ids corresponding to the bounding boxes.segmentation: a list of layout segmentation polygons.area: Area of the bboxes.pdf_cells: a list of lists corresponding to bbox. Each list contains the PDF cells (content) inside the bbox.metadata: page and document metadetails.pdf: Binary blob with the original PDF image.Bounding boxes classes / categories:
1: Caption
2: Footnote
3: Formula
4: List-item
5: Page-footer
6: Page-header
7: Picture
8: Section-header
9: Table
10: Text
11: Title
The ["metadata"]["doc_category"] field uses one of the following constants:
* financial_reports,
* scientific_articles,
* laws_and_regulations,
* government_tenders,
* manuals,
* patents
The dataset provides three splits
trainvaltestThe labeling guideline used for training of the annotation experts are available at DocLayNet_Labeling_Guide_Public.pdf.
Annotations are crowdsourced.
The dataset is curated by the Deep Search team at IBM Research. You can contact us at deepsearch-core@zurich.ibm.com.
Curators:
License: CDLA-Permissive-1.0
@article{doclaynet2022,
title = {DocLayNet: A Large Human-Annotated Dataset for Document-Layout Segmentation},
doi = {10.1145/3534678.353904},
url = {https://doi.org/10.1145/3534678.3539043},
author = {Pfitzmann, Birgit and Auer, Christoph and Dolfi, Michele and Nassar, Ahmed S and Staar, Peter W J},
year = {2022},
isbn = {9781450393850},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
booktitle = {Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining},
pages = {3743–3751},
numpages = {9},
location = {Washington DC, USA},
series = {KDD '22}
}
This dataset is an extention of the original DocLayNet dataset which embeds the PDF files of the document images inside a binary column.
DocLayNet provides page-by-page layout segmentation ground-truth using bounding-boxes for 11 distinct class labels on 80863 unique pages from 6 document categories. It provides several unique features compared to related work such as PubLayNet or DocBank:
This dataset is structured differently from the other repository ds4sd/DocLayNet, as this one includes the content (PDF cells) of the detections, and abandons the COCO format.
image: page PIL image.bboxes: a list of layout bounding boxes.category_id: a list of class ids corresponding to the bounding boxes.segmentation: a list of layout segmentation polygons.area: Area of the bboxes.pdf_cells: a list of lists corresponding to bbox. Each list contains the PDF cells (content) inside the bbox.metadata: page and document metadetails.pdf: Binary blob with the original PDF image.Bounding boxes classes / categories:
1: Caption
2: Footnote
3: Formula
4: List-item
5: Page-footer
6: Page-header
7: Picture
8: Section-header
9: Table
10: Text
11: Title
The ["metadata"]["doc_category"] field uses one of the following constants:
* financial_reports,
* scientific_articles,
* laws_and_regulations,
* government_tenders,
* manuals,
* patents
The dataset provides three splits
trainvaltestThe labeling guideline used for training of the annotation experts are available at DocLayNet_Labeling_Guide_Public.pdf.
Annotations are crowdsourced.
The dataset is curated by the Deep Search team at IBM Research. You can contact us at deepsearch-core@zurich.ibm.com.
Curators:
License: CDLA-Permissive-1.0
@article{doclaynet2022,
title = {DocLayNet: A Large Human-Annotated Dataset for Document-Layout Segmentation},
doi = {10.1145/3534678.353904},
url = {https://doi.org/10.1145/3534678.3539043},
author = {Pfitzmann, Birgit and Auer, Christoph and Dolfi, Michele and Nassar, Ahmed S and Staar, Peter W J},
year = {2022},
isbn = {9781450393850},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
booktitle = {Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining},
pages = {3743–3751},
numpages = {9},
location = {Washington DC, USA},
series = {KDD '22}
}