alenisaw/turkicocr-cyrillic

Dataset

TurkicOCR Synthetic Cyrillic Dataset

0

19 commits

1 linked in READMEs

updated Aug 15, 2026

See the code

README

TurkicOCR Synthetic Cyrillic Dataset

A large-scale synthetic dataset for document AI research in underrepresented Turkic languages — Kazakh and Kyrgyz. Built to cover the full document understanding pipeline: text detection, recognition (OCR), layout analysis, and visual document understanding (VDU).

Pages span 29 authentic document archetypes across administrative, educational, and commercial domains, rendered with 7 procedural degradation profiles that simulate real-world capture conditions — from clean office prints to aged paper, phone photographs, and official ink stamps.

Three nested configs (tiny / medium / large) enable progressive training scale without re-downloading data.

from datasets import load_dataset
ds = load_dataset("alenisaw/turkicocr-cyrillic", name="large")

Configs

ConfigTotalTrainValidationTest
tiny25,00022,5001,2501,250
medium50,00045,0002,5002,500
large100,00090,0005,0005,000

tiny ⊂ medium ⊂ large — deterministic nested views of the same generation. Images are stored as JPEG inside packed TAR shards; parquet indexes reference each page by page_id and tar_path.

Document Layouts

29 layouts across 5 categories:

CategoryLayouts
Administrative & OfficialOfficial letters, memos, meeting minutes, official statements, archival notifications, certificates
Forms & RegistriesApplication forms, simple forms, registry extracts
Books & ProseSingle/two-column book pages, dictionary entries, glossaries, indexes, academic abstracts, bulletins, historical newspapers
Educational & SpecializedSyllabi, lecture notes, exam sheets, exam registers, worksheets
Tables & TransactionalInvoices, receipts, catalog entries, attendance/schedule/simple/wide-schedule tables, inventory sheets

Degradation Profiles

7 procedurally generated visual effect profiles:

ProfileSimulates
cleanNo degradation
low_dpi_scanLow-resolution scan artifacts
office_scanOffice scanner noise and banding
official_stampedRound/rectangular ink stamps and handwritten signatures
old_paperAging, yellowing, water stains, blotches
phone_photoPerspective distortion, lens blur, camera projection
photocopyRepeated photocopy erosion and thresholding

Intended Use

For training and evaluating OCR, document layout analysis, and VDU models (LayoutLM, Donut, Pix2Struct, ColPali). The dataset is synthetic — validate on real-world documents before deployment.

Limitations

  • Synthetic content: Text is procedurally generated from corpus sources. Semantic coherence between entities (e.g. name–address binding on forms) is not guaranteed. Optimized for visual/geometric recognition, not semantic NLP tasks.
  • Domain gap: Real-world generalization should be verified on actual scanned or photographed documents.

Acknowledgements

The author would like to thank the Research and Innovation Center "CyberTech" at Astana IT University for their support and resources during the creation of this dataset.

Citation

If you use this dataset or the accompanying TurkicOCR-SVTRv2-B recognizer in your research, please cite:

@inproceedings{issayev2026turkicocr,
  title={TurkicOCR-SVTRv2-B: Lightweight Line-Grounded Recognizer for Kazakh and Kyrgyz Optical Character Recognition},
  author={Issayev, Alen and Zhalgas, Aidana},
  booktitle={Analysis of Images, Social Networks and Texts (AIST 2026)},
  series={Lecture Notes in Computer Science (LNCS)},
  publisher={Springer},
  year={2026},
  doi={10.1007/978-3-031-XXXXX-X_XX}
}

@misc{issayev_2026_turkicocr_cyrillic,
  author       = {Issayev, Alen},
  title        = {TurkicOCR-Cyrillic},
  year         = {2026},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/9255},
  url          = {https://huggingface.co/datasets/alenisaw/turkicocr-cyrillic},
  note         = {Synthetic Cyrillic OCR and document-understanding dataset}
}
cyrillic
document-ocr
kazakh
kyrgyz
synthetic-data
turkic

alenisaw/turkicocr-cyrillic

Dataset

TurkicOCR Synthetic Cyrillic Dataset

0

19 commits

1 linked in READMEs

updated Aug 15, 2026

See the code

README

TurkicOCR Synthetic Cyrillic Dataset

A large-scale synthetic dataset for document AI research in underrepresented Turkic languages — Kazakh and Kyrgyz. Built to cover the full document understanding pipeline: text detection, recognition (OCR), layout analysis, and visual document understanding (VDU).

Pages span 29 authentic document archetypes across administrative, educational, and commercial domains, rendered with 7 procedural degradation profiles that simulate real-world capture conditions — from clean office prints to aged paper, phone photographs, and official ink stamps.

Three nested configs (tiny / medium / large) enable progressive training scale without re-downloading data.

from datasets import load_dataset
ds = load_dataset("alenisaw/turkicocr-cyrillic", name="large")

Configs

ConfigTotalTrainValidationTest
tiny25,00022,5001,2501,250
medium50,00045,0002,5002,500
large100,00090,0005,0005,000

tiny ⊂ medium ⊂ large — deterministic nested views of the same generation. Images are stored as JPEG inside packed TAR shards; parquet indexes reference each page by page_id and tar_path.

Document Layouts

29 layouts across 5 categories:

CategoryLayouts
Administrative & OfficialOfficial letters, memos, meeting minutes, official statements, archival notifications, certificates
Forms & RegistriesApplication forms, simple forms, registry extracts
Books & ProseSingle/two-column book pages, dictionary entries, glossaries, indexes, academic abstracts, bulletins, historical newspapers
Educational & SpecializedSyllabi, lecture notes, exam sheets, exam registers, worksheets
Tables & TransactionalInvoices, receipts, catalog entries, attendance/schedule/simple/wide-schedule tables, inventory sheets

Degradation Profiles

7 procedurally generated visual effect profiles:

ProfileSimulates
cleanNo degradation
low_dpi_scanLow-resolution scan artifacts
office_scanOffice scanner noise and banding
official_stampedRound/rectangular ink stamps and handwritten signatures
old_paperAging, yellowing, water stains, blotches
phone_photoPerspective distortion, lens blur, camera projection
photocopyRepeated photocopy erosion and thresholding

Intended Use

For training and evaluating OCR, document layout analysis, and VDU models (LayoutLM, Donut, Pix2Struct, ColPali). The dataset is synthetic — validate on real-world documents before deployment.

Limitations

  • Synthetic content: Text is procedurally generated from corpus sources. Semantic coherence between entities (e.g. name–address binding on forms) is not guaranteed. Optimized for visual/geometric recognition, not semantic NLP tasks.
  • Domain gap: Real-world generalization should be verified on actual scanned or photographed documents.

Acknowledgements

The author would like to thank the Research and Innovation Center "CyberTech" at Astana IT University for their support and resources during the creation of this dataset.

Citation

If you use this dataset or the accompanying TurkicOCR-SVTRv2-B recognizer in your research, please cite:

@inproceedings{issayev2026turkicocr,
  title={TurkicOCR-SVTRv2-B: Lightweight Line-Grounded Recognizer for Kazakh and Kyrgyz Optical Character Recognition},
  author={Issayev, Alen and Zhalgas, Aidana},
  booktitle={Analysis of Images, Social Networks and Texts (AIST 2026)},
  series={Lecture Notes in Computer Science (LNCS)},
  publisher={Springer},
  year={2026},
  doi={10.1007/978-3-031-XXXXX-X_XX}
}

@misc{issayev_2026_turkicocr_cyrillic,
  author       = {Issayev, Alen},
  title        = {TurkicOCR-Cyrillic},
  year         = {2026},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/9255},
  url          = {https://huggingface.co/datasets/alenisaw/turkicocr-cyrillic},
  note         = {Synthetic Cyrillic OCR and document-understanding dataset}
}
cyrillic
document-ocr
kazakh
kyrgyz
synthetic-data
turkic