amalia-llm/PorTEXTO

Dataset

1

stars

2

commits

2

linked in READMEs

Aug 31, 2026

updated

benchmark
european-portuguese
ocr
visual-text-extraction

README

AMALIA Paper Data Main Repo Eval Repo

PorTEXTO

PorTEXTO is the first benchmark for contemporary and culturally relevant European Portuguese (pt-PT) visual text extraction.

While existing OCR benchmarks focus on historical artifacts or high-resource languages, PorTEXTO targets modern, real-world Portuguese content — handwritten notes, in-the-wild scene text, and synthetic images — providing a challenging evaluation suite for OCR and large vision-language models.

Key Findings

  • A sharp performance drop from synthetic to real-world samples in most models.
  • Specialized multilingual data is a better driver for pt-PT performance than model size or resolution budget.

Dataset Structure

ConfigDescriptionSamples
handwrittenHandwritten text regions cropped from scanned documents.351
handwritten_full_pageFull-page scans of handwritten documents.121
syntheticSynthetically generated text images for pt-PT OCR evaluation.200
in_the_wildReal-world scene text captured in natural environments.145

Total: 817 samples

Schema

FieldTypeDescription
imageImageThe input image (full-page scan or cropped region)
questionstringThe canonical pt-PT transcription prompt
answerstringThe ground-truth transcription

Licensing

This dataset is released under cc-by-nc-4.0 (Attribution–NonCommercial). The images were collected for this benchmark, and the ground-truth transcriptions were produced with gemini-3.1-pro. You may use the dataset for non-commercial purposes with attribution.

Citation

If you use PorTEXTO in your work, please cite:

@article{cardeira2026portexto,
  title={PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction},
  author={Cardeira, Jo{\~a}o and Gl{\'o}ria-Silva, Diogo and da Luz, Manuel Letras and Ferreira, Rafael and Tavares, Diogo and Semedo, David and Magalh{\~a}es, Jo{\~a}o},
  journal={arXiv preprint arXiv:2606.19096},
  year={2026}
}

Contributors

joao-cardeira

1 commits

JC

amalia-llm/PorTEXTO

Dataset

1

stars

2

commits

2

linked in READMEs

Aug 31, 2026

updated

benchmark
european-portuguese
ocr
visual-text-extraction

README

AMALIA Paper Data Main Repo Eval Repo

PorTEXTO

PorTEXTO is the first benchmark for contemporary and culturally relevant European Portuguese (pt-PT) visual text extraction.

While existing OCR benchmarks focus on historical artifacts or high-resource languages, PorTEXTO targets modern, real-world Portuguese content — handwritten notes, in-the-wild scene text, and synthetic images — providing a challenging evaluation suite for OCR and large vision-language models.

Key Findings

  • A sharp performance drop from synthetic to real-world samples in most models.
  • Specialized multilingual data is a better driver for pt-PT performance than model size or resolution budget.

Dataset Structure

ConfigDescriptionSamples
handwrittenHandwritten text regions cropped from scanned documents.351
handwritten_full_pageFull-page scans of handwritten documents.121
syntheticSynthetically generated text images for pt-PT OCR evaluation.200
in_the_wildReal-world scene text captured in natural environments.145

Total: 817 samples

Schema

FieldTypeDescription
imageImageThe input image (full-page scan or cropped region)
questionstringThe canonical pt-PT transcription prompt
answerstringThe ground-truth transcription

Licensing

This dataset is released under cc-by-nc-4.0 (Attribution–NonCommercial). The images were collected for this benchmark, and the ground-truth transcriptions were produced with gemini-3.1-pro. You may use the dataset for non-commercial purposes with attribution.

Citation

If you use PorTEXTO in your work, please cite:

@article{cardeira2026portexto,
  title={PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction},
  author={Cardeira, Jo{\~a}o and Gl{\'o}ria-Silva, Diogo and da Luz, Manuel Letras and Ferreira, Rafael and Tavares, Diogo and Semedo, David and Magalh{\~a}es, Jo{\~a}o},
  journal={arXiv preprint arXiv:2606.19096},
  year={2026}
}

Contributors

joao-cardeira

1 commits

JC