biglam/europeana_newspapers_images

Dataset

Europeana Newspapers: Page Images with OCR Layout

10

13 commits

updated Sep 30, 2026

See the code

README

Europeana Newspapers: Page Images with OCR Layout

98,877 newspaper page images, 1700s–1940s, in 10 languages (plus pages marked multi-language or unidentified), each paired with the OCR produced when the page was digitised: full text, the ALTO XML, and every line and word box already converted to image pixels with the OCR engine's per-word confidence. The pages come from eight European libraries via Europeana Newspapers and are a sample of the 5.9 million pages in biglam/europeana_newspapers, joined on id.

Status: proof of concept. This is a first test release to see whether page images plus the original OCR layout are useful as a Hub dataset. The sample is designed to grow (see Sampling); the structure may still change.

What it looks like

All three images are from one page: Hamburger Nachrichten, 22 January 1840, page 7 (Hamburg State Library).

OCR line boxes drawn on the page image lines[].bbox drawn on image: the OCR layout, already in image pixels.

Word boxes coloured by OCR confidence words[].bbox coloured by words[].wc, from red (low confidence) to green (high).

One line image with its OCR text One line cut from the page with its OCR text. Even at a mean confidence of 0.79 the OCR reads "Prineipals. Rsflecttrende" for "Principals. Reflectirende": treat the text as silver, not gold.

What's in a row

ColumnDescription
imageFull-size page image as served by the library's IIIF server (not re-encoded)
idEuropeana page id; joins to biglam/europeana_newspapers (data and alto configs)
title, date, language, decade, collection, data_providerNewspaper title, issue date, language (from the source dataset), Europeana collection id, holding library
textPage text as extracted in biglam/europeana_newspapers
mean_ocrMean OCR word confidence for the page (from the source dataset)
linesOne entry per ALTO TextLine: text, bbox [x0, y0, x1, y1] in image pixels, mean_wc
wordsOne entry per ALTO String: text, bbox in image pixels, wc (OCR confidence 0–1), line index
alto_xmlThe raw ALTO v2 XML for the page
width, heightImage size in pixels
alto_unit, alto_page_sizeALTO measurement unit (pixel or mm10) and page size in ALTO units
box_alignment, alto_image_aspect_diffWhether the boxes can be trusted to land on the image; see below
page_iiif_urlIIIF URL of the page image
rightsRights statement from the Europeana record (all Public Domain Mark 1.0)
stratum, stratum_rank, tierSampling bookkeeping (see Sampling)
from datasets import load_dataset

# stream: rows carry full-size scans (~3 MB each)
ds = load_dataset("biglam/europeana_newspapers_images", split="train", streaming=True)
row = next(iter(ds))
row["image"].crop(row["lines"][0]["bbox"])  # first OCR line, cut from the page

Pick pages before downloading. The metadata config is a 7 MB file with one row per page and no images, text or ALTO (id, title, date, language, collection, mean_ocr, mean_wc, box_alignment, image size, line and word counts, page_iiif_url, and data_file, the parquet shard the page is in). Filter it first, then read images only from the shards you need:

import duckdb
from datasets import Dataset, Image

repo = "hf://datasets/biglam/europeana_newspapers_images"
# 1. choose pages from the small metadata file
picked = duckdb.sql(f"""
    select id, data_file from '{repo}/metadata/metadata.parquet'
    where "language" = 'el' and box_alignment = 'ok'
    order by mean_wc limit 5""").df()
# 2. read those pages from their shards (each page costs about one 60 MB row group)
table = duckdb.execute(
    "select * from read_parquet(?) where id in (select unnest(?))",
    [[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
).to_arrow_table()
ds = Dataset(table).cast_column("image", Image())  # a normal datasets.Dataset

This holds the selected pages in memory (roughly 3–4 MB per page). For more than about a thousand pages, write them to a local parquet file instead and load that; load_dataset memory-maps it rather than holding it in RAM:

from datasets import load_dataset

duckdb.execute(
    "copy (select * from read_parquet(?) where id in (select unnest(?))) to 'subset.parquet'",
    [[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
)
# COPY drops the image feature type, so cast it back
ds = load_dataset("parquet", data_files="subset.parquet", split="train").cast_column("image", Image())

Possible uses

  • OCR training and evaluation. Crop lines[].bbox from image and pair with lines[].text for line-level OCR data in 10 languages and several scripts (Fraktur, Cyrillic, Greek). The targets are the original OCR, so they are silver, not gold: good for pre-training or for finding hard pages, not as a benchmark truth.
  • OCR quality estimation. words[].wc and lines[].mean_wc give the original engine's confidence at word level, so bad regions of a page can be found and re-OCR'd, or left out, rather than scoring the page as a whole.
  • Re-OCR comparisons. Run a modern OCR or vision-language model on the same pages and compare against the historical OCR, per language, decade or collection.
  • Layout and reading order. Line and word boxes plus the block structure in alto_xml give page layout for multi-column historical newspapers.

Things to know before using it

The OCR is historical and uneven. It was produced by the libraries when the pages were digitised (the ALTO names ABBYY FineReader Engine for 65,636 pages and the CCS docWorks workflow for 33,241) and was taken from Europeana's 2019 full-text dumps. Quality varies by collection, typeface and paper. Some Serbian Cyrillic pages were OCR'd as Latin characters, which leaves unreadable text with plausible-looking boxes.

OCR confidence is not comparable across collections. The same print quality gets very different wc values in different collections (the engine settings differed), so compare confidence within a collection, not across them.

Image resolution depends on the library. "Full size" is whatever each IIIF server provides:

CollectionLibraryPagesMedian image width (px)10th–90th percentile
9200356National Library of Estonia22,5414,0002,460–6,127
9200300Austrian National Library16,2782,0401,527–3,206
9200339University of Belgrade13,9091,3931,019–1,834
9200338Hamburg State Library13,3434,0462,629–5,112
9200301National Library of Finland11,2772,5001,856–2,812
9200355Berlin State Library8,3433,7022,536–4,733
9200396National Library of Luxembourg6,7761,2561,256–1,256
9200357National Library of Poland6,4102,3021,621–4,031

Check box_alignment before cropping. Boxes are ALTO coordinates scaled by image size / ALTO page size. That is right when the image and the ALTO describe the same page frame. For 637 pages (610 of them Austrian National Library scans, some of which include the facing page or wide margins) the image and ALTO aspect ratios differ by more than 2%, and the boxes drift off the text; these are marked box_alignment = "unverified". In a visual review of 40 pages, the boxes on the other pages landed on the text.

Some pages were left out on purpose. Pages from collection 9200357 (National Library of Poland) that are in Izraelita or dated 1939 are excluded: for these, the OCR often belongs to a neighbouring page of the issue rather than the image. Pages whose images are no longer served (1,123 pages, about 1% of those tried, almost all from the Austrian National Library) are missing too, which is why there are 98,877 rows rather than 100,000.

Sampling

The sample is not representative of the source dataset, by design. Pages were allocated by language first (share proportional to the square root of each language's page count, with a floor so small languages are included), then across collection Γ— decade within each language, with at most 5% of pages from any one title and, where possible, one page per issue. Pages are ranked deterministically within each stratum, so a larger sample is a strict superset of this one (tier 0 is the original 1,000-page pilot).

LanguagePages
German (de)34,052
Estonian (et)12,296
Serbian (sr)11,416
French (fr)6,776
Polish (pl)6,410
Finnish (fi)6,241
multi-language5,762
Swedish (sv)5,036
Russian (ru)4,198
Greek (el)3,715
no language found2,493
Croatian (hr)482

Only pages whose page image could be matched reliably were eligible. The source dataset's item_iiif_url points to each issue's first page; the page-level image links here come from Europeana's 2019 metadata dump (edm:isShownBy + edm:hasView), checked against the live IIIF manifests and the images themselves.

How it was made

Built on Hugging Face Jobs: selection and text/ALTO extraction with DuckDB over biglam/europeana_newspapers, then images fetched from the libraries' IIIF servers at a rate each server was comfortable with (about 3–4 images/s for Europeana's server), staged in an HF bucket, and assembled into parquet. The fetcher identified itself with a User-Agent and backed off when a server slowed down.

Licence

Europeana's records mark all these items Public Domain Mark 1.0, as asserted by the holding libraries (2019 metadata). Some material from the 1910s–1940s may still be in copyright in some jurisdictions despite that mark; check before commercial reuse.

Credit

All newspapers, images and OCR were created by the Austrian National Library, National Library of Finland, Hamburg State Library, University of Belgrade, Berlin State Library, National Library of Estonia, National Library of Poland and National Library of Luxembourg, and aggregated by Europeana Newspapers. Sampled and repackaged by Daniel van Strien (Machine Learning Librarian, Hugging Face).

alto
document-layout
glam
historical
iiif
newspapers

biglam/europeana_newspapers_images

Dataset

Europeana Newspapers: Page Images with OCR Layout

10

13 commits

updated Sep 30, 2026

See the code

README

Europeana Newspapers: Page Images with OCR Layout

98,877 newspaper page images, 1700s–1940s, in 10 languages (plus pages marked multi-language or unidentified), each paired with the OCR produced when the page was digitised: full text, the ALTO XML, and every line and word box already converted to image pixels with the OCR engine's per-word confidence. The pages come from eight European libraries via Europeana Newspapers and are a sample of the 5.9 million pages in biglam/europeana_newspapers, joined on id.

Status: proof of concept. This is a first test release to see whether page images plus the original OCR layout are useful as a Hub dataset. The sample is designed to grow (see Sampling); the structure may still change.

What it looks like

All three images are from one page: Hamburger Nachrichten, 22 January 1840, page 7 (Hamburg State Library).

OCR line boxes drawn on the page image lines[].bbox drawn on image: the OCR layout, already in image pixels.

Word boxes coloured by OCR confidence words[].bbox coloured by words[].wc, from red (low confidence) to green (high).

One line image with its OCR text One line cut from the page with its OCR text. Even at a mean confidence of 0.79 the OCR reads "Prineipals. Rsflecttrende" for "Principals. Reflectirende": treat the text as silver, not gold.

What's in a row

ColumnDescription
imageFull-size page image as served by the library's IIIF server (not re-encoded)
idEuropeana page id; joins to biglam/europeana_newspapers (data and alto configs)
title, date, language, decade, collection, data_providerNewspaper title, issue date, language (from the source dataset), Europeana collection id, holding library
textPage text as extracted in biglam/europeana_newspapers
mean_ocrMean OCR word confidence for the page (from the source dataset)
linesOne entry per ALTO TextLine: text, bbox [x0, y0, x1, y1] in image pixels, mean_wc
wordsOne entry per ALTO String: text, bbox in image pixels, wc (OCR confidence 0–1), line index
alto_xmlThe raw ALTO v2 XML for the page
width, heightImage size in pixels
alto_unit, alto_page_sizeALTO measurement unit (pixel or mm10) and page size in ALTO units
box_alignment, alto_image_aspect_diffWhether the boxes can be trusted to land on the image; see below
page_iiif_urlIIIF URL of the page image
rightsRights statement from the Europeana record (all Public Domain Mark 1.0)
stratum, stratum_rank, tierSampling bookkeeping (see Sampling)
from datasets import load_dataset

# stream: rows carry full-size scans (~3 MB each)
ds = load_dataset("biglam/europeana_newspapers_images", split="train", streaming=True)
row = next(iter(ds))
row["image"].crop(row["lines"][0]["bbox"])  # first OCR line, cut from the page

Pick pages before downloading. The metadata config is a 7 MB file with one row per page and no images, text or ALTO (id, title, date, language, collection, mean_ocr, mean_wc, box_alignment, image size, line and word counts, page_iiif_url, and data_file, the parquet shard the page is in). Filter it first, then read images only from the shards you need:

import duckdb
from datasets import Dataset, Image

repo = "hf://datasets/biglam/europeana_newspapers_images"
# 1. choose pages from the small metadata file
picked = duckdb.sql(f"""
    select id, data_file from '{repo}/metadata/metadata.parquet'
    where "language" = 'el' and box_alignment = 'ok'
    order by mean_wc limit 5""").df()
# 2. read those pages from their shards (each page costs about one 60 MB row group)
table = duckdb.execute(
    "select * from read_parquet(?) where id in (select unnest(?))",
    [[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
).to_arrow_table()
ds = Dataset(table).cast_column("image", Image())  # a normal datasets.Dataset

This holds the selected pages in memory (roughly 3–4 MB per page). For more than about a thousand pages, write them to a local parquet file instead and load that; load_dataset memory-maps it rather than holding it in RAM:

from datasets import load_dataset

duckdb.execute(
    "copy (select * from read_parquet(?) where id in (select unnest(?))) to 'subset.parquet'",
    [[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
)
# COPY drops the image feature type, so cast it back
ds = load_dataset("parquet", data_files="subset.parquet", split="train").cast_column("image", Image())

Possible uses

  • OCR training and evaluation. Crop lines[].bbox from image and pair with lines[].text for line-level OCR data in 10 languages and several scripts (Fraktur, Cyrillic, Greek). The targets are the original OCR, so they are silver, not gold: good for pre-training or for finding hard pages, not as a benchmark truth.
  • OCR quality estimation. words[].wc and lines[].mean_wc give the original engine's confidence at word level, so bad regions of a page can be found and re-OCR'd, or left out, rather than scoring the page as a whole.
  • Re-OCR comparisons. Run a modern OCR or vision-language model on the same pages and compare against the historical OCR, per language, decade or collection.
  • Layout and reading order. Line and word boxes plus the block structure in alto_xml give page layout for multi-column historical newspapers.

Things to know before using it

The OCR is historical and uneven. It was produced by the libraries when the pages were digitised (the ALTO names ABBYY FineReader Engine for 65,636 pages and the CCS docWorks workflow for 33,241) and was taken from Europeana's 2019 full-text dumps. Quality varies by collection, typeface and paper. Some Serbian Cyrillic pages were OCR'd as Latin characters, which leaves unreadable text with plausible-looking boxes.

OCR confidence is not comparable across collections. The same print quality gets very different wc values in different collections (the engine settings differed), so compare confidence within a collection, not across them.

Image resolution depends on the library. "Full size" is whatever each IIIF server provides:

CollectionLibraryPagesMedian image width (px)10th–90th percentile
9200356National Library of Estonia22,5414,0002,460–6,127
9200300Austrian National Library16,2782,0401,527–3,206
9200339University of Belgrade13,9091,3931,019–1,834
9200338Hamburg State Library13,3434,0462,629–5,112
9200301National Library of Finland11,2772,5001,856–2,812
9200355Berlin State Library8,3433,7022,536–4,733
9200396National Library of Luxembourg6,7761,2561,256–1,256
9200357National Library of Poland6,4102,3021,621–4,031

Check box_alignment before cropping. Boxes are ALTO coordinates scaled by image size / ALTO page size. That is right when the image and the ALTO describe the same page frame. For 637 pages (610 of them Austrian National Library scans, some of which include the facing page or wide margins) the image and ALTO aspect ratios differ by more than 2%, and the boxes drift off the text; these are marked box_alignment = "unverified". In a visual review of 40 pages, the boxes on the other pages landed on the text.

Some pages were left out on purpose. Pages from collection 9200357 (National Library of Poland) that are in Izraelita or dated 1939 are excluded: for these, the OCR often belongs to a neighbouring page of the issue rather than the image. Pages whose images are no longer served (1,123 pages, about 1% of those tried, almost all from the Austrian National Library) are missing too, which is why there are 98,877 rows rather than 100,000.

Sampling

The sample is not representative of the source dataset, by design. Pages were allocated by language first (share proportional to the square root of each language's page count, with a floor so small languages are included), then across collection Γ— decade within each language, with at most 5% of pages from any one title and, where possible, one page per issue. Pages are ranked deterministically within each stratum, so a larger sample is a strict superset of this one (tier 0 is the original 1,000-page pilot).

LanguagePages
German (de)34,052
Estonian (et)12,296
Serbian (sr)11,416
French (fr)6,776
Polish (pl)6,410
Finnish (fi)6,241
multi-language5,762
Swedish (sv)5,036
Russian (ru)4,198
Greek (el)3,715
no language found2,493
Croatian (hr)482

Only pages whose page image could be matched reliably were eligible. The source dataset's item_iiif_url points to each issue's first page; the page-level image links here come from Europeana's 2019 metadata dump (edm:isShownBy + edm:hasView), checked against the live IIIF manifests and the images themselves.

How it was made

Built on Hugging Face Jobs: selection and text/ALTO extraction with DuckDB over biglam/europeana_newspapers, then images fetched from the libraries' IIIF servers at a rate each server was comfortable with (about 3–4 images/s for Europeana's server), staged in an HF bucket, and assembled into parquet. The fetcher identified itself with a User-Agent and backed off when a server slowed down.

Licence

Europeana's records mark all these items Public Domain Mark 1.0, as asserted by the holding libraries (2019 metadata). Some material from the 1910s–1940s may still be in copyright in some jurisdictions despite that mark; check before commercial reuse.

Credit

All newspapers, images and OCR were created by the Austrian National Library, National Library of Finland, Hamburg State Library, University of Belgrade, Berlin State Library, National Library of Estonia, National Library of Poland and National Library of Luxembourg, and aggregated by Europeana Newspapers. Sampled and repackaged by Daniel van Strien (Machine Learning Librarian, Hugging Face).

alto
document-layout
glam
historical
iiif
newspapers