98,877 newspaper page images, 1700sβ1940s, in 10 languages (plus pages marked multi-language or unidentified), each paired with the OCR produced when the page was digitised: full text, the ALTO XML, and every line and word box already converted to image pixels with the OCR engine's per-word confidence. The pages come from eight European libraries via Europeana Newspapers and are a sample of the 5.9 million pages in biglam/europeana_newspapers, joined on id.
Status: proof of concept. This is a first test release to see whether page images plus the original OCR layout are useful as a Hub dataset. The sample is designed to grow (see Sampling); the structure may still change.
All three images are from one page: Hamburger Nachrichten, 22 January 1840, page 7 (Hamburg State Library).
lines[].bbox drawn on image: the OCR layout, already in image pixels.
words[].bbox coloured by words[].wc, from red (low confidence) to green (high).
One line cut from the page with its OCR text. Even at a mean confidence of 0.79 the OCR reads "Prineipals. Rsflecttrende" for "Principals. Reflectirende": treat the text as silver, not gold.
| Column | Description |
|---|---|
image | Full-size page image as served by the library's IIIF server (not re-encoded) |
id | Europeana page id; joins to biglam/europeana_newspapers (data and alto configs) |
title, date, language, decade, collection, data_provider | Newspaper title, issue date, language (from the source dataset), Europeana collection id, holding library |
text | Page text as extracted in biglam/europeana_newspapers |
mean_ocr | Mean OCR word confidence for the page (from the source dataset) |
lines | One entry per ALTO TextLine: text, bbox [x0, y0, x1, y1] in image pixels, mean_wc |
words | One entry per ALTO String: text, bbox in image pixels, wc (OCR confidence 0β1), line index |
alto_xml | The raw ALTO v2 XML for the page |
width, height | Image size in pixels |
alto_unit, alto_page_size | ALTO measurement unit (pixel or mm10) and page size in ALTO units |
box_alignment, alto_image_aspect_diff | Whether the boxes can be trusted to land on the image; see below |
page_iiif_url | IIIF URL of the page image |
rights | Rights statement from the Europeana record (all Public Domain Mark 1.0) |
stratum, stratum_rank, tier | Sampling bookkeeping (see Sampling) |
from datasets import load_dataset
# stream: rows carry full-size scans (~3 MB each)
ds = load_dataset("biglam/europeana_newspapers_images", split="train", streaming=True)
row = next(iter(ds))
row["image"].crop(row["lines"][0]["bbox"]) # first OCR line, cut from the page
Pick pages before downloading. The metadata config is a 7 MB file with one row per page and no images, text or ALTO (id, title, date, language, collection, mean_ocr, mean_wc, box_alignment, image size, line and word counts, page_iiif_url, and data_file, the parquet shard the page is in). Filter it first, then read images only from the shards you need:
import duckdb
from datasets import Dataset, Image
repo = "hf://datasets/biglam/europeana_newspapers_images"
# 1. choose pages from the small metadata file
picked = duckdb.sql(f"""
select id, data_file from '{repo}/metadata/metadata.parquet'
where "language" = 'el' and box_alignment = 'ok'
order by mean_wc limit 5""").df()
# 2. read those pages from their shards (each page costs about one 60 MB row group)
table = duckdb.execute(
"select * from read_parquet(?) where id in (select unnest(?))",
[[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
).to_arrow_table()
ds = Dataset(table).cast_column("image", Image()) # a normal datasets.Dataset
This holds the selected pages in memory (roughly 3β4 MB per page). For more than about a thousand pages, write them to a local parquet file instead and load that; load_dataset memory-maps it rather than holding it in RAM:
from datasets import load_dataset
duckdb.execute(
"copy (select * from read_parquet(?) where id in (select unnest(?))) to 'subset.parquet'",
[[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
)
# COPY drops the image feature type, so cast it back
ds = load_dataset("parquet", data_files="subset.parquet", split="train").cast_column("image", Image())
lines[].bbox from image and pair with lines[].text for line-level OCR data in 10 languages and several scripts (Fraktur, Cyrillic, Greek). The targets are the original OCR, so they are silver, not gold: good for pre-training or for finding hard pages, not as a benchmark truth.words[].wc and lines[].mean_wc give the original engine's confidence at word level, so bad regions of a page can be found and re-OCR'd, or left out, rather than scoring the page as a whole.alto_xml give page layout for multi-column historical newspapers.The OCR is historical and uneven. It was produced by the libraries when the pages were digitised (the ALTO names ABBYY FineReader Engine for 65,636 pages and the CCS docWorks workflow for 33,241) and was taken from Europeana's 2019 full-text dumps. Quality varies by collection, typeface and paper. Some Serbian Cyrillic pages were OCR'd as Latin characters, which leaves unreadable text with plausible-looking boxes.
OCR confidence is not comparable across collections. The same print quality gets very different wc values in different collections (the engine settings differed), so compare confidence within a collection, not across them.
Image resolution depends on the library. "Full size" is whatever each IIIF server provides:
| Collection | Library | Pages | Median image width (px) | 10thβ90th percentile |
|---|---|---|---|---|
| 9200356 | National Library of Estonia | 22,541 | 4,000 | 2,460β6,127 |
| 9200300 | Austrian National Library | 16,278 | 2,040 | 1,527β3,206 |
| 9200339 | University of Belgrade | 13,909 | 1,393 | 1,019β1,834 |
| 9200338 | Hamburg State Library | 13,343 | 4,046 | 2,629β5,112 |
| 9200301 | National Library of Finland | 11,277 | 2,500 | 1,856β2,812 |
| 9200355 | Berlin State Library | 8,343 | 3,702 | 2,536β4,733 |
| 9200396 | National Library of Luxembourg | 6,776 | 1,256 | 1,256β1,256 |
| 9200357 | National Library of Poland | 6,410 | 2,302 | 1,621β4,031 |
Check box_alignment before cropping. Boxes are ALTO coordinates scaled by image size / ALTO page size. That is right when the image and the ALTO describe the same page frame. For 637 pages (610 of them Austrian National Library scans, some of which include the facing page or wide margins) the image and ALTO aspect ratios differ by more than 2%, and the boxes drift off the text; these are marked box_alignment = "unverified". In a visual review of 40 pages, the boxes on the other pages landed on the text.
Some pages were left out on purpose. Pages from collection 9200357 (National Library of Poland) that are in Izraelita or dated 1939 are excluded: for these, the OCR often belongs to a neighbouring page of the issue rather than the image. Pages whose images are no longer served (1,123 pages, about 1% of those tried, almost all from the Austrian National Library) are missing too, which is why there are 98,877 rows rather than 100,000.
The sample is not representative of the source dataset, by design. Pages were allocated by language first (share proportional to the square root of each language's page count, with a floor so small languages are included), then across collection Γ decade within each language, with at most 5% of pages from any one title and, where possible, one page per issue. Pages are ranked deterministically within each stratum, so a larger sample is a strict superset of this one (tier 0 is the original 1,000-page pilot).
| Language | Pages |
|---|---|
German (de) | 34,052 |
Estonian (et) | 12,296 |
Serbian (sr) | 11,416 |
French (fr) | 6,776 |
Polish (pl) | 6,410 |
Finnish (fi) | 6,241 |
| multi-language | 5,762 |
Swedish (sv) | 5,036 |
Russian (ru) | 4,198 |
Greek (el) | 3,715 |
| no language found | 2,493 |
Croatian (hr) | 482 |
Only pages whose page image could be matched reliably were eligible. The source dataset's item_iiif_url points to each issue's first page; the page-level image links here come from Europeana's 2019 metadata dump (edm:isShownBy + edm:hasView), checked against the live IIIF manifests and the images themselves.
Built on Hugging Face Jobs: selection and text/ALTO extraction with DuckDB over biglam/europeana_newspapers, then images fetched from the libraries' IIIF servers at a rate each server was comfortable with (about 3β4 images/s for Europeana's server), staged in an HF bucket, and assembled into parquet. The fetcher identified itself with a User-Agent and backed off when a server slowed down.
Europeana's records mark all these items Public Domain Mark 1.0, as asserted by the holding libraries (2019 metadata). Some material from the 1910sβ1940s may still be in copyright in some jurisdictions despite that mark; check before commercial reuse.
All newspapers, images and OCR were created by the Austrian National Library, National Library of Finland, Hamburg State Library, University of Belgrade, Berlin State Library, National Library of Estonia, National Library of Poland and National Library of Luxembourg, and aggregated by Europeana Newspapers. Sampled and repackaged by Daniel van Strien (Machine Learning Librarian, Hugging Face).
98,877 newspaper page images, 1700sβ1940s, in 10 languages (plus pages marked multi-language or unidentified), each paired with the OCR produced when the page was digitised: full text, the ALTO XML, and every line and word box already converted to image pixels with the OCR engine's per-word confidence. The pages come from eight European libraries via Europeana Newspapers and are a sample of the 5.9 million pages in biglam/europeana_newspapers, joined on id.
Status: proof of concept. This is a first test release to see whether page images plus the original OCR layout are useful as a Hub dataset. The sample is designed to grow (see Sampling); the structure may still change.
All three images are from one page: Hamburger Nachrichten, 22 January 1840, page 7 (Hamburg State Library).
lines[].bbox drawn on image: the OCR layout, already in image pixels.
words[].bbox coloured by words[].wc, from red (low confidence) to green (high).
One line cut from the page with its OCR text. Even at a mean confidence of 0.79 the OCR reads "Prineipals. Rsflecttrende" for "Principals. Reflectirende": treat the text as silver, not gold.
| Column | Description |
|---|---|
image | Full-size page image as served by the library's IIIF server (not re-encoded) |
id | Europeana page id; joins to biglam/europeana_newspapers (data and alto configs) |
title, date, language, decade, collection, data_provider | Newspaper title, issue date, language (from the source dataset), Europeana collection id, holding library |
text | Page text as extracted in biglam/europeana_newspapers |
mean_ocr | Mean OCR word confidence for the page (from the source dataset) |
lines | One entry per ALTO TextLine: text, bbox [x0, y0, x1, y1] in image pixels, mean_wc |
words | One entry per ALTO String: text, bbox in image pixels, wc (OCR confidence 0β1), line index |
alto_xml | The raw ALTO v2 XML for the page |
width, height | Image size in pixels |
alto_unit, alto_page_size | ALTO measurement unit (pixel or mm10) and page size in ALTO units |
box_alignment, alto_image_aspect_diff | Whether the boxes can be trusted to land on the image; see below |
page_iiif_url | IIIF URL of the page image |
rights | Rights statement from the Europeana record (all Public Domain Mark 1.0) |
stratum, stratum_rank, tier | Sampling bookkeeping (see Sampling) |
from datasets import load_dataset
# stream: rows carry full-size scans (~3 MB each)
ds = load_dataset("biglam/europeana_newspapers_images", split="train", streaming=True)
row = next(iter(ds))
row["image"].crop(row["lines"][0]["bbox"]) # first OCR line, cut from the page
Pick pages before downloading. The metadata config is a 7 MB file with one row per page and no images, text or ALTO (id, title, date, language, collection, mean_ocr, mean_wc, box_alignment, image size, line and word counts, page_iiif_url, and data_file, the parquet shard the page is in). Filter it first, then read images only from the shards you need:
import duckdb
from datasets import Dataset, Image
repo = "hf://datasets/biglam/europeana_newspapers_images"
# 1. choose pages from the small metadata file
picked = duckdb.sql(f"""
select id, data_file from '{repo}/metadata/metadata.parquet'
where "language" = 'el' and box_alignment = 'ok'
order by mean_wc limit 5""").df()
# 2. read those pages from their shards (each page costs about one 60 MB row group)
table = duckdb.execute(
"select * from read_parquet(?) where id in (select unnest(?))",
[[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
).to_arrow_table()
ds = Dataset(table).cast_column("image", Image()) # a normal datasets.Dataset
This holds the selected pages in memory (roughly 3β4 MB per page). For more than about a thousand pages, write them to a local parquet file instead and load that; load_dataset memory-maps it rather than holding it in RAM:
from datasets import load_dataset
duckdb.execute(
"copy (select * from read_parquet(?) where id in (select unnest(?))) to 'subset.parquet'",
[[f"{repo}/{f}" for f in picked.data_file.unique()], picked.id.tolist()],
)
# COPY drops the image feature type, so cast it back
ds = load_dataset("parquet", data_files="subset.parquet", split="train").cast_column("image", Image())
lines[].bbox from image and pair with lines[].text for line-level OCR data in 10 languages and several scripts (Fraktur, Cyrillic, Greek). The targets are the original OCR, so they are silver, not gold: good for pre-training or for finding hard pages, not as a benchmark truth.words[].wc and lines[].mean_wc give the original engine's confidence at word level, so bad regions of a page can be found and re-OCR'd, or left out, rather than scoring the page as a whole.alto_xml give page layout for multi-column historical newspapers.The OCR is historical and uneven. It was produced by the libraries when the pages were digitised (the ALTO names ABBYY FineReader Engine for 65,636 pages and the CCS docWorks workflow for 33,241) and was taken from Europeana's 2019 full-text dumps. Quality varies by collection, typeface and paper. Some Serbian Cyrillic pages were OCR'd as Latin characters, which leaves unreadable text with plausible-looking boxes.
OCR confidence is not comparable across collections. The same print quality gets very different wc values in different collections (the engine settings differed), so compare confidence within a collection, not across them.
Image resolution depends on the library. "Full size" is whatever each IIIF server provides:
| Collection | Library | Pages | Median image width (px) | 10thβ90th percentile |
|---|---|---|---|---|
| 9200356 | National Library of Estonia | 22,541 | 4,000 | 2,460β6,127 |
| 9200300 | Austrian National Library | 16,278 | 2,040 | 1,527β3,206 |
| 9200339 | University of Belgrade | 13,909 | 1,393 | 1,019β1,834 |
| 9200338 | Hamburg State Library | 13,343 | 4,046 | 2,629β5,112 |
| 9200301 | National Library of Finland | 11,277 | 2,500 | 1,856β2,812 |
| 9200355 | Berlin State Library | 8,343 | 3,702 | 2,536β4,733 |
| 9200396 | National Library of Luxembourg | 6,776 | 1,256 | 1,256β1,256 |
| 9200357 | National Library of Poland | 6,410 | 2,302 | 1,621β4,031 |
Check box_alignment before cropping. Boxes are ALTO coordinates scaled by image size / ALTO page size. That is right when the image and the ALTO describe the same page frame. For 637 pages (610 of them Austrian National Library scans, some of which include the facing page or wide margins) the image and ALTO aspect ratios differ by more than 2%, and the boxes drift off the text; these are marked box_alignment = "unverified". In a visual review of 40 pages, the boxes on the other pages landed on the text.
Some pages were left out on purpose. Pages from collection 9200357 (National Library of Poland) that are in Izraelita or dated 1939 are excluded: for these, the OCR often belongs to a neighbouring page of the issue rather than the image. Pages whose images are no longer served (1,123 pages, about 1% of those tried, almost all from the Austrian National Library) are missing too, which is why there are 98,877 rows rather than 100,000.
The sample is not representative of the source dataset, by design. Pages were allocated by language first (share proportional to the square root of each language's page count, with a floor so small languages are included), then across collection Γ decade within each language, with at most 5% of pages from any one title and, where possible, one page per issue. Pages are ranked deterministically within each stratum, so a larger sample is a strict superset of this one (tier 0 is the original 1,000-page pilot).
| Language | Pages |
|---|---|
German (de) | 34,052 |
Estonian (et) | 12,296 |
Serbian (sr) | 11,416 |
French (fr) | 6,776 |
Polish (pl) | 6,410 |
Finnish (fi) | 6,241 |
| multi-language | 5,762 |
Swedish (sv) | 5,036 |
Russian (ru) | 4,198 |
Greek (el) | 3,715 |
| no language found | 2,493 |
Croatian (hr) | 482 |
Only pages whose page image could be matched reliably were eligible. The source dataset's item_iiif_url points to each issue's first page; the page-level image links here come from Europeana's 2019 metadata dump (edm:isShownBy + edm:hasView), checked against the live IIIF manifests and the images themselves.
Built on Hugging Face Jobs: selection and text/ALTO extraction with DuckDB over biglam/europeana_newspapers, then images fetched from the libraries' IIIF servers at a rate each server was comfortable with (about 3β4 images/s for Europeana's server), staged in an HF bucket, and assembled into parquet. The fetcher identified itself with a User-Agent and backed off when a server slowed down.
Europeana's records mark all these items Public Domain Mark 1.0, as asserted by the holding libraries (2019 metadata). Some material from the 1910sβ1940s may still be in copyright in some jurisdictions despite that mark; check before commercial reuse.
All newspapers, images and OCR were created by the Austrian National Library, National Library of Finland, Hamburg State Library, University of Belgrade, Berlin State Library, National Library of Estonia, National Library of Poland and National Library of Luxembourg, and aggregated by Europeana Newspapers. Sampled and repackaged by Daniel van Strien (Machine Learning Librarian, Hugging Face).