20min-XD (20 Minuten cross-lingual document-level) is a comparable corpus of Swiss news articles in German and French, collected from the online editions of 20 Minuten and 20 minutes between 2015 and 2024. The dataset consists of 15,000 semantically aligned German and French article pairs. Unlike parallel corpora, 20min-XD captures a broad spectrum of cross-lingual similarity, ranging from near-translations to related articles covering the same event. This dataset is intended for non-commercial research use only – please refer to the accompanying license/copyright notice for details.
Document level:
id_de/fr: unique article IDscore: cosine similarity scorecontent_id_de/fr: content IDpubtime_de/fr: time of publicationarticle_link_de/fr: link to the online news articlemedium_code_de/fr: abbreviation for the name of the news outletmedium_name_de/fr: full name of the news outletlanguage_de/fr: language codechar_count_de/fr: character count of the contenthead_de/fr: article head textsubhead_de/fr: article subhead textcontent_de/fr: article content textSentence level:
id_de/fr: unique sentence IDscore: cosine similarity scoresentence_id_de/fr: document internal sentence IDaligned_article_id_de/fr: document IDsentence_de/fr: sentencechar_count_de/fr: character count of the sentencechar_count_diff: absolute difference between character counts of German and French sentenceIf you use 20min-XD in your research, please cite:
@inproceedings{wastl-et-al-2025-20min,
title = "20min-{XD}: A Comparable Corpus of {S}wiss News Articles",
author = "Wastl, Michelle and
Vamvas, Jannis and
Calleri, Selena and
Sennrich, Rico",
editor = {Gerber, Jonathan and
Cieliebak, Mark and
Tuggener, Don and
H{\"u}rlimann, Manuela},
booktitle = "Proceedings of the 10th edition of the Swiss Text Analytics Conference",
month = may,
year = "2025",
address = "Winterthur, Switzerland",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.swisstext-1.1/",
pages = "1--10"
}
20min-XD (20 Minuten cross-lingual document-level) is a comparable corpus of Swiss news articles in German and French, collected from the online editions of 20 Minuten and 20 minutes between 2015 and 2024. The dataset consists of 15,000 semantically aligned German and French article pairs. Unlike parallel corpora, 20min-XD captures a broad spectrum of cross-lingual similarity, ranging from near-translations to related articles covering the same event. This dataset is intended for non-commercial research use only – please refer to the accompanying license/copyright notice for details.
Document level:
id_de/fr: unique article IDscore: cosine similarity scorecontent_id_de/fr: content IDpubtime_de/fr: time of publicationarticle_link_de/fr: link to the online news articlemedium_code_de/fr: abbreviation for the name of the news outletmedium_name_de/fr: full name of the news outletlanguage_de/fr: language codechar_count_de/fr: character count of the contenthead_de/fr: article head textsubhead_de/fr: article subhead textcontent_de/fr: article content textSentence level:
id_de/fr: unique sentence IDscore: cosine similarity scoresentence_id_de/fr: document internal sentence IDaligned_article_id_de/fr: document IDsentence_de/fr: sentencechar_count_de/fr: character count of the sentencechar_count_diff: absolute difference between character counts of German and French sentenceIf you use 20min-XD in your research, please cite:
@inproceedings{wastl-et-al-2025-20min,
title = "20min-{XD}: A Comparable Corpus of {S}wiss News Articles",
author = "Wastl, Michelle and
Vamvas, Jannis and
Calleri, Selena and
Sennrich, Rico",
editor = {Gerber, Jonathan and
Cieliebak, Mark and
Tuggener, Don and
H{\"u}rlimann, Manuela},
booktitle = "Proceedings of the 10th edition of the Swiss Text Analytics Conference",
month = may,
year = "2025",
address = "Winterthur, Switzerland",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.swisstext-1.1/",
pages = "1--10"
}