You can find the main data card on the GEM Website.
The XWikis Corpus provides datasets with different language pairs and directions for cross-lingual and multi-lingual abstractive document summarisation.
You can load the dataset via:
import datasets
data = datasets.load_dataset('GEM/xwikis')
The data loader can be found here.
https://arxiv.org/abs/2202.09583
Laura Perez-Beltrachini (University of Edinburgh)
https://arxiv.org/abs/2202.09583
@InProceedings{clads-emnlp,
author = "Laura Perez-Beltrachini and Mirella Lapata",
title = "Models and Datasets for Cross-Lingual Summarisation",
booktitle = "Proceedings of The 2021 Conference on Empirical Methods in Natural Language Processing ",
year = "2021",
address = "Punta Cana, Dominican Republic",
}
Laura Perez-Beltrachini
no
yes
German, English, French, Czech, Chinese
cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0 International
Cross-lingual and Multi-lingual single long input document abstractive summarisation.
Summarization
Entity descriptive summarisation, that is, generate a summary that conveys the most salient facts of a document related to a given entity.
academic
Laura Perez-Beltrachini (University of Edinburgh)
Laura Perez-Beltrachini (University of Edinburgh) and Ronald Cardenas (University of Edinburgh)
For each language pair and direction there exists a train/valid/test split. The test split is a sample of size 7k from the intersection of titles existing in the four languages (cs,fr,en,de). Train/valid are randomly split.
no
no
no
ROUGE
yes
ROUGE-1/2/L
no
Found
Single website
other
not filtered
found
no
The input documents have section structure information.
validated by another rater
Bilingual annotators assessed the content overlap of source document and target summaries.
no
no PII
no
no
no
no
public domain
public domain
18 commits
5 commits
2 commits
2 commits
You can find the main data card on the GEM Website.
The XWikis Corpus provides datasets with different language pairs and directions for cross-lingual and multi-lingual abstractive document summarisation.
You can load the dataset via:
import datasets
data = datasets.load_dataset('GEM/xwikis')
The data loader can be found here.
https://arxiv.org/abs/2202.09583
Laura Perez-Beltrachini (University of Edinburgh)
https://arxiv.org/abs/2202.09583
@InProceedings{clads-emnlp,
author = "Laura Perez-Beltrachini and Mirella Lapata",
title = "Models and Datasets for Cross-Lingual Summarisation",
booktitle = "Proceedings of The 2021 Conference on Empirical Methods in Natural Language Processing ",
year = "2021",
address = "Punta Cana, Dominican Republic",
}
Laura Perez-Beltrachini
no
yes
German, English, French, Czech, Chinese
cc-by-sa-4.0: Creative Commons Attribution Share Alike 4.0 International
Cross-lingual and Multi-lingual single long input document abstractive summarisation.
Summarization
Entity descriptive summarisation, that is, generate a summary that conveys the most salient facts of a document related to a given entity.
academic
Laura Perez-Beltrachini (University of Edinburgh)
Laura Perez-Beltrachini (University of Edinburgh) and Ronald Cardenas (University of Edinburgh)
For each language pair and direction there exists a train/valid/test split. The test split is a sample of size 7k from the intersection of titles existing in the four languages (cs,fr,en,de). Train/valid are randomly split.
no
no
no
ROUGE
yes
ROUGE-1/2/L
no
Found
Single website
other
not filtered
found
no
The input documents have section structure information.
validated by another rater
Bilingual annotators assessed the content overlap of source document and target summaries.
no
no PII
no
no
no
no
public domain
public domain
18 commits
5 commits
2 commits
2 commits