csebuetnlp/xlsum

Dataset

159

stars

12

commits

2

linked in READMEs

Apr 18, 2023

updated

conditional-text-generation

README

Dataset Card for "XL-Sum"

Table of Contents

Dataset Description

Dataset Summary

We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.

Supported Tasks and Leaderboards

More information needed

Languages

  • amharic
  • arabic
  • azerbaijani
  • bengali
  • burmese
  • chinese_simplified
  • chinese_traditional
  • english
  • french
  • gujarati
  • hausa
  • hindi
  • igbo
  • indonesian
  • japanese
  • kirundi
  • korean
  • kyrgyz
  • marathi
  • nepali
  • oromo
  • pashto
  • persian
  • pidgin
  • portuguese
  • punjabi
  • russian
  • scottish_gaelic
  • serbian_cyrillic
  • serbian_latin
  • sinhala
  • somali
  • spanish
  • swahili
  • tamil
  • telugu
  • thai
  • tigrinya
  • turkish
  • ukrainian
  • urdu
  • uzbek
  • vietnamese
  • welsh
  • yoruba

Dataset Structure

Data Instances

One example from the English dataset is given below in JSON format.

{
  "id": "technology-17657859",
  "url": "https://www.bbc.com/news/technology-17657859",
  "title": "Yahoo files e-book advert system patent applications",
  "summary": "Yahoo has signalled it is investigating e-book adverts as a way to stimulate its earnings.",
  "text": "Yahoo's patents suggest users could weigh the type of ads against the sizes of discount before purchase. It says in two US patent applications that ads for digital book readers have been \"less than optimal\" to date. The filings suggest that users could be offered titles at a variety of prices depending on the ads' prominence They add that the products shown could be determined by the type of book being read, or even the contents of a specific chapter, phrase or word. The paperwork was published by the US Patent and Trademark Office late last week and relates to work carried out at the firm's headquarters in Sunnyvale, California. \"Greater levels of advertising, which may be more valuable to an advertiser and potentially more distracting to an e-book reader, may warrant higher discounts,\" it states. Free books It suggests users could be offered ads as hyperlinks based within the book's text, in-laid text or even \"dynamic content\" such as video. Another idea suggests boxes at the bottom of a page could trail later chapters or quotes saying \"brought to you by Company A\". It adds that the more willing the customer is to see the ads, the greater the potential discount. \"Higher frequencies... may even be great enough to allow the e-book to be obtained for free,\" it states. The authors write that the type of ad could influence the value of the discount, with \"lower class advertising... such as teeth whitener advertisements\" offering a cheaper price than \"high\" or \"middle class\" adverts, for things like pizza. The inventors also suggest that ads could be linked to the mood or emotional state the reader is in as a they progress through a title. For example, they say if characters fall in love or show affection during a chapter, then ads for flowers or entertainment could be triggered. The patents also suggest this could applied to children's books - giving the Tom Hanks animated film Polar Express as an example. It says a scene showing a waiter giving the protagonists hot drinks \"may be an excellent opportunity to show an advertisement for hot cocoa, or a branded chocolate bar\". Another example states: \"If the setting includes young characters, a Coke advertisement could be provided, inviting the reader to enjoy a glass of Coke with his book, and providing a graphic of a cool glass.\" It adds that such targeting could be further enhanced by taking account of previous titles the owner has bought. 'Advertising-free zone' At present, several Amazon and Kobo e-book readers offer full-screen adverts when the device is switched off and show smaller ads on their menu screens, but the main text of the titles remains free of marketing. Yahoo does not currently provide ads to these devices, and a move into the area could boost its shrinking revenues. However, Philip Jones, deputy editor of the Bookseller magazine, said that the internet firm might struggle to get some of its ideas adopted. \"This has been mooted before and was fairly well decried,\" he said. \"Perhaps in a limited context it could work if the merchandise was strongly related to the title and was kept away from the text. \"But readers - particularly parents - like the fact that reading is an advertising-free zone. Authors would also want something to say about ads interrupting their narrative flow.\""
}

Data Fields

  • 'id': A string representing the article ID.
  • 'url': A string representing the article URL.
  • 'title': A string containing the article title.
  • 'summary': A string containing the article summary.
  • 'text' : A string containing the article text.

Data Splits

We used a 80%-10%-10% split for all languages with a few exceptions. English was split 93%-3.5%-3.5% for the evaluation set size to resemble that of CNN/DM and XSum; Scottish Gaelic, Kyrgyz and Sinhala had relatively fewer samples, their evaluation sets were increased to 500 samples for more reliable evaluation. Same articles were used for evaluation in the two variants of Chinese and Serbian to prevent data leakage in multilingual training. Individual dataset download links with train-dev-test example counts are given below:

LanguageISO 639-1 CodeBBC subdomain(s)TrainDevTestTotal
Amharicamhttps://www.bbc.com/amharic57617197197199
Arabicarhttps://www.bbc.com/arabic375194689468946897
Azerbaijaniazhttps://www.bbc.com/azeri64788098098096
Bengalibnhttps://www.bbc.com/bengali81021012101210126
Burmesemyhttps://www.bbc.com/burmese45695705705709
Chinese (Simplified)zh-CNhttps://www.bbc.com/ukchina/simp, https://www.bbc.com/zhongwen/simp373624670467046702
Chinese (Traditional)zh-TWhttps://www.bbc.com/ukchina/trad, https://www.bbc.com/zhongwen/trad373734670467046713
Englishenhttps://www.bbc.com/english, https://www.bbc.com/sinhala *3065221153511535329592
Frenchfrhttps://www.bbc.com/afrique86971086108610869
Gujaratiguhttps://www.bbc.com/gujarati91191139113911397
Hausahahttps://www.bbc.com/hausa64188028028022
Hindihihttps://www.bbc.com/hindi707788847884788472
Igboighttps://www.bbc.com/igbo41835225225227
Indonesianidhttps://www.bbc.com/indonesia382424780478047802
Japanesejahttps://www.bbc.com/japanese71138898898891
Kirundirnhttps://www.bbc.com/gahuza57467187187182
Koreankohttps://www.bbc.com/korean44075505505507
Kyrgyzkyhttps://www.bbc.com/kyrgyz22665005003266
Marathimrhttps://www.bbc.com/marathi109031362136213627
Nepalinphttps://www.bbc.com/nepali58087257257258
Oromoomhttps://www.bbc.com/afaanoromoo60637577577577
Pashtopshttps://www.bbc.com/pashto143531794179417941
Persianfahttps://www.bbc.com/persian472515906590659063
Pidgin**n/ahttps://www.bbc.com/pidgin92081151115111510
Portuguesepthttps://www.bbc.com/portuguese574027175717571752
Punjabipahttps://www.bbc.com/punjabi82151026102610267
Russianruhttps://www.bbc.com/russian, https://www.bbc.com/ukrainian *622437780778077803
Scottish Gaelicgdhttps://www.bbc.com/naidheachdan13135005002313
Serbian (Cyrillic)srhttps://www.bbc.com/serbian/cyr72759099099093
Serbian (Latin)srhttps://www.bbc.com/serbian/lat72769099099094
Sinhalasihttps://www.bbc.com/sinhala32495005004249
Somalisohttps://www.bbc.com/somali59627457457452
Spanisheshttps://www.bbc.com/mundo381104763476347636
Swahiliswhttps://www.bbc.com/swahili78989879879872
Tamiltahttps://www.bbc.com/tamil162222027202720276
Telugutehttps://www.bbc.com/telugu104211302130213025
Thaithhttps://www.bbc.com/thai66168268268268
Tigrinyatihttps://www.bbc.com/tigrinya54516816816813
Turkishtrhttps://www.bbc.com/turkce271763397339733970
Ukrainianukhttps://www.bbc.com/ukrainian432015399539953999
Urduurhttps://www.bbc.com/urdu676658458845884581
Uzbekuzhttps://www.bbc.com/uzbek47285905905908
Vietnamesevihttps://www.bbc.com/vietnamese321114013401340137
Welshcyhttps://www.bbc.com/cymrufyw97321216121612164
Yorubayohttps://www.bbc.com/yoruba63507937937936

* A lot of articles in BBC Sinhala and BBC Ukrainian were written in English and Russian respectively. They were identified using Fasttext and moved accordingly.

** West African Pidgin English

Dataset Creation

Curation Rationale

More information needed

Source Data

BBC News

Initial Data Collection and Normalization

Detailed in the paper

Who are the source language producers?

Detailed in the paper

Annotations

Detailed in the paper

Annotation process

Detailed in the paper

Who are the annotators?

Detailed in the paper

Personal and Sensitive Information

More information needed

Considerations for Using the Data

Social Impact of Dataset

More information needed

Discussion of Biases

More information needed

Other Known Limitations

More information needed

Additional Information

Dataset Curators

More information needed

Licensing Information

Contents of this repository are restricted to only non-commercial research purposes under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). Copyright of the dataset contents belongs to the original copyright holders.

Citation Information

If you use any of the datasets, models or code modules, please cite the following paper:

@inproceedings{hasan-etal-2021-xl,
    title = "{XL}-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages",
    author = "Hasan, Tahmid  and
      Bhattacharjee, Abhik  and
      Islam, Md. Saiful  and
      Mubasshir, Kazi  and
      Li, Yuan-Fang  and
      Kang, Yong-Bin  and
      Rahman, M. Sohel  and
      Shahriyar, Rifat",
    booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021",
    month = aug,
    year = "2021",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.findings-acl.413",
    pages = "4693--4703",
}

Contributions

Thanks to @abhik1505040 and @Tahmid for adding this dataset.

Contributors

abhik1505040

7 commits

Tahmid

2 commits

julien-c

1 commits

csebuetnlp/xlsum

Dataset

159

stars

12

commits

2

linked in READMEs

Apr 18, 2023

updated

conditional-text-generation

README

Dataset Card for "XL-Sum"

Table of Contents

Dataset Description

Dataset Summary

We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.

Supported Tasks and Leaderboards

More information needed

Languages

  • amharic
  • arabic
  • azerbaijani
  • bengali
  • burmese
  • chinese_simplified
  • chinese_traditional
  • english
  • french
  • gujarati
  • hausa
  • hindi
  • igbo
  • indonesian
  • japanese
  • kirundi
  • korean
  • kyrgyz
  • marathi
  • nepali
  • oromo
  • pashto
  • persian
  • pidgin
  • portuguese
  • punjabi
  • russian
  • scottish_gaelic
  • serbian_cyrillic
  • serbian_latin
  • sinhala
  • somali
  • spanish
  • swahili
  • tamil
  • telugu
  • thai
  • tigrinya
  • turkish
  • ukrainian
  • urdu
  • uzbek
  • vietnamese
  • welsh
  • yoruba

Dataset Structure

Data Instances

One example from the English dataset is given below in JSON format.

{
  "id": "technology-17657859",
  "url": "https://www.bbc.com/news/technology-17657859",
  "title": "Yahoo files e-book advert system patent applications",
  "summary": "Yahoo has signalled it is investigating e-book adverts as a way to stimulate its earnings.",
  "text": "Yahoo's patents suggest users could weigh the type of ads against the sizes of discount before purchase. It says in two US patent applications that ads for digital book readers have been \"less than optimal\" to date. The filings suggest that users could be offered titles at a variety of prices depending on the ads' prominence They add that the products shown could be determined by the type of book being read, or even the contents of a specific chapter, phrase or word. The paperwork was published by the US Patent and Trademark Office late last week and relates to work carried out at the firm's headquarters in Sunnyvale, California. \"Greater levels of advertising, which may be more valuable to an advertiser and potentially more distracting to an e-book reader, may warrant higher discounts,\" it states. Free books It suggests users could be offered ads as hyperlinks based within the book's text, in-laid text or even \"dynamic content\" such as video. Another idea suggests boxes at the bottom of a page could trail later chapters or quotes saying \"brought to you by Company A\". It adds that the more willing the customer is to see the ads, the greater the potential discount. \"Higher frequencies... may even be great enough to allow the e-book to be obtained for free,\" it states. The authors write that the type of ad could influence the value of the discount, with \"lower class advertising... such as teeth whitener advertisements\" offering a cheaper price than \"high\" or \"middle class\" adverts, for things like pizza. The inventors also suggest that ads could be linked to the mood or emotional state the reader is in as a they progress through a title. For example, they say if characters fall in love or show affection during a chapter, then ads for flowers or entertainment could be triggered. The patents also suggest this could applied to children's books - giving the Tom Hanks animated film Polar Express as an example. It says a scene showing a waiter giving the protagonists hot drinks \"may be an excellent opportunity to show an advertisement for hot cocoa, or a branded chocolate bar\". Another example states: \"If the setting includes young characters, a Coke advertisement could be provided, inviting the reader to enjoy a glass of Coke with his book, and providing a graphic of a cool glass.\" It adds that such targeting could be further enhanced by taking account of previous titles the owner has bought. 'Advertising-free zone' At present, several Amazon and Kobo e-book readers offer full-screen adverts when the device is switched off and show smaller ads on their menu screens, but the main text of the titles remains free of marketing. Yahoo does not currently provide ads to these devices, and a move into the area could boost its shrinking revenues. However, Philip Jones, deputy editor of the Bookseller magazine, said that the internet firm might struggle to get some of its ideas adopted. \"This has been mooted before and was fairly well decried,\" he said. \"Perhaps in a limited context it could work if the merchandise was strongly related to the title and was kept away from the text. \"But readers - particularly parents - like the fact that reading is an advertising-free zone. Authors would also want something to say about ads interrupting their narrative flow.\""
}

Data Fields

  • 'id': A string representing the article ID.
  • 'url': A string representing the article URL.
  • 'title': A string containing the article title.
  • 'summary': A string containing the article summary.
  • 'text' : A string containing the article text.

Data Splits

We used a 80%-10%-10% split for all languages with a few exceptions. English was split 93%-3.5%-3.5% for the evaluation set size to resemble that of CNN/DM and XSum; Scottish Gaelic, Kyrgyz and Sinhala had relatively fewer samples, their evaluation sets were increased to 500 samples for more reliable evaluation. Same articles were used for evaluation in the two variants of Chinese and Serbian to prevent data leakage in multilingual training. Individual dataset download links with train-dev-test example counts are given below:

LanguageISO 639-1 CodeBBC subdomain(s)TrainDevTestTotal
Amharicamhttps://www.bbc.com/amharic57617197197199
Arabicarhttps://www.bbc.com/arabic375194689468946897
Azerbaijaniazhttps://www.bbc.com/azeri64788098098096
Bengalibnhttps://www.bbc.com/bengali81021012101210126
Burmesemyhttps://www.bbc.com/burmese45695705705709
Chinese (Simplified)zh-CNhttps://www.bbc.com/ukchina/simp, https://www.bbc.com/zhongwen/simp373624670467046702
Chinese (Traditional)zh-TWhttps://www.bbc.com/ukchina/trad, https://www.bbc.com/zhongwen/trad373734670467046713
Englishenhttps://www.bbc.com/english, https://www.bbc.com/sinhala *3065221153511535329592
Frenchfrhttps://www.bbc.com/afrique86971086108610869
Gujaratiguhttps://www.bbc.com/gujarati91191139113911397
Hausahahttps://www.bbc.com/hausa64188028028022
Hindihihttps://www.bbc.com/hindi707788847884788472
Igboighttps://www.bbc.com/igbo41835225225227
Indonesianidhttps://www.bbc.com/indonesia382424780478047802
Japanesejahttps://www.bbc.com/japanese71138898898891
Kirundirnhttps://www.bbc.com/gahuza57467187187182
Koreankohttps://www.bbc.com/korean44075505505507
Kyrgyzkyhttps://www.bbc.com/kyrgyz22665005003266
Marathimrhttps://www.bbc.com/marathi109031362136213627
Nepalinphttps://www.bbc.com/nepali58087257257258
Oromoomhttps://www.bbc.com/afaanoromoo60637577577577
Pashtopshttps://www.bbc.com/pashto143531794179417941
Persianfahttps://www.bbc.com/persian472515906590659063
Pidgin**n/ahttps://www.bbc.com/pidgin92081151115111510
Portuguesepthttps://www.bbc.com/portuguese574027175717571752
Punjabipahttps://www.bbc.com/punjabi82151026102610267
Russianruhttps://www.bbc.com/russian, https://www.bbc.com/ukrainian *622437780778077803
Scottish Gaelicgdhttps://www.bbc.com/naidheachdan13135005002313
Serbian (Cyrillic)srhttps://www.bbc.com/serbian/cyr72759099099093
Serbian (Latin)srhttps://www.bbc.com/serbian/lat72769099099094
Sinhalasihttps://www.bbc.com/sinhala32495005004249
Somalisohttps://www.bbc.com/somali59627457457452
Spanisheshttps://www.bbc.com/mundo381104763476347636
Swahiliswhttps://www.bbc.com/swahili78989879879872
Tamiltahttps://www.bbc.com/tamil162222027202720276
Telugutehttps://www.bbc.com/telugu104211302130213025
Thaithhttps://www.bbc.com/thai66168268268268
Tigrinyatihttps://www.bbc.com/tigrinya54516816816813
Turkishtrhttps://www.bbc.com/turkce271763397339733970
Ukrainianukhttps://www.bbc.com/ukrainian432015399539953999
Urduurhttps://www.bbc.com/urdu676658458845884581
Uzbekuzhttps://www.bbc.com/uzbek47285905905908
Vietnamesevihttps://www.bbc.com/vietnamese321114013401340137
Welshcyhttps://www.bbc.com/cymrufyw97321216121612164
Yorubayohttps://www.bbc.com/yoruba63507937937936

* A lot of articles in BBC Sinhala and BBC Ukrainian were written in English and Russian respectively. They were identified using Fasttext and moved accordingly.

** West African Pidgin English

Dataset Creation

Curation Rationale

More information needed

Source Data

BBC News

Initial Data Collection and Normalization

Detailed in the paper

Who are the source language producers?

Detailed in the paper

Annotations

Detailed in the paper

Annotation process

Detailed in the paper

Who are the annotators?

Detailed in the paper

Personal and Sensitive Information

More information needed

Considerations for Using the Data

Social Impact of Dataset

More information needed

Discussion of Biases

More information needed

Other Known Limitations

More information needed

Additional Information

Dataset Curators

More information needed

Licensing Information

Contents of this repository are restricted to only non-commercial research purposes under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). Copyright of the dataset contents belongs to the original copyright holders.

Citation Information

If you use any of the datasets, models or code modules, please cite the following paper:

@inproceedings{hasan-etal-2021-xl,
    title = "{XL}-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages",
    author = "Hasan, Tahmid  and
      Bhattacharjee, Abhik  and
      Islam, Md. Saiful  and
      Mubasshir, Kazi  and
      Li, Yuan-Fang  and
      Kang, Yong-Bin  and
      Rahman, M. Sohel  and
      Shahriyar, Rifat",
    booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021",
    month = aug,
    year = "2021",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.findings-acl.413",
    pages = "4693--4703",
}

Contributions

Thanks to @abhik1505040 and @Tahmid for adding this dataset.

Contributors

abhik1505040

7 commits

Tahmid

2 commits

julien-c

1 commits