Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
38
283 commits
1 linked in READMEs
updated Oct 16, 2023
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we are trying to centralize these dispersed resources into a single, comprehensive repository.
Encompassing a wide spectrum of content, ranging from social media conversations to literary masterpieces, MAD captures the rich tapestry of Arabic communication, including both standard Arabic and regional dialects.
This corpus offers comprehensive insights into the linguistic diversity and cultural nuances of Arabic expression.
If you want to use this dataset you pick one among the available configs:
Ara--MBZUAI--Bactrian-X | Ara--OpenAssistant--oasst1 | Ary--AbderrahmanSkiredj1--Darija-Wikipedia
Ara--Wikipedia | Ary--Wikipedia | Arz--Wikipedia
Ary--Ali-C137--Darija-Stories-Dataset | Ara--Ali-C137--Hindawi-Books-dataset | ``
Example of usage:
dataset = load_dataset('M-A-D/Mixed-Arabic-Datasets-Repo', 'Ara--MBZUAI--Bactrian-X')
If you loaded multiple datasets and wanted to merge them together then you can simply laverage concatenate_datasets() from datasets
dataset3 = concatenate_datasets([dataset1['train'], dataset2['train']])
Note : proccess the datasets before merging in order to make sure you have a new dataset that is consistent
The Mixed Arabic Datasets (MAD) is a dynamic and evolving collection, with its size fluctuating as new datasets are added or removed. As MAD continuously expands, it becomes a living resource that adapts to the ever-changing landscape of Arabic language datasets.
Dataset List
MAD draws from a diverse array of sources, each contributing to its richness and breadth. While the collection is constantly evolving, some of the datasets that are poised to join MAD in the near future include:
The Mixed Arabic Datasets (MAD) holds the potential to catalyze a multitude of groundbreaking applications:
MAD's access mechanism is unique: while it doesn't carry a general license itself, each constituent dataset within the corpus retains its individual license. By accessing the dataset details through the provided links in the "Dataset List" section above, users can understand the specific licensing terms for each dataset.
For discussions, contributions, and community interactions, join us on Discord!
Want to contribute to the Mixed Arabic Datasets project? Follow our comprehensive guide on Google Colab for step-by-step instructions: Contribution Guide.
Note: If you'd like to test a contribution before submitting it, feel free to do so on the MAD Test Dataset.
@dataset{
title = {Mixed Arabic Datasets (MAD)},
author = {MAD Community},
howpublished = {Dataset},
url = {https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo},
year = {2023},
}
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
38
283 commits
1 linked in READMEs
updated Oct 16, 2023
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we are trying to centralize these dispersed resources into a single, comprehensive repository.
Encompassing a wide spectrum of content, ranging from social media conversations to literary masterpieces, MAD captures the rich tapestry of Arabic communication, including both standard Arabic and regional dialects.
This corpus offers comprehensive insights into the linguistic diversity and cultural nuances of Arabic expression.
If you want to use this dataset you pick one among the available configs:
Ara--MBZUAI--Bactrian-X | Ara--OpenAssistant--oasst1 | Ary--AbderrahmanSkiredj1--Darija-Wikipedia
Ara--Wikipedia | Ary--Wikipedia | Arz--Wikipedia
Ary--Ali-C137--Darija-Stories-Dataset | Ara--Ali-C137--Hindawi-Books-dataset | ``
Example of usage:
dataset = load_dataset('M-A-D/Mixed-Arabic-Datasets-Repo', 'Ara--MBZUAI--Bactrian-X')
If you loaded multiple datasets and wanted to merge them together then you can simply laverage concatenate_datasets() from datasets
dataset3 = concatenate_datasets([dataset1['train'], dataset2['train']])
Note : proccess the datasets before merging in order to make sure you have a new dataset that is consistent
The Mixed Arabic Datasets (MAD) is a dynamic and evolving collection, with its size fluctuating as new datasets are added or removed. As MAD continuously expands, it becomes a living resource that adapts to the ever-changing landscape of Arabic language datasets.
Dataset List
MAD draws from a diverse array of sources, each contributing to its richness and breadth. While the collection is constantly evolving, some of the datasets that are poised to join MAD in the near future include:
The Mixed Arabic Datasets (MAD) holds the potential to catalyze a multitude of groundbreaking applications:
MAD's access mechanism is unique: while it doesn't carry a general license itself, each constituent dataset within the corpus retains its individual license. By accessing the dataset details through the provided links in the "Dataset List" section above, users can understand the specific licensing terms for each dataset.
For discussions, contributions, and community interactions, join us on Discord!
Want to contribute to the Mixed Arabic Datasets project? Follow our comprehensive guide on Google Colab for step-by-step instructions: Contribution Guide.
Note: If you'd like to test a contribution before submitting it, feel free to do so on the MAD Test Dataset.
@dataset{
title = {Mixed Arabic Datasets (MAD)},
author = {MAD Community},
howpublished = {Dataset},
url = {https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo},
year = {2023},
}