Arabic Machine-Generated Social Media Posts Dataset
1
13 commits
1 linked in READMEs
updated May 22, 2026
This dataset contains machine-generated Arabic social media posts using a text polishing approach across multiple Large Language Model (LLMs). It was created as part of the research paper: "Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation" (https://www.sciencedirect.com/science/article/abs/pii/S0957417425042599).
The dataset addresses the need for comprehensive Arabic machine-generated text resources in informal/social media contexts, enabling research in detection systems, stylometric analysis, and cross-model generalization studies in casual Arabic writing.
by_polishing - Text refinement approach where models polish existing human social media posts while preserving dialectal expressions, diacritical marks, and the informal nature of social media writing.
Each sample contains:
original_post: The original human-written Arabic social media post{model}_generated_post: Machine-generated polished version from each model
allam_generated_postjais_generated_postllama_generated_postopenai_generated_post| Model | Size | Domain Focus | Source |
|---|---|---|---|
| ALLaM | 7B | Arabic-focused | Open |
| Jais | 70B | Arabic-focused | Open |
| Llama 3.1 | 70B | General | Open |
| OpenAI GPT-4 | - | General | Closed |
| Metric | Value | Notes |
|---|---|---|
| Total Samples | 3,318 | Single generation method (polishing) |
| Language | Arabic (MSA + ~Dialectal/informal) | Modern Standard Arabic with informal elements |
| Domain | Social Media Posts | Book and hotel reviews |
| Source Datasets | BRAD + HARD | Book Reviews (BRAD) + Hotel Reviews (HARD) |
| Generation Method | Polishing only | Text refinement preserving style |
| Metric | Value | Notes |
|---|---|---|
| BRAD Samples | 3,000 | Book reviews from Goodreads.com |
| HARD Samples | 500 | Hotel reviews from Booking.com |
| Human Post Length | 867.4 words (avg) | Range: 135-1,546 words |
| BRAD Length Range | 724-1,500 words | Naturally longer book reviews |
| HARD Length Range | 150-614 words | Shorter hotel reviews |
| Total samples (after filtering) | 3318 |
The dataset was constructed from two prominent Arabic review collections: BRAD (Book Reviews in Arabic Dataset) collected from Goodreads.com and HARD (Hotel Arabic Reviews Dataset) collected from Booking.com. Both datasets primarily contain Modern Standard Arabic text with loose language and informal tone. The selection focused on obtaining longer-form reviews suitable for meaningful linguistic analysis. From BRAD, 3,000 reviews were selected containing between 724-1,500 words per review, while from HARD, 500 reviews ranging from 150-614 words were extracted. Preprocessing steps included removal of special characters and non-printable text, normalization of Arabic text through tatweel removal, and standardization of repeated punctuation marks (limiting repetitions to a maximum of 3). After generation, we filter samples by removing invalid generated posts, setting a minimum threshold of 50 words per post, and dropping duplicated samples (if there are any). The resulting final dataset after this filtration contains 3,318 samples. What is currently available in this dataset repository represents the final processed outcome of this entire pipeline. For access to the original papers, metadata, and links, please visit The paper repo (requires cloning with Git LFS enabled due to large file sizes). For further details, please refer to the full paper.
| Model | Generated Length (avg) | % of Human Length | Max Length | Notes |
|---|---|---|---|---|
| Human | 867.4 words | 100% | 1,546 words | Baseline |
| ALLaM | 627.4 words | 72% | 2,705 words | Closest to human length |
| OpenAI | 449.5 words | 52% | 1,761 words | |
| Jais | 305.3 words | 35% | 409 words | |
| Llama | 225.3 words | 26% | 443 words |
These are only spotlights, please refer to the paper for more details.
from datasets import load_dataset
# Load the complete dataset
dataset = load_dataset("KFUPM-JRCAI/arabic-generated-social-media-posts")
# Example: Get a sample
sample = dataset[0]
print("Original:", sample["original_post"])
print("ALLaM:", sample["allam_generated_post"])
print("Jais:", sample["jais_generated_post"])
If you use this dataset in your research, please cite:
The Expert Systems with Applications Journal paper:
@article{al2025arabic,
title={Arabic Machine-Generated Text Detection: Stylometric Analysis and Cross-Model Evaluation},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={Expert Systems with Applications},
pages={130644},
year={2025},
publisher={Elsevier}
}
The preprint (arxiv):
@article{al2025arabic,
title={The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={arXiv preprint arXiv:2505.23276},
year={2025}
}
This work was supported by:
This dataset is intended for research purposes to:
Please use responsibly and in accordance with your institution's research ethics guidelines.
Arabic Machine-Generated Social Media Posts Dataset
1
13 commits
1 linked in READMEs
updated May 22, 2026
This dataset contains machine-generated Arabic social media posts using a text polishing approach across multiple Large Language Model (LLMs). It was created as part of the research paper: "Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation" (https://www.sciencedirect.com/science/article/abs/pii/S0957417425042599).
The dataset addresses the need for comprehensive Arabic machine-generated text resources in informal/social media contexts, enabling research in detection systems, stylometric analysis, and cross-model generalization studies in casual Arabic writing.
by_polishing - Text refinement approach where models polish existing human social media posts while preserving dialectal expressions, diacritical marks, and the informal nature of social media writing.
Each sample contains:
original_post: The original human-written Arabic social media post{model}_generated_post: Machine-generated polished version from each model
allam_generated_postjais_generated_postllama_generated_postopenai_generated_post| Model | Size | Domain Focus | Source |
|---|---|---|---|
| ALLaM | 7B | Arabic-focused | Open |
| Jais | 70B | Arabic-focused | Open |
| Llama 3.1 | 70B | General | Open |
| OpenAI GPT-4 | - | General | Closed |
| Metric | Value | Notes |
|---|---|---|
| Total Samples | 3,318 | Single generation method (polishing) |
| Language | Arabic (MSA + ~Dialectal/informal) | Modern Standard Arabic with informal elements |
| Domain | Social Media Posts | Book and hotel reviews |
| Source Datasets | BRAD + HARD | Book Reviews (BRAD) + Hotel Reviews (HARD) |
| Generation Method | Polishing only | Text refinement preserving style |
| Metric | Value | Notes |
|---|---|---|
| BRAD Samples | 3,000 | Book reviews from Goodreads.com |
| HARD Samples | 500 | Hotel reviews from Booking.com |
| Human Post Length | 867.4 words (avg) | Range: 135-1,546 words |
| BRAD Length Range | 724-1,500 words | Naturally longer book reviews |
| HARD Length Range | 150-614 words | Shorter hotel reviews |
| Total samples (after filtering) | 3318 |
The dataset was constructed from two prominent Arabic review collections: BRAD (Book Reviews in Arabic Dataset) collected from Goodreads.com and HARD (Hotel Arabic Reviews Dataset) collected from Booking.com. Both datasets primarily contain Modern Standard Arabic text with loose language and informal tone. The selection focused on obtaining longer-form reviews suitable for meaningful linguistic analysis. From BRAD, 3,000 reviews were selected containing between 724-1,500 words per review, while from HARD, 500 reviews ranging from 150-614 words were extracted. Preprocessing steps included removal of special characters and non-printable text, normalization of Arabic text through tatweel removal, and standardization of repeated punctuation marks (limiting repetitions to a maximum of 3). After generation, we filter samples by removing invalid generated posts, setting a minimum threshold of 50 words per post, and dropping duplicated samples (if there are any). The resulting final dataset after this filtration contains 3,318 samples. What is currently available in this dataset repository represents the final processed outcome of this entire pipeline. For access to the original papers, metadata, and links, please visit The paper repo (requires cloning with Git LFS enabled due to large file sizes). For further details, please refer to the full paper.
| Model | Generated Length (avg) | % of Human Length | Max Length | Notes |
|---|---|---|---|---|
| Human | 867.4 words | 100% | 1,546 words | Baseline |
| ALLaM | 627.4 words | 72% | 2,705 words | Closest to human length |
| OpenAI | 449.5 words | 52% | 1,761 words | |
| Jais | 305.3 words | 35% | 409 words | |
| Llama | 225.3 words | 26% | 443 words |
These are only spotlights, please refer to the paper for more details.
from datasets import load_dataset
# Load the complete dataset
dataset = load_dataset("KFUPM-JRCAI/arabic-generated-social-media-posts")
# Example: Get a sample
sample = dataset[0]
print("Original:", sample["original_post"])
print("ALLaM:", sample["allam_generated_post"])
print("Jais:", sample["jais_generated_post"])
If you use this dataset in your research, please cite:
The Expert Systems with Applications Journal paper:
@article{al2025arabic,
title={Arabic Machine-Generated Text Detection: Stylometric Analysis and Cross-Model Evaluation},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={Expert Systems with Applications},
pages={130644},
year={2025},
publisher={Elsevier}
}
The preprint (arxiv):
@article{al2025arabic,
title={The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={arXiv preprint arXiv:2505.23276},
year={2025}
}
This work was supported by:
This dataset is intended for research purposes to:
Please use responsibly and in accordance with your institution's research ethics guidelines.