This dataset contains machine-generated Arabic text across multiple generation methods, and Large Language Model (LLMs). It was created as part of the research paper "Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation" (https://www.sciencedirect.com/science/article/abs/pii/S0957417425042599).
The dataset addresses the need for comprehensive Arabic machine-generated text resources, enabling research in detection systems, stylometric analysis, and cross-model generalization studies.
The dataset is organized into different subsets based on generation methods:
by_polishing - Text refinement approach where models polish existing human abstractsfrom_title - Free-form generation from paper titles onlyfrom_title_and_content - Content-aware generation using both title and paper contentEach sample contains:
original_abstract: The original human-written Arabic abstract{model}_generated_abstract: Machine-generated version from each model
allam_generated_abstractjais_generated_abstractllama_generated_abstractopenai_generated_abstract| Model | Size | Domain Focus | Source |
|---|---|---|---|
| ALLaM | 7B | Arabic-focused | Open |
| Jais | 70B | Arabic-focused | Open |
| Llama 3.1 | 70B | General | Open |
| OpenAI GPT-4 | - | General | Closed |
| Generation Method | Samples | Description |
|---|---|---|
by_polishing | 2,851 | Text refinement of existing human abstracts |
from_title | 2,963 | Free-form generation from paper titles only |
from_title_and_content | 2,574 | Content-aware generation using title + paper content |
| Total | 8,388 | Across all generation methods |
| Metric | Value | Notes |
|---|---|---|
| Language | Arabic (MSA) | Modern Standard Arabic |
| Domain | Academic Abstracts | Algerian Scientific Journals |
| Source Platform | ASJP | Algerian Scientific Journals Platform |
| Time Period | 2010-2022 | Pre-AI era to avoid contamination |
| Source Papers | 2500-3,000 | Original human-written abstracts, see the paper for more details |
| Human Abstract Length | 120 words (avg) | Range: 75-294 words |
The dataset was constructed through a comprehensive collection and processing pipeline from the Algerian Scientific Journals Platform (ASJP). The process involved web scraping papers to extract metadata including titles, journal names, volumes, publication dates, and abstracts. Custom scripts using statistical analysis and rule-based methods were developed to segment multilingual abstracts (Arabic, English, French). PDF text extraction was performed using PyPDF2 with extensive preprocessing to handle Arabic script formatting challenges. Text normalization included Unicode standardization, removal of headers/footers, and whitespace standardization. Quality filtering removed generated abstracts containing error messages or falling below a 30-word threshold. The scraping and processing code is available at https://github.com/KFUPM-JRCAI/arabs-dataset. What is currently available in this dataset repository represents the final processed outcome of this entire pipeline. For access to the original papers, metadata, and links, please visit The paper repo (requires cloning with Git LFS enabled due to large file sizes). For detailed methodology and preprocessing steps, please refer to the full paper.
| Model | Title-Only | Title+Content | Polishing | Notes |
|---|---|---|---|---|
| Human | 120 words | 120 words | 120 words | Baseline |
| ALLaM | 77.2 words | 95.3 words | 104.3 words | |
| Jais | 62.3 words | 105.7 words | 68.5 words | Shortest overall |
| Llama | 99.9 words | 103.2 words | 102.3 words | Most consistent across methods |
| OpenAI | 123.3 words | 113.9 words | 165.1 words | Longest in polishing method |
These are only spotlights, please refer to the paper for more details.
from datasets import load_dataset
# Load the complete dataset
dataset = load_dataset("KFUPM-JRCAI/arabic-generated-abstracts")
# Access different generation methods
by_polishing = dataset["by_polishing"]
from_title = dataset["from_title"]
from_title_and_content = dataset["from_title_and_content"]
# Example: Get a sample
sample = dataset["by_polishing"][0]
print("Original:", sample["original_abstract"])
print("ALLaM:", sample["allam_generated_abstract"])
If you use this dataset in your research, please cite:
The Expert Systems with Applications Journal paper:
@article{al2025arabic,
title={Arabic Machine-Generated Text Detection: Stylometric Analysis and Cross-Model Evaluation},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={Expert Systems with Applications},
pages={130644},
year={2025},
publisher={Elsevier}
}
The preprint (arxiv):
@article{al2025arabic,
title={The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={arXiv preprint arXiv:2505.23276},
year={2025}
}
This work was supported by:
This dataset is intended for research purposes to:
This dataset contains machine-generated Arabic text across multiple generation methods, and Large Language Model (LLMs). It was created as part of the research paper "Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation" (https://www.sciencedirect.com/science/article/abs/pii/S0957417425042599).
The dataset addresses the need for comprehensive Arabic machine-generated text resources, enabling research in detection systems, stylometric analysis, and cross-model generalization studies.
The dataset is organized into different subsets based on generation methods:
by_polishing - Text refinement approach where models polish existing human abstractsfrom_title - Free-form generation from paper titles onlyfrom_title_and_content - Content-aware generation using both title and paper contentEach sample contains:
original_abstract: The original human-written Arabic abstract{model}_generated_abstract: Machine-generated version from each model
allam_generated_abstractjais_generated_abstractllama_generated_abstractopenai_generated_abstract| Model | Size | Domain Focus | Source |
|---|---|---|---|
| ALLaM | 7B | Arabic-focused | Open |
| Jais | 70B | Arabic-focused | Open |
| Llama 3.1 | 70B | General | Open |
| OpenAI GPT-4 | - | General | Closed |
| Generation Method | Samples | Description |
|---|---|---|
by_polishing | 2,851 | Text refinement of existing human abstracts |
from_title | 2,963 | Free-form generation from paper titles only |
from_title_and_content | 2,574 | Content-aware generation using title + paper content |
| Total | 8,388 | Across all generation methods |
| Metric | Value | Notes |
|---|---|---|
| Language | Arabic (MSA) | Modern Standard Arabic |
| Domain | Academic Abstracts | Algerian Scientific Journals |
| Source Platform | ASJP | Algerian Scientific Journals Platform |
| Time Period | 2010-2022 | Pre-AI era to avoid contamination |
| Source Papers | 2500-3,000 | Original human-written abstracts, see the paper for more details |
| Human Abstract Length | 120 words (avg) | Range: 75-294 words |
The dataset was constructed through a comprehensive collection and processing pipeline from the Algerian Scientific Journals Platform (ASJP). The process involved web scraping papers to extract metadata including titles, journal names, volumes, publication dates, and abstracts. Custom scripts using statistical analysis and rule-based methods were developed to segment multilingual abstracts (Arabic, English, French). PDF text extraction was performed using PyPDF2 with extensive preprocessing to handle Arabic script formatting challenges. Text normalization included Unicode standardization, removal of headers/footers, and whitespace standardization. Quality filtering removed generated abstracts containing error messages or falling below a 30-word threshold. The scraping and processing code is available at https://github.com/KFUPM-JRCAI/arabs-dataset. What is currently available in this dataset repository represents the final processed outcome of this entire pipeline. For access to the original papers, metadata, and links, please visit The paper repo (requires cloning with Git LFS enabled due to large file sizes). For detailed methodology and preprocessing steps, please refer to the full paper.
| Model | Title-Only | Title+Content | Polishing | Notes |
|---|---|---|---|---|
| Human | 120 words | 120 words | 120 words | Baseline |
| ALLaM | 77.2 words | 95.3 words | 104.3 words | |
| Jais | 62.3 words | 105.7 words | 68.5 words | Shortest overall |
| Llama | 99.9 words | 103.2 words | 102.3 words | Most consistent across methods |
| OpenAI | 123.3 words | 113.9 words | 165.1 words | Longest in polishing method |
These are only spotlights, please refer to the paper for more details.
from datasets import load_dataset
# Load the complete dataset
dataset = load_dataset("KFUPM-JRCAI/arabic-generated-abstracts")
# Access different generation methods
by_polishing = dataset["by_polishing"]
from_title = dataset["from_title"]
from_title_and_content = dataset["from_title_and_content"]
# Example: Get a sample
sample = dataset["by_polishing"][0]
print("Original:", sample["original_abstract"])
print("ALLaM:", sample["allam_generated_abstract"])
If you use this dataset in your research, please cite:
The Expert Systems with Applications Journal paper:
@article{al2025arabic,
title={Arabic Machine-Generated Text Detection: Stylometric Analysis and Cross-Model Evaluation},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={Expert Systems with Applications},
pages={130644},
year={2025},
publisher={Elsevier}
}
The preprint (arxiv):
@article{al2025arabic,
title={The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text},
author={Al-Shaibani, Maged S and Ahmed, Moataz},
journal={arXiv preprint arXiv:2505.23276},
year={2025}
}
This work was supported by:
This dataset is intended for research purposes to: