ArabicText-Large: High-Quality Arabic Corpus for LLM Training
69
154 commits
1 linked in READMEs
updated Oct 27, 2025

ArabicText-Large is a comprehensive, high-quality Arabic text corpus comprising 743,288 articles with over 244 million words, specifically curated for Large Language Model (LLM) training and fine-tuning. This dataset represents one of the largest publicly available Arabic text collections for machine learning research.
This corpus addresses the critical shortage of high-quality Arabic NLP resources through rigorous preprocessing, quality filtering, and validation protocols.
Built by RightNow AI, the first GPU-native AI code editor.
Dataset DOI: https://doi.org/10.57967/hf/6685
| Metric | Value |
|---|---|
| Total Articles | 743,288 |
| Total Words | 244,153,780 |
| Total Sentences | 12,392,064 |
| Unique Words | 1,529,064 |
| Average Words/Article | 328.5 |
| Average Sentences/Article | 16.7 |
| Average Words/Sentence | 19.7 |
| Vocabulary Richness | 0.0063 |
| Dataset Size | 2.8 GB (compressed) |
| Arabic Content Purity | 94.2% |
| Topic Category | Articles | Percentage |
|---|---|---|
| History & Culture | 156,090 | 21.0% |
| Science & Technology | 148,657 | 20.0% |
| Geography & Places | 133,792 | 18.0% |
| Biography | 111,493 | 15.0% |
| Arts & Literature | 89,194 | 12.0% |
| Politics & Society | 74,329 | 10.0% |
| Religion | 66,863 | 9.0% |
| Sports | 51,830 | 7.0% |
| Other Topics | 22,298 | 3.0% |
| Quality Tier | Articles | Percentage |
|---|---|---|
| Excellent (≥80%) | 130,373 | 17.5% |
| Good (60-80%) | 306,526 | 41.2% |
| Fair (40-60%) | 306,389 | 41.2% |
Average Quality Score: 58.3% High-Quality Articles (≥60%): 58.7%
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("Jr23xd23/ArabicText-Large")
# Access the training split
train_data = dataset["train"]
print(f"Total articles: {len(train_data)}")
# Access a single article
article = train_data[0]
print(f"Title: {article['title']}")
print(f"Text: {article['text'][:200]}...")
import json
articles = []
with open('data.jsonl', 'r', encoding='utf-8') as f:
for line in f:
article = json.loads(line)
articles.append(article)
print(f"Loaded {len(articles)} articles")
Each entry in the dataset follows this structure:
{
"id": "unique_article_identifier",
"title": "Article Title in Arabic",
"text": "Full cleaned Arabic text content...",
"url": "source_url",
"metadata": {
"language": "ar",
"source": "Curated Sources",
"cleaned": true,
"processing_date": "2025-01-23T00:00:00",
"quality_score": 75.5
}
}
Our multi-stage processing ensures the highest quality:
Articles are retained only if they meet all criteria:
Article Lengths:
Sentence Lengths:
Word Lengths:
Most Frequent Words:
| Rank | Word (Arabic) | Translation | Frequency | Percentage |
|---|---|---|---|---|
| 1 | في | in | 9,778,012 | 4.01% |
| 2 | من | from | 7,346,952 | 3.01% |
| 3 | على | on | 3,324,220 | 1.36% |
| 4 | إلى | to | 2,453,720 | 1.01% |
| 5 | أن | that | 1,595,356 | 0.65% |
| Dataset | Words | Articles | Domain | Quality | Year | License |
|---|---|---|---|---|---|---|
| Arabic Gigaword | 848M | N/A | News | Moderate | 2011 | LDC |
| AraBERT Corpus | 70M | N/A | Mixed | Good | 2020 | MIT |
| OSCAR-Arabic | 22B | N/A | Web | Variable | 2019 | CC0 |
| mC4-Arabic | 42B | N/A | Web | Variable | 2021 | ODC-BY |
| ArabicText-Large | 244M | 743K | Encyclopedia | High | 2025 | Apache 2.0 |
Planned improvements include:
This dataset is released under the Apache License 2.0.
Copyright 2025 Jaber Jaber
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
If you use this dataset in your research, please cite:
@misc{jaber_2025,
author = {Jaber, Jaber},
title = {ArabicText-Large: A High-Quality 244-Million-Word Corpus for Arabic Language Model Training},
year = 2025,
url = {https://huggingface.co/datasets/Jr23xd23/ArabicText-Large},
doi = {10.57967/hf/6685},
publisher = {Hugging Face}
}
Research Paper:
@article{jaber2025arabictext,
title={ArabicText-Large: A High-Quality 244-Million-Word Corpus for Arabic Language Model Training},
author={Jaber, Jaber},
journal={Journal of Open Humanities Data},
year={2025},
doi={10.57967/hf/6685},
url={https://huggingface.co/datasets/Jr23xd23/ArabicText-Large}
}
We welcome community contributions:
For questions, collaborations, or research inquiries:
Author: Jaber Jaber Organization: RightNow AI Email: jaber@rightnowai.co Website: https://www.rightnowai.co
We extend our gratitude to:
Dataset Homepage: ArabicText-Large on Hugging Face DOI: https://doi.org/10.57967/hf/6685 License: Apache 2.0 Author: Jaber Jaber Year: 2025
Advancing Arabic NLP research and development
ArabicText-Large: High-Quality Arabic Corpus for LLM Training
69
154 commits
1 linked in READMEs
updated Oct 27, 2025

ArabicText-Large is a comprehensive, high-quality Arabic text corpus comprising 743,288 articles with over 244 million words, specifically curated for Large Language Model (LLM) training and fine-tuning. This dataset represents one of the largest publicly available Arabic text collections for machine learning research.
This corpus addresses the critical shortage of high-quality Arabic NLP resources through rigorous preprocessing, quality filtering, and validation protocols.
Built by RightNow AI, the first GPU-native AI code editor.
Dataset DOI: https://doi.org/10.57967/hf/6685
| Metric | Value |
|---|---|
| Total Articles | 743,288 |
| Total Words | 244,153,780 |
| Total Sentences | 12,392,064 |
| Unique Words | 1,529,064 |
| Average Words/Article | 328.5 |
| Average Sentences/Article | 16.7 |
| Average Words/Sentence | 19.7 |
| Vocabulary Richness | 0.0063 |
| Dataset Size | 2.8 GB (compressed) |
| Arabic Content Purity | 94.2% |
| Topic Category | Articles | Percentage |
|---|---|---|
| History & Culture | 156,090 | 21.0% |
| Science & Technology | 148,657 | 20.0% |
| Geography & Places | 133,792 | 18.0% |
| Biography | 111,493 | 15.0% |
| Arts & Literature | 89,194 | 12.0% |
| Politics & Society | 74,329 | 10.0% |
| Religion | 66,863 | 9.0% |
| Sports | 51,830 | 7.0% |
| Other Topics | 22,298 | 3.0% |
| Quality Tier | Articles | Percentage |
|---|---|---|
| Excellent (≥80%) | 130,373 | 17.5% |
| Good (60-80%) | 306,526 | 41.2% |
| Fair (40-60%) | 306,389 | 41.2% |
Average Quality Score: 58.3% High-Quality Articles (≥60%): 58.7%
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("Jr23xd23/ArabicText-Large")
# Access the training split
train_data = dataset["train"]
print(f"Total articles: {len(train_data)}")
# Access a single article
article = train_data[0]
print(f"Title: {article['title']}")
print(f"Text: {article['text'][:200]}...")
import json
articles = []
with open('data.jsonl', 'r', encoding='utf-8') as f:
for line in f:
article = json.loads(line)
articles.append(article)
print(f"Loaded {len(articles)} articles")
Each entry in the dataset follows this structure:
{
"id": "unique_article_identifier",
"title": "Article Title in Arabic",
"text": "Full cleaned Arabic text content...",
"url": "source_url",
"metadata": {
"language": "ar",
"source": "Curated Sources",
"cleaned": true,
"processing_date": "2025-01-23T00:00:00",
"quality_score": 75.5
}
}
Our multi-stage processing ensures the highest quality:
Articles are retained only if they meet all criteria:
Article Lengths:
Sentence Lengths:
Word Lengths:
Most Frequent Words:
| Rank | Word (Arabic) | Translation | Frequency | Percentage |
|---|---|---|---|---|
| 1 | في | in | 9,778,012 | 4.01% |
| 2 | من | from | 7,346,952 | 3.01% |
| 3 | على | on | 3,324,220 | 1.36% |
| 4 | إلى | to | 2,453,720 | 1.01% |
| 5 | أن | that | 1,595,356 | 0.65% |
| Dataset | Words | Articles | Domain | Quality | Year | License |
|---|---|---|---|---|---|---|
| Arabic Gigaword | 848M | N/A | News | Moderate | 2011 | LDC |
| AraBERT Corpus | 70M | N/A | Mixed | Good | 2020 | MIT |
| OSCAR-Arabic | 22B | N/A | Web | Variable | 2019 | CC0 |
| mC4-Arabic | 42B | N/A | Web | Variable | 2021 | ODC-BY |
| ArabicText-Large | 244M | 743K | Encyclopedia | High | 2025 | Apache 2.0 |
Planned improvements include:
This dataset is released under the Apache License 2.0.
Copyright 2025 Jaber Jaber
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
If you use this dataset in your research, please cite:
@misc{jaber_2025,
author = {Jaber, Jaber},
title = {ArabicText-Large: A High-Quality 244-Million-Word Corpus for Arabic Language Model Training},
year = 2025,
url = {https://huggingface.co/datasets/Jr23xd23/ArabicText-Large},
doi = {10.57967/hf/6685},
publisher = {Hugging Face}
}
Research Paper:
@article{jaber2025arabictext,
title={ArabicText-Large: A High-Quality 244-Million-Word Corpus for Arabic Language Model Training},
author={Jaber, Jaber},
journal={Journal of Open Humanities Data},
year={2025},
doi={10.57967/hf/6685},
url={https://huggingface.co/datasets/Jr23xd23/ArabicText-Large}
}
We welcome community contributions:
For questions, collaborations, or research inquiries:
Author: Jaber Jaber Organization: RightNow AI Email: jaber@rightnowai.co Website: https://www.rightnowai.co
We extend our gratitude to:
Dataset Homepage: ArabicText-Large on Hugging Face DOI: https://doi.org/10.57967/hf/6685 License: Apache 2.0 Author: Jaber Jaber Year: 2025
Advancing Arabic NLP research and development