This dataset contains 4,322 parallel Egyptian Arabic-English dialogue pairs with automatic domain classification. The data is extracted from TV series subtitles and features natural conversational Egyptian Arabic dialect (العامية المصرية).
Egyptian Arabic is one of the most widely spoken Arabic dialects, used by over 100 million speakers. This dataset provides:
Each entry contains:
{
"id": "ep01_line0001",
"arabic": "خلاويص؟",
"english": "Ready or not?",
"episode": 1,
"dialect": "egyptian",
"language": "ar",
"language_variant": "ar_EG",
"genre": "dialogue",
"domain": "general"
}
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier (format: epXX_lineYYYY) |
arabic | string | Egyptian Arabic text |
english | string | English translation |
episode | int | Episode number (for context) |
dialect | string | Dialect identifier (always "egyptian") |
language | string | ISO language code (always "ar") |
language_variant | string | Specific variant code (always "ar_EG") |
genre | string | Content genre (dialogue/narration) |
domain | string | Auto-detected content domain |
| Domain | Count | Percentage |
|---|---|---|
| general | 2,143 | 49.6% |
| technology | 531 | 12.3% |
| family | 368 | 8.5% |
| horror | 281 | 6.5% |
| medical | 233 | 5.4% |
| romance | 136 | 3.1% |
| weather | 115 | 2.7% |
| food | 104 | 2.4% |
| paranormal | 86 | 2.0% |
| social | 55 | 1.3% |
| Episode | Entries |
|---|---|
| Episode 1 | 889 |
| Episode 2 | 782 |
| Episode 3 | 584 |
| Episode 4 | 907 |
| Episode 5 | 554 |
| Episode 6 | 606 |
This dataset includes automatic domain classification using keyword-based detection:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("fr3on/egyptian-dialogue")
# Access the data
print(dataset['train'][0])
# Filter by domain
medical_data = dataset['train'].filter(lambda x: x['domain'] == 'medical')
# Filter by episode
episode_1 = dataset['train'].filter(lambda x: x['episode'] == 1)
import pandas as pd
# Load Parquet file directly
df = pd.read_parquet("data/train-00000-of-00001.parquet")
# Analyze domains
print(df['domain'].value_counts())
# Filter and export
medical_df = df[df['domain'] == 'medical']
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, Seq2SeqTrainer
# Load dataset
dataset = load_dataset("fr3on/egyptian-dialogue")
# Load model for Arabic-English translation
model_name = "Helsinki-NLP/opus-mt-ar-en"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
# Tokenize
def preprocess(examples):
inputs = tokenizer(examples['arabic'], truncation=True, max_length=128)
targets = tokenizer(examples['english'], truncation=True, max_length=128)
inputs['labels'] = targets['input_ids']
return inputs
tokenized = dataset.map(preprocess, batched=True)
# Train
trainer = Seq2SeqTrainer(
model=model,
train_dataset=tokenized['train'],
eval_dataset=tokenized['test']
)
trainer.train()
from datasets import load_dataset
dataset = load_dataset("fr3on/egyptian-dialogue")
# Train separate models per domain
for domain in ['medical', 'legal', 'technology']:
domain_data = dataset['train'].filter(lambda x: x['domain'] == domain)
# Train domain-specific model
print(f"Training {domain} model with {len(domain_data)} examples")
Egyptian Arabic differs significantly from Modern Standard Arabic (MSA):
This dataset is released under the CC BY 4.0 License.
If you use this dataset in your research, please cite:
@dataset{egyptian_dialogue_2026,
title={Egyptian Arabic Dialogue Dataset},
author={fr3on},
year={2025},
publisher={Hugging Face},
url={https://huggingface.co/datasets/fr3on/egyptian-dialogue}
}
Keywords: Egyptian Arabic, ar_EG, dialect, colloquial, translation, dialogue, domain classification, NLP, machine translation, Arabic dialects, conversational AI, parquet
Dataset Size: 4,322 examples | Format: Parquet | License: CC BY 4.0
This dataset contains 4,322 parallel Egyptian Arabic-English dialogue pairs with automatic domain classification. The data is extracted from TV series subtitles and features natural conversational Egyptian Arabic dialect (العامية المصرية).
Egyptian Arabic is one of the most widely spoken Arabic dialects, used by over 100 million speakers. This dataset provides:
Each entry contains:
{
"id": "ep01_line0001",
"arabic": "خلاويص؟",
"english": "Ready or not?",
"episode": 1,
"dialect": "egyptian",
"language": "ar",
"language_variant": "ar_EG",
"genre": "dialogue",
"domain": "general"
}
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier (format: epXX_lineYYYY) |
arabic | string | Egyptian Arabic text |
english | string | English translation |
episode | int | Episode number (for context) |
dialect | string | Dialect identifier (always "egyptian") |
language | string | ISO language code (always "ar") |
language_variant | string | Specific variant code (always "ar_EG") |
genre | string | Content genre (dialogue/narration) |
domain | string | Auto-detected content domain |
| Domain | Count | Percentage |
|---|---|---|
| general | 2,143 | 49.6% |
| technology | 531 | 12.3% |
| family | 368 | 8.5% |
| horror | 281 | 6.5% |
| medical | 233 | 5.4% |
| romance | 136 | 3.1% |
| weather | 115 | 2.7% |
| food | 104 | 2.4% |
| paranormal | 86 | 2.0% |
| social | 55 | 1.3% |
| Episode | Entries |
|---|---|
| Episode 1 | 889 |
| Episode 2 | 782 |
| Episode 3 | 584 |
| Episode 4 | 907 |
| Episode 5 | 554 |
| Episode 6 | 606 |
This dataset includes automatic domain classification using keyword-based detection:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("fr3on/egyptian-dialogue")
# Access the data
print(dataset['train'][0])
# Filter by domain
medical_data = dataset['train'].filter(lambda x: x['domain'] == 'medical')
# Filter by episode
episode_1 = dataset['train'].filter(lambda x: x['episode'] == 1)
import pandas as pd
# Load Parquet file directly
df = pd.read_parquet("data/train-00000-of-00001.parquet")
# Analyze domains
print(df['domain'].value_counts())
# Filter and export
medical_df = df[df['domain'] == 'medical']
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, Seq2SeqTrainer
# Load dataset
dataset = load_dataset("fr3on/egyptian-dialogue")
# Load model for Arabic-English translation
model_name = "Helsinki-NLP/opus-mt-ar-en"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
# Tokenize
def preprocess(examples):
inputs = tokenizer(examples['arabic'], truncation=True, max_length=128)
targets = tokenizer(examples['english'], truncation=True, max_length=128)
inputs['labels'] = targets['input_ids']
return inputs
tokenized = dataset.map(preprocess, batched=True)
# Train
trainer = Seq2SeqTrainer(
model=model,
train_dataset=tokenized['train'],
eval_dataset=tokenized['test']
)
trainer.train()
from datasets import load_dataset
dataset = load_dataset("fr3on/egyptian-dialogue")
# Train separate models per domain
for domain in ['medical', 'legal', 'technology']:
domain_data = dataset['train'].filter(lambda x: x['domain'] == domain)
# Train domain-specific model
print(f"Training {domain} model with {len(domain_data)} examples")
Egyptian Arabic differs significantly from Modern Standard Arabic (MSA):
This dataset is released under the CC BY 4.0 License.
If you use this dataset in your research, please cite:
@dataset{egyptian_dialogue_2026,
title={Egyptian Arabic Dialogue Dataset},
author={fr3on},
year={2025},
publisher={Hugging Face},
url={https://huggingface.co/datasets/fr3on/egyptian-dialogue}
}
Keywords: Egyptian Arabic, ar_EG, dialect, colloquial, translation, dialogue, domain classification, NLP, machine translation, Arabic dialects, conversational AI, parquet
Dataset Size: 4,322 examples | Format: Parquet | License: CC BY 4.0