Cleaned, section-chunked training corpus for a small bilingual LLM (Arabic + English), combining a curated Egyptian-history collection with general Wikipedia coverage from CohereLabs/wikipedia-2023-11-embed-multilingual-v3.
Every pretrain record is a contiguous span of 120-1500 characters with the
article title and section headings removed. Provenance is in ds_source.
| Config | Split | Rows | Training phase | Purpose |
|---|---|---|---|---|
pretrain | train | 7,053,893 | Pretrain | Bilingual Wikipedia corpus, AR+EN, shuffled |
pretrain | eval | 114,259 | Pretrain | Held-out perplexity eval (leakage-safe by article) |
pretrain | phase1_train | 175,476 | Legacy | Previous phase-1 split, kept for compatibility (see below) |
pretrain | phase1_eval | 2,000 | Legacy | Previous phase-1 eval split |
pretrain | phase2_train | 94,486 | Legacy | Previous phase-2 (Egypt-domain) split |
pretrain | phase2_eval | 2,001 | Legacy | Previous phase-2 eval split |
finetune | train | 57,488 | Finetune | SFT pairs, loss-masked to answer at train time |
finetune | eval | 6,379 | Finetune | Held-out QA eval, AR+EN (filter via language) |
from datasets import load_dataset
REPO = "bakrianoo/jabarti-llm-dataset"
pretrain = load_dataset(REPO, "pretrain", split="train") # or split="eval"
finetune = load_dataset(REPO, "finetune", split="train") # or split="eval"
# Optional: keep only the curated Egyptian-history records
jabarti = pretrain.filter(lambda r: r["ds_source"] == "jabarti")
pretrain)ds_source | language | train records | share of characters |
|---|---|---|---|
cohere-wiki | en | 3,515,381 | 50.3% |
cohere-wiki | ar | 3,225,427 | 39.5% |
jabarti | en | 220,445 | 7.2% |
jabarti | ar | 92,640 | 3.0% |
jabarti records carry the full curation metadata below; cohere-wiki
records are null on those fields and populate only text, title, url,
language, record_id, page_id, ds_source and sample_weight.
Because jabarti is ~4% of records, use ds_source and egypt_relevance
with sampling weights if you want to emphasise the Egyptian-history domain.
Eval is held out by whole article across both sources β an article is
entirely in train or entirely in eval, never split across the boundary.
Verified: 0 of 21,989 eval articles appear in train. Articles are matched
on the normalised Wikipedia page, so the same article under different internal
ids cannot straddle the split.
pretrain config)text β record body; article title and section headings removedds_source β jabarti | cohere-wikirecord_id β stable per-record idpage_id β groups records from the same article (used for the eval holdout)sample_weight β 1/sqrt(records from this article); corrects the bias
toward long articles. Recomputed for this chunking.ortho_stripped β Arabic written without hamza forms (see caveat below)language β ar or enchunk_index, chunk_total β position within the source articlechunk_strategy β section | paragraph | wholeegypt_relevance β core | related | incidental | nonecontent_quality β high | medium | low | unusableis_person β whether the article is a biographytopic_domain β history | biography | sports | arts | science | politics | geography | otherimportance β high | medium | low (based on Wikipedia pageviews)ortho_convention β native | stripped | mixed (AR only; EN rows are empty string)source_type β native_wiki | llm_generated (jabarti only)pretrain)phase1_train, phase1_eval, phase2_train and phase2_eval re-present the
previous phase-organized pretrain splits, unchanged in content, from revision
2cad63a,
so existing pipelines keep working on the latest revision:
p1_train = load_dataset(REPO, "pretrain", split="phase1_train")
They share the current schema; differences from train / eval:
text is the original chunk with title prefix and section headings.sample_weight keeps the original value, 1/sqrt(chunk_total).ds_source is jabarti; section_index = 0, section_total = 1 (each
record is a whole original chunk).record_id is <split>.p<phase>.<lang>_<article#>_c<chunk>, distinct from
the ids in train / eval.ortho_stripped is recomputed per record (Arabic, >=50 Arabic letters,
<12 hamza forms per 1k chars), which approximates the current flag.section_title column is dropped. The heading is already present in
text for every record that had one.phase2_train includes phase-1 articles by design, and the legacy
splits overlap with train / eval. Do not mix them with the new splits.llm_generated portion of jabarti was written
without hamza forms (~5 per 1k chars against ~28 in native Arabic Wikipedia)
and essentially without diacritics. Those records are flagged
ortho_stripped = true; filter or down-weight them if standard orthography
matters for your model.jabarti and cohere-wiki both draw on Wikipedia,
so some articles appear under both sources with different chunk boundaries.
Exact-duplicate text is removed; re-chunked duplicates are retained
deliberately, acting as upsampling of that content.finetune config)question, answer β bilingual Q&A pairlanguage β ar or entype β factual | temporal | analytical | synthesis | unanswerabledifficulty β easy | medium | hardsource_article_id β links back to the phase_2 source articlearticle_title β title of the source articleCleaned, section-chunked training corpus for a small bilingual LLM (Arabic + English), combining a curated Egyptian-history collection with general Wikipedia coverage from CohereLabs/wikipedia-2023-11-embed-multilingual-v3.
Every pretrain record is a contiguous span of 120-1500 characters with the
article title and section headings removed. Provenance is in ds_source.
| Config | Split | Rows | Training phase | Purpose |
|---|---|---|---|---|
pretrain | train | 7,053,893 | Pretrain | Bilingual Wikipedia corpus, AR+EN, shuffled |
pretrain | eval | 114,259 | Pretrain | Held-out perplexity eval (leakage-safe by article) |
pretrain | phase1_train | 175,476 | Legacy | Previous phase-1 split, kept for compatibility (see below) |
pretrain | phase1_eval | 2,000 | Legacy | Previous phase-1 eval split |
pretrain | phase2_train | 94,486 | Legacy | Previous phase-2 (Egypt-domain) split |
pretrain | phase2_eval | 2,001 | Legacy | Previous phase-2 eval split |
finetune | train | 57,488 | Finetune | SFT pairs, loss-masked to answer at train time |
finetune | eval | 6,379 | Finetune | Held-out QA eval, AR+EN (filter via language) |
from datasets import load_dataset
REPO = "bakrianoo/jabarti-llm-dataset"
pretrain = load_dataset(REPO, "pretrain", split="train") # or split="eval"
finetune = load_dataset(REPO, "finetune", split="train") # or split="eval"
# Optional: keep only the curated Egyptian-history records
jabarti = pretrain.filter(lambda r: r["ds_source"] == "jabarti")
pretrain)ds_source | language | train records | share of characters |
|---|---|---|---|
cohere-wiki | en | 3,515,381 | 50.3% |
cohere-wiki | ar | 3,225,427 | 39.5% |
jabarti | en | 220,445 | 7.2% |
jabarti | ar | 92,640 | 3.0% |
jabarti records carry the full curation metadata below; cohere-wiki
records are null on those fields and populate only text, title, url,
language, record_id, page_id, ds_source and sample_weight.
Because jabarti is ~4% of records, use ds_source and egypt_relevance
with sampling weights if you want to emphasise the Egyptian-history domain.
Eval is held out by whole article across both sources β an article is
entirely in train or entirely in eval, never split across the boundary.
Verified: 0 of 21,989 eval articles appear in train. Articles are matched
on the normalised Wikipedia page, so the same article under different internal
ids cannot straddle the split.
pretrain config)text β record body; article title and section headings removedds_source β jabarti | cohere-wikirecord_id β stable per-record idpage_id β groups records from the same article (used for the eval holdout)sample_weight β 1/sqrt(records from this article); corrects the bias
toward long articles. Recomputed for this chunking.ortho_stripped β Arabic written without hamza forms (see caveat below)language β ar or enchunk_index, chunk_total β position within the source articlechunk_strategy β section | paragraph | wholeegypt_relevance β core | related | incidental | nonecontent_quality β high | medium | low | unusableis_person β whether the article is a biographytopic_domain β history | biography | sports | arts | science | politics | geography | otherimportance β high | medium | low (based on Wikipedia pageviews)ortho_convention β native | stripped | mixed (AR only; EN rows are empty string)source_type β native_wiki | llm_generated (jabarti only)pretrain)phase1_train, phase1_eval, phase2_train and phase2_eval re-present the
previous phase-organized pretrain splits, unchanged in content, from revision
2cad63a,
so existing pipelines keep working on the latest revision:
p1_train = load_dataset(REPO, "pretrain", split="phase1_train")
They share the current schema; differences from train / eval:
text is the original chunk with title prefix and section headings.sample_weight keeps the original value, 1/sqrt(chunk_total).ds_source is jabarti; section_index = 0, section_total = 1 (each
record is a whole original chunk).record_id is <split>.p<phase>.<lang>_<article#>_c<chunk>, distinct from
the ids in train / eval.ortho_stripped is recomputed per record (Arabic, >=50 Arabic letters,
<12 hamza forms per 1k chars), which approximates the current flag.section_title column is dropped. The heading is already present in
text for every record that had one.phase2_train includes phase-1 articles by design, and the legacy
splits overlap with train / eval. Do not mix them with the new splits.llm_generated portion of jabarti was written
without hamza forms (~5 per 1k chars against ~28 in native Arabic Wikipedia)
and essentially without diacritics. Those records are flagged
ortho_stripped = true; filter or down-weight them if standard orthography
matters for your model.jabarti and cohere-wiki both draw on Wikipedia,
so some articles appear under both sources with different chunk boundaries.
Exact-duplicate text is removed; re-chunked duplicates are retained
deliberately, acting as upsampling of that content.finetune config)question, answer β bilingual Q&A pairlanguage β ar or entype β factual | temporal | analytical | synthesis | unanswerabledifficulty β easy | medium | hardsource_article_id β links back to the phase_2 source articlearticle_title β title of the source article