SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance.
SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise.
SYNTH differs from existing open synthetic dataset in being:
SYNTH is not only the name of a dataset but an initiative for open synthetic data and open environment led by AI Alliance and Pleias that aims to address the critical gap in open-source AI development by creating a cutting-edge, open-source data corpus for training sovereign AI models and advanced AI agents.
At its core, SYNTH is a fully synthetic and engineered corpus derived from a sample of 50,000 pages curated by the Wikipedia community. Throughout the past two decades, thousands of contributors selected a collection of core topics that every encyclopedia should have, Wikipedia:Vital articles. It’s a concentric selection starting at level 1 (10 articles) up to level 5 (50,000 articles). SYNTH includes as its starting point all articles featured in level 5.
SYNTH further expands on this core nucleus with three additional seed collections:
This content act as the SYNTH memory base and has been amplified at a minimum 100 times (about 10000 times for recent/self knowledge). Our amplification strategy relies on a new synthetic pipeline, partly inspired by RAG applications:
The approach has been originally explored by Pleias for retrieval-augmented generation. It has been extended to virtually most of the expected use case of small reasoning models:
While the final training data is fully synthetic, it relied on seeds collected from three data sources:
The dataset aims to support data efficient training of small reasoning model. It provide a generalist, self-sufficient collection of multilingual amplified encyclopedic texts along with synthetic reasoning traces, as well as synthetic tasks that reinforce most of the expected capacities of small model.
In contrast with organic pretraining dataset, SYNTH allows for fast convergence to the existing SOTA (about 100 billion tokens). Furthermore, SYNTH is fully releasable, only use sourced text under free license.
Overall, SYNTH aims to support an emerging ecosystem of small training model by providing a reusable generalist foundational dataset.
Direct use include:
Current out-of-scope use include:
Yet, SYNTH is a live resources and we intend to cover some of these use cases in future releases.
| Field | Type | Description |
|---|---|---|
| synth_id | string | Unique synthetic identifier for each generated sample. |
| language | string | Language of the text sample (e.g., "en", "fr", "it", "es", "de", "pl", "nl", "la"). |
| exercise | string | Type of synthetic exercise (e.g., reasoning, writing, retrieval, arithmetic). Describes the synthetic task context. |
| model | string | Finetuned model used to generate the synthetic sample |
| query | string | Backtranslated query. |
| query_seed_url | string | URL of the Wikipedia or Wikibooks section that served as the seed for query generation. |
| query_seed_text | string | Extend text used as seed for query generation. |
| additional_seed_url | string | Optional additional URL(s) used as supplementary seed |
| seed_license | string | License of the seed text (most of the time "CC-BY-SA 4.0"). |
| constraints | string | Generation constraints applied to answer generation. Varies depending on the exercise |
| script | string | Internal template or script identifier defining the structure of the synthetic exercise. |
| synthetic_reasoning | string | Generated reasoning draft. |
| synthetic_answer | string | Final generated answer or output corresponding to the query. |
| words | int64 | Word count of the full generated text sample (query + draft + answer) |
SYNTH is structured around a “memory core”, the Wikipedia vital articles.. Throughout the past two decades, thousands of contributors selected a collection of core topics that every encyclopedia should have: it’s a concentric selection starting at level 1 (10 articles) up to level 5 (50,000 articles). SYNTH includes as its starting point all articles featured in level 5. It further expands on this selection by increasing coverage of more specialized domains (physics, chemistry, law…) through targeted expansion of wikidata knowledge graphs.
The 58,698 Wikipedia articles were collected thanks to ''Structured Wikipedia'', a project from Wikimedia Enterprise that parsed directly rendered Wikipedia articles in html. Structured Wikipedia fixed most of the formatting issues linked with the mediawiki syntax and provides a clean, section-based version of all Wikipedia pages.
We additionally extracted 3,000 cooking recipes from Wikibooks using the standard API method from Wikimedia.
The main sourced dataset used for synthetic amplification was curated by the English Wikipedia communities throughout nearly 2 decades. Rationale for selection are available on the relevant talk pages of Wikipedia:Vital articles.
The selection reflect similar bias for "canon" general knowledge in English-speaking countries than major LLM benchmarks like MMLU (drawn from high school exams).
The dataset only contain encyclopedic information on highly well-known historical people. No PII curation was needed.
The dataset was created from a collection of 50,000 Wikipedia articles curated by the community (Wikipedia:Vital Articles).
On top of the well documented structural bias in Wikipedia contribution and editing, the selection has been intently made from the perspective of western US/European culture.
Due to systematic Wikipedia grounding, the data presents a very low risk of toxic or problematic content, as well as poor or highly hallucinated information.
SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance.
SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise.
SYNTH differs from existing open synthetic dataset in being:
SYNTH is not only the name of a dataset but an initiative for open synthetic data and open environment led by AI Alliance and Pleias that aims to address the critical gap in open-source AI development by creating a cutting-edge, open-source data corpus for training sovereign AI models and advanced AI agents.
At its core, SYNTH is a fully synthetic and engineered corpus derived from a sample of 50,000 pages curated by the Wikipedia community. Throughout the past two decades, thousands of contributors selected a collection of core topics that every encyclopedia should have, Wikipedia:Vital articles. It’s a concentric selection starting at level 1 (10 articles) up to level 5 (50,000 articles). SYNTH includes as its starting point all articles featured in level 5.
SYNTH further expands on this core nucleus with three additional seed collections:
This content act as the SYNTH memory base and has been amplified at a minimum 100 times (about 10000 times for recent/self knowledge). Our amplification strategy relies on a new synthetic pipeline, partly inspired by RAG applications:
The approach has been originally explored by Pleias for retrieval-augmented generation. It has been extended to virtually most of the expected use case of small reasoning models:
While the final training data is fully synthetic, it relied on seeds collected from three data sources:
The dataset aims to support data efficient training of small reasoning model. It provide a generalist, self-sufficient collection of multilingual amplified encyclopedic texts along with synthetic reasoning traces, as well as synthetic tasks that reinforce most of the expected capacities of small model.
In contrast with organic pretraining dataset, SYNTH allows for fast convergence to the existing SOTA (about 100 billion tokens). Furthermore, SYNTH is fully releasable, only use sourced text under free license.
Overall, SYNTH aims to support an emerging ecosystem of small training model by providing a reusable generalist foundational dataset.
Direct use include:
Current out-of-scope use include:
Yet, SYNTH is a live resources and we intend to cover some of these use cases in future releases.
| Field | Type | Description |
|---|---|---|
| synth_id | string | Unique synthetic identifier for each generated sample. |
| language | string | Language of the text sample (e.g., "en", "fr", "it", "es", "de", "pl", "nl", "la"). |
| exercise | string | Type of synthetic exercise (e.g., reasoning, writing, retrieval, arithmetic). Describes the synthetic task context. |
| model | string | Finetuned model used to generate the synthetic sample |
| query | string | Backtranslated query. |
| query_seed_url | string | URL of the Wikipedia or Wikibooks section that served as the seed for query generation. |
| query_seed_text | string | Extend text used as seed for query generation. |
| additional_seed_url | string | Optional additional URL(s) used as supplementary seed |
| seed_license | string | License of the seed text (most of the time "CC-BY-SA 4.0"). |
| constraints | string | Generation constraints applied to answer generation. Varies depending on the exercise |
| script | string | Internal template or script identifier defining the structure of the synthetic exercise. |
| synthetic_reasoning | string | Generated reasoning draft. |
| synthetic_answer | string | Final generated answer or output corresponding to the query. |
| words | int64 | Word count of the full generated text sample (query + draft + answer) |
SYNTH is structured around a “memory core”, the Wikipedia vital articles.. Throughout the past two decades, thousands of contributors selected a collection of core topics that every encyclopedia should have: it’s a concentric selection starting at level 1 (10 articles) up to level 5 (50,000 articles). SYNTH includes as its starting point all articles featured in level 5. It further expands on this selection by increasing coverage of more specialized domains (physics, chemistry, law…) through targeted expansion of wikidata knowledge graphs.
The 58,698 Wikipedia articles were collected thanks to ''Structured Wikipedia'', a project from Wikimedia Enterprise that parsed directly rendered Wikipedia articles in html. Structured Wikipedia fixed most of the formatting issues linked with the mediawiki syntax and provides a clean, section-based version of all Wikipedia pages.
We additionally extracted 3,000 cooking recipes from Wikibooks using the standard API method from Wikimedia.
The main sourced dataset used for synthetic amplification was curated by the English Wikipedia communities throughout nearly 2 decades. Rationale for selection are available on the relevant talk pages of Wikipedia:Vital articles.
The selection reflect similar bias for "canon" general knowledge in English-speaking countries than major LLM benchmarks like MMLU (drawn from high school exams).
The dataset only contain encyclopedic information on highly well-known historical people. No PII curation was needed.
The dataset was created from a collection of 50,000 Wikipedia articles curated by the community (Wikipedia:Vital Articles).
On top of the well documented structural bias in Wikipedia contribution and editing, the selection has been intently made from the perspective of western US/European culture.
Due to systematic Wikipedia grounding, the data presents a very low risk of toxic or problematic content, as well as poor or highly hallucinated information.