
This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
By Kurakura AI: Dataset Link.
Built from scraped subjects of French Baccalauréat and Preparatory Class (CPGE) entrance exams in mathematics, computer science, and physics.
By AllenAI: Dataset Link.
Translated prompts to French, generated new answers with Qwen3-32B, then filtered the dataset.
By AllenAI: Dataset Link.
Translated prompts to French, generated new answers with Qwen3-32B, then filtered the dataset.
By HuggingFaceTB: Dataset Link.
Extracted French samples only.
By CohereLabs: Dataset Link.
Extracted French samples only.
By legmlai: Dataset Link.
Filtered the dataset.
By Manuel Faysse: Dataset Link.
Extracted French samples only.
@misc{luth2025kurakurai,
title = {Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer},
author = {Lasbordes, Maxence and Gad, Sinoué},
year = {2025},
howpublished = {\url{https://arxiv.org/abs/2510.05846}},
note = {arXiv:2510.05846}
}

This dataset includes all the data used to fine-tune Luth-0.6B-Instruct and Luth-1.7B-Instruct, enhancing their French capabilities on tasks such as instruction following, mathematics, and general knowledge. The models also improved in English thanks to knowledge transfer between the two languages.
It contains ~338M tokens in French. Our data scripts are available on GitHub.
By Kurakura AI: Dataset Link.
Built from scraped subjects of French Baccalauréat and Preparatory Class (CPGE) entrance exams in mathematics, computer science, and physics.
By AllenAI: Dataset Link.
Translated prompts to French, generated new answers with Qwen3-32B, then filtered the dataset.
By AllenAI: Dataset Link.
Translated prompts to French, generated new answers with Qwen3-32B, then filtered the dataset.
By HuggingFaceTB: Dataset Link.
Extracted French samples only.
By CohereLabs: Dataset Link.
Extracted French samples only.
By legmlai: Dataset Link.
Filtered the dataset.
By Manuel Faysse: Dataset Link.
Extracted French samples only.
@misc{luth2025kurakurai,
title = {Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer},
author = {Lasbordes, Maxence and Gad, Sinoué},
year = {2025},
howpublished = {\url{https://arxiv.org/abs/2510.05846}},
note = {arXiv:2510.05846}
}