🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub
The Ling-Coder Dataset comprises the following components:
This dataset is part of the synthetic data used during the training of the Ling-Coder-lite-base model. It comprises over 24 million English and Chinese samples and covers various task topics such as text2code, text2sql, code execution reasoning, code fixing, and testcase generation.
The dataset was generated using the Magpie method with open-source code LLMs such as Qwen2.5-Coder and DeepSeek-Coder-v2. During the synthesis process, we configured multiple sampling strategies, such as varying temperatures, and set multiple task themes. The synthesized data was subjected to a pipeline of rule-based cleansing, detoxification, deduplication, quality inspection, and ablation testing to meet our admission criteria. For more detailed information on the construction process, please refer to the Ling-Coder Lite technique report.
Please consider citing our technique report Ling-Coder-TR if you find this dataset useful:
@misc{codefuse2025samplemattersleveragingmixtureofexperts,
title={Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM},
author={Codefuse and Ling Team},
year={2025},
eprint={2503.17793},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.17793},
}
19 commits
1 commits
🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub
The Ling-Coder Dataset comprises the following components:
This dataset is part of the synthetic data used during the training of the Ling-Coder-lite-base model. It comprises over 24 million English and Chinese samples and covers various task topics such as text2code, text2sql, code execution reasoning, code fixing, and testcase generation.
The dataset was generated using the Magpie method with open-source code LLMs such as Qwen2.5-Coder and DeepSeek-Coder-v2. During the synthesis process, we configured multiple sampling strategies, such as varying temperatures, and set multiple task themes. The synthesized data was subjected to a pipeline of rule-based cleansing, detoxification, deduplication, quality inspection, and ablation testing to meet our admission criteria. For more detailed information on the construction process, please refer to the Ling-Coder Lite technique report.
Please consider citing our technique report Ling-Coder-TR if you find this dataset useful:
@misc{codefuse2025samplemattersleveragingmixtureofexperts,
title={Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM},
author={Codefuse and Ling Team},
year={2025},
eprint={2503.17793},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.17793},
}
19 commits
1 commits