π€ Hugging Face π€ ModelScope π₯οΈ GitHub
The Ling-Coder Dataset comprises the following components:
This is a portion of the code DPO data used during the training of the Ling-Coder Lite model, consisting of 250K samples and encompassing several popular programming languages.
The dataset was derived from the open-source dataset code_contests by selecting questions and positive/negative samples based on code test case pass rates, PPL distribution, and reward model scoring.
Please consider citing our technique report Ling-Coder-TR if you find this dataset useful:
@misc{codefuse2025samplemattersleveragingmixtureofexperts,
title={Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM},
author={Codefuse and Ling Team},
year={2025},
eprint={2503.17793},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.17793},
}
10 commits
1 commits
π€ Hugging Face π€ ModelScope π₯οΈ GitHub
The Ling-Coder Dataset comprises the following components:
This is a portion of the code DPO data used during the training of the Ling-Coder Lite model, consisting of 250K samples and encompassing several popular programming languages.
The dataset was derived from the open-source dataset code_contests by selecting questions and positive/negative samples based on code test case pass rates, PPL distribution, and reward model scoring.
Please consider citing our technique report Ling-Coder-TR if you find this dataset useful:
@misc{codefuse2025samplemattersleveragingmixtureofexperts,
title={Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM},
author={Codefuse and Ling Team},
year={2025},
eprint={2503.17793},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.17793},
}
10 commits
1 commits