[NeurIPS 2026] DataFlex: A Unified Benchmark and Evaluation Platform for Data-Centric Training of Large Language Models
See the code
Data Select · Mix · Reweight — Right in the LLM Training Loop
DataFlex is an advanced dynamic training framework built on top of LLaMA-Factory.
It intelligently schedules training data during optimization and integrates several difficult-to-reproduce repositories into a unified framework. The system provides reproducible implementations of Data Selection, Data Mixture, and Data Reweighting, thereby improving both experimental reproducibility and final model performance.
DataFlex integrates seamlessly with LLaMA-Factory, offering researchers and developers more flexible and powerful training control. For goals and design philosophy, please refer to DataFlex-Doc. We summarize repositories related to Data Selection, Data Mixture, and Data Reweighting. ❌ indicates that no official repository is available; ✅ indicates that an official repository is available; ⚠️ indicates that an official repository exists but contains issues.
| Method | Category | Requires Model-in-the-Loop? | Official Repo |
|---|---|---|---|
| LESS | Gradient-Based | ✅ Yes | ⚠️official code |
| NICE | Gradient-Based | ✅ Yes | ⚠️official code |
| Loss | Loss-Based | ✅ Yes | ❌ |
| Delta Loss | Loss-Based | ✅ Yes | ❌ |
| NEAR | Data Distribution-Based | ❌ No | ❌ |
| TSDS | Data Distribution-Based | ❌ No | ✅official code |
| Static | No Selection | ❌ No | ❌ |
| Random | Random Sampling | ❌ No | ❌ |
| Method | Category | Requires Model-in-the-Loop? | Official Repo |
|---|---|---|---|
| DOREMI | Offline Mixture | ✅ Yes | ⚠️official code |
| ODM | Online Mixture | ✅ Yes | ⚠️official code |
| Method | Category | Requires Model-in-the-Loop? | Official Repo |
|---|---|---|---|
| Loss Reweighting | Loss-Based | ✅ Yes | ❌ |
| Joint-Update-Aware Reweighting | Batch-Aware | ✅ Yes | ❌ |
Please use the following commands for environment setup and installation👇
Install torch first, so pip does not resolve a different build and then have to replace it:
pip install --index-url https://download.pytorch.org/whl/cu124 \
torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
Then install DataFlex:
pip install dataflex
Or install from source for development:
git clone https://github.com/OpenDCAI/DataFlex.git
cd DataFlex
pip install -e .
Every config under examples/deepspeed needs DeepSpeed, which ships as an optional extra:
pip install "dataflex[deepspeed]" # from source: pip install -e ".[deepspeed]"
Note: Requires Python 3.11+ and LlamaFactory 0.9.5+, installed automatically along with the other core dependencies. Works on transformers 4.55 through 5.6; the newer model families (Qwen3.5, Gemma 4) need transformers 5.5+. We recommend transformers 5.3+ if you need
train_from_scratchunder DeepSpeed ZeRO-3.The
deepspeedextra stays below 0.17 on purpose: 0.17+ fails to import on the torch pinned above. On a newer torch you are free to lift that cap.The LESS selector needs TRAK, which is another optional extra:
pip install dataflex[less].
The launch command is similar to LLaMA-Factory. Below is an example using LESS :
dataflex-cli train examples/train_lora/selectors/less.yaml
Unlike vanilla LLaMA-Factory, your .yaml config file must also include DataFlex-specific parameters. For details, please refer to DataFlex-Doc.
Using DataFlex can improve performance over the default LLaMA-Factory training.
We use a subset of Open-Hermes-2.5 as the training dataset. The data selection algorithms and data reweighting algorithm outperform the random selector baseline on the MMLU benchmark subset relevant to the training dataset. For the Less and Nice algorithm, we set the validation set as the MMLU-Validation-Set, using a GPT-5-generated trajectory.
We use subsets of SlimPajama-627B for data mixture. The data mixture algorithms outperform the baseline (default data mixture) on MMLU accuracy while also achieving lower perplexity across different data domains.
| Method | Acc ↑ | Perplexity (PPL) ↓ | |||||||
|---|---|---|---|---|---|---|---|---|---|
| MMLU | ALL | CC | C4 | SE | Wiki | GitHub | ArXiv | Book | |
| Slim-Pajama-6B | |||||||||
| Baseline | 25.27 | 4.217 | 4.278 | 4.532 | 3.402 | 3.546 | 2.640 | 3.508 | 4.778 |
| DoReMi | 25.84 | 4.134 | 4.108 | 4.358 | 3.788 | 3.997 | 3.420 | 3.413 | 4.661 |
| ODM | 26.04 | 4.244 | 4.326 | 4.555 | 3.243 | 3.699 | 2.704 | 2.904 | 4.613 |
| Slim-Pajama-30B | |||||||||
| Baseline | 25.51 | 3.584 | 3.723 | 3.505 | 2.850 | 3.215 | 3.163 | 4.540 | 5.329 |
| DoReMi | 25.97 | 3.562 | 3.731 | 3.503 | 2.706 | 2.985 | 2.973 | 4.441 | 5.214 |
| ODM | 25.63 | 3.429 | 3.598 | 3.519 | 2.382 | 2.713 | 2.255 | 3.487 | 4.746 |
DataFlex focuses on data scheduling during training. For a complete pipeline starting from raw data, it pairs well with DataFlow:
DataFlow converts raw files into LLM training data through composable operator pipelines — document parsing, knowledge cleaning, QA / CoT synthesis, and training format conversion. The output JSON can be fed directly into DataFlex. The two projects are independent with no code dependency, connected only by standard data formats. DataFlex accepts training data from any source — DataFlow, manual annotation, HuggingFace datasets, or custom processing scripts.
We thank LLaMA-Factory for offering an efficient and user-friendly framework for large model fine-tuning, which greatly facilitated rapid iteration in our training and experimentation workflows.
We thank Zhongguancun Academy for their API and GPU support.
Our gratitude extends to all contributors in the open-source community—their efforts collectively drive the development of DataFlex.
The DataFlex paper has been accepted to the NeurIPS 2026 Evaluation & Dataset Track. Until the official proceedings citation is available, please cite the arXiv version:
@article{liang2026dataflex,
title={DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models},
author={Liang, Hao and Zhao, Zhengyang and Qiang, Meiyi and Chen, Mingrui and Ma, Lu and Yu, Rongyi and Feng, Hengyi and Sun, Shixuan and Meng, Zimo and Ma, Xiaochen and others},
journal={arXiv preprint arXiv:2603.26164},
year={2026}
}
@article{liang2026towards,
title={Towards Next-Generation LLM Training: From the Data-Centric Perspective},
author={Liang, Hao and Zhao, Zhengyang and Han, Zhaoyang and Qiang, Meiyi and Ma, Xiaochen and Zeng, Bohan and Cai, Qifeng and Li, Zhiyu and Tang, Linpeng and Zhang, Wentao and others},
journal={arXiv preprint arXiv:2603.14712},
year={2026}
}
We welcome contributions of new trainers and selectors! Please ensure code formatting is consistent with the existing style before submitting a PR.
We also welcome you to join the DataFlex and DataFlow open-source community to ask questions, share ideas, and collaborate with other developers!
• 📮 GitHub Issues: Report bugs or suggest features
• 🔧 GitHub Pull Requests: Contribute code improvements
• 💬 Join our community groups to connect with us and other contributors!
26 followers · starred Apr 2026
98 followers · starred Sep 2026
9 followers · starred Apr 2026
[NeurIPS 2026] DataFlex: A Unified Benchmark and Evaluation Platform for Data-Centric Training of Large Language Models
See the code
Data Select · Mix · Reweight — Right in the LLM Training Loop
DataFlex is an advanced dynamic training framework built on top of LLaMA-Factory.
It intelligently schedules training data during optimization and integrates several difficult-to-reproduce repositories into a unified framework. The system provides reproducible implementations of Data Selection, Data Mixture, and Data Reweighting, thereby improving both experimental reproducibility and final model performance.
DataFlex integrates seamlessly with LLaMA-Factory, offering researchers and developers more flexible and powerful training control. For goals and design philosophy, please refer to DataFlex-Doc. We summarize repositories related to Data Selection, Data Mixture, and Data Reweighting. ❌ indicates that no official repository is available; ✅ indicates that an official repository is available; ⚠️ indicates that an official repository exists but contains issues.
| Method | Category | Requires Model-in-the-Loop? | Official Repo |
|---|---|---|---|
| LESS | Gradient-Based | ✅ Yes | ⚠️official code |
| NICE | Gradient-Based | ✅ Yes | ⚠️official code |
| Loss | Loss-Based | ✅ Yes | ❌ |
| Delta Loss | Loss-Based | ✅ Yes | ❌ |
| NEAR | Data Distribution-Based | ❌ No | ❌ |
| TSDS | Data Distribution-Based | ❌ No | ✅official code |
| Static | No Selection | ❌ No | ❌ |
| Random | Random Sampling | ❌ No | ❌ |
| Method | Category | Requires Model-in-the-Loop? | Official Repo |
|---|---|---|---|
| DOREMI | Offline Mixture | ✅ Yes | ⚠️official code |
| ODM | Online Mixture | ✅ Yes | ⚠️official code |
| Method | Category | Requires Model-in-the-Loop? | Official Repo |
|---|---|---|---|
| Loss Reweighting | Loss-Based | ✅ Yes | ❌ |
| Joint-Update-Aware Reweighting | Batch-Aware | ✅ Yes | ❌ |
Please use the following commands for environment setup and installation👇
Install torch first, so pip does not resolve a different build and then have to replace it:
pip install --index-url https://download.pytorch.org/whl/cu124 \
torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
Then install DataFlex:
pip install dataflex
Or install from source for development:
git clone https://github.com/OpenDCAI/DataFlex.git
cd DataFlex
pip install -e .
Every config under examples/deepspeed needs DeepSpeed, which ships as an optional extra:
pip install "dataflex[deepspeed]" # from source: pip install -e ".[deepspeed]"
Note: Requires Python 3.11+ and LlamaFactory 0.9.5+, installed automatically along with the other core dependencies. Works on transformers 4.55 through 5.6; the newer model families (Qwen3.5, Gemma 4) need transformers 5.5+. We recommend transformers 5.3+ if you need
train_from_scratchunder DeepSpeed ZeRO-3.The
deepspeedextra stays below 0.17 on purpose: 0.17+ fails to import on the torch pinned above. On a newer torch you are free to lift that cap.The LESS selector needs TRAK, which is another optional extra:
pip install dataflex[less].
The launch command is similar to LLaMA-Factory. Below is an example using LESS :
dataflex-cli train examples/train_lora/selectors/less.yaml
Unlike vanilla LLaMA-Factory, your .yaml config file must also include DataFlex-specific parameters. For details, please refer to DataFlex-Doc.
Using DataFlex can improve performance over the default LLaMA-Factory training.
We use a subset of Open-Hermes-2.5 as the training dataset. The data selection algorithms and data reweighting algorithm outperform the random selector baseline on the MMLU benchmark subset relevant to the training dataset. For the Less and Nice algorithm, we set the validation set as the MMLU-Validation-Set, using a GPT-5-generated trajectory.
We use subsets of SlimPajama-627B for data mixture. The data mixture algorithms outperform the baseline (default data mixture) on MMLU accuracy while also achieving lower perplexity across different data domains.
| Method | Acc ↑ | Perplexity (PPL) ↓ | |||||||
|---|---|---|---|---|---|---|---|---|---|
| MMLU | ALL | CC | C4 | SE | Wiki | GitHub | ArXiv | Book | |
| Slim-Pajama-6B | |||||||||
| Baseline | 25.27 | 4.217 | 4.278 | 4.532 | 3.402 | 3.546 | 2.640 | 3.508 | 4.778 |
| DoReMi | 25.84 | 4.134 | 4.108 | 4.358 | 3.788 | 3.997 | 3.420 | 3.413 | 4.661 |
| ODM | 26.04 | 4.244 | 4.326 | 4.555 | 3.243 | 3.699 | 2.704 | 2.904 | 4.613 |
| Slim-Pajama-30B | |||||||||
| Baseline | 25.51 | 3.584 | 3.723 | 3.505 | 2.850 | 3.215 | 3.163 | 4.540 | 5.329 |
| DoReMi | 25.97 | 3.562 | 3.731 | 3.503 | 2.706 | 2.985 | 2.973 | 4.441 | 5.214 |
| ODM | 25.63 | 3.429 | 3.598 | 3.519 | 2.382 | 2.713 | 2.255 | 3.487 | 4.746 |
DataFlex focuses on data scheduling during training. For a complete pipeline starting from raw data, it pairs well with DataFlow:
DataFlow converts raw files into LLM training data through composable operator pipelines — document parsing, knowledge cleaning, QA / CoT synthesis, and training format conversion. The output JSON can be fed directly into DataFlex. The two projects are independent with no code dependency, connected only by standard data formats. DataFlex accepts training data from any source — DataFlow, manual annotation, HuggingFace datasets, or custom processing scripts.
We thank LLaMA-Factory for offering an efficient and user-friendly framework for large model fine-tuning, which greatly facilitated rapid iteration in our training and experimentation workflows.
We thank Zhongguancun Academy for their API and GPU support.
Our gratitude extends to all contributors in the open-source community—their efforts collectively drive the development of DataFlex.
The DataFlex paper has been accepted to the NeurIPS 2026 Evaluation & Dataset Track. Until the official proceedings citation is available, please cite the arXiv version:
@article{liang2026dataflex,
title={DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models},
author={Liang, Hao and Zhao, Zhengyang and Qiang, Meiyi and Chen, Mingrui and Ma, Lu and Yu, Rongyi and Feng, Hengyi and Sun, Shixuan and Meng, Zimo and Ma, Xiaochen and others},
journal={arXiv preprint arXiv:2603.26164},
year={2026}
}
@article{liang2026towards,
title={Towards Next-Generation LLM Training: From the Data-Centric Perspective},
author={Liang, Hao and Zhao, Zhengyang and Han, Zhaoyang and Qiang, Meiyi and Ma, Xiaochen and Zeng, Bohan and Cai, Qifeng and Li, Zhiyu and Tang, Linpeng and Zhang, Wentao and others},
journal={arXiv preprint arXiv:2603.14712},
year={2026}
}
We welcome contributions of new trainers and selectors! Please ensure code formatting is consistent with the existing style before submitting a PR.
We also welcome you to join the DataFlex and DataFlow open-source community to ask questions, share ideas, and collaborate with other developers!
• 📮 GitHub Issues: Report bugs or suggest features
• 🔧 GitHub Pull Requests: Contribute code improvements
• 💬 Join our community groups to connect with us and other contributors!
26 followers · starred Apr 2026
98 followers · starred Sep 2026
9 followers · starred Apr 2026