zjunlp/DataMind

[ICLR/AAAI/KDD/EMNLP2026] Open-Source LLM-Based Data Analysis Agents

137

stars

60

commits

Python

primary language

Sep 6, 2026

updated

zjunlp.github.io/DataMind/
agent
artificial-intelligence
data-analysis
datamind
dataprm
data-science
large-language-models
longds-bench
natural-language-processing
reinforcement-learning

README

DataMind

๐Ÿ“„arXiv โ€ข ๐Ÿค—HuggingFace

Awesome License

โญ If you like our project, please give us a star on GitHub for the latest updates!

DataMind develops open-source LLM-based data analysis agents from multiple perspectives, including empirical diagnosis, generalist agent scaling, process-level supervision, long-horizon evaluation, and unsupervised skill discovery. Together, these works provide a systematic path toward more capable, scalable, and reliable data-analytic agents.

๐Ÿ“ข News

๐Ÿงญ Project Navigation

This repository hosts multiple data analysis projects. The table below provides an overview and links to each project's documentation:

ProjectDescriptionPaperDocumentation
DataMind-AnalysisEmpirical diagnosis and targeted training for understanding why open-source LLMs struggle with data analysisAAAI 2026DataMind-Analysis.md
DataMindScalable data synthesis and agent training recipe for building generalist data-analytic agentsICLR 2026DataMind.md
DataPRMEnvironment-aware process reward model for reliable multi-step data analysisKDD 2026DataPRM.md
LongDS-BenchLong-horizon benchmark for evaluating analytical state management in multi-turn data analysisEMNLP 2026LongDS-Bench.md
DataCOPEUnsupervised verifier-guided skill discovery framework for data-analytic agents. Meanwhile, we also provide a User-Friendly Framework for generating data analysis skills.arXivDataCOPE.md

๐ŸŽ‰Contributors

We deeply appreciate the collaborative efforts of everyone involved. We will continue to enhance and maintain this repository over the long term. If you encounter any issues, feel free to submit them to us!

โœ๏ธ Citation

If you find our work helpful, please use the following citations.

@misc{qiu2026unsupervisedskilldiscoveryagentic,
      title={Unsupervised Skill Discovery for Agentic Data Analysis}, 
      author={Zhisong Qiu and Kangqi Song and Shengwei Tang and Shuofei Qiao and Lei Liang and Huajun Chen and Shumin Deng},
      year={2026},
      eprint={2606.06416},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2606.06416}, 
}

@misc{xu2026longdsbench,
      title={LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis}, 
      author={Kewei Xu and Xiaoben Lu and Shuofei Qiao and Zihan Ding and Haoming Xu and Lei Liang and Ningyu Zhang},
      year={2026},
      eprint={2605.30434},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2605.30434}, 
}

@article{qiu2026rewarding,
  title={Rewarding the scientific process: Process-level reward modeling for agentic data analysis},
  author={Qiu, Zhisong and Qiao, Shuofei and Xu, Kewei and Zhu, Yuqi and Du, Lun and Zhang, Ningyu and Chen, Huajun},
  journal={arXiv preprint arXiv:2604.24198},
  year={2026}
}

@article{qiao2025scaling,
  title={Scaling Generalist Data-Analytic Agents},
  author={Qiao, Shuofei and Zhao, Yanqiu and Qiu, Zhisong and Wang, Xiaobin and Zhang, Jintian and Bin, Zhao and Zhang, Ningyu and Jiang, Yong and Xie, Pengjun and Huang, Fei and others},
  journal={arXiv preprint arXiv:2509.25084},
  year={2025}
}

@article{zhu2025open,
  title={Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study},
  author={Zhu, Yuqi and Zhong, Yi and Zhang, Jintian and Zhang, Ziheng and Qiao, Shuofei and Luo, Yujie and Du, Lun and Zheng, Da and Chen, Huajun and Zhang, Ningyu},
  journal={arXiv preprint arXiv:2506.19794},
  year={2025}
}

Contributors

xkw666

29 commits

consultantQ

21 commits

healer-666

4 commits

LesileZ

3 commits

zjunlp/DataMind

[ICLR/AAAI/KDD/EMNLP2026] Open-Source LLM-Based Data Analysis Agents

137

stars

60

commits

Python

primary language

Sep 6, 2026

updated

zjunlp.github.io/DataMind/
agent
artificial-intelligence
data-analysis
datamind
dataprm
data-science
large-language-models
longds-bench
natural-language-processing
reinforcement-learning

README

DataMind

๐Ÿ“„arXiv โ€ข ๐Ÿค—HuggingFace

Awesome License

โญ If you like our project, please give us a star on GitHub for the latest updates!

DataMind develops open-source LLM-based data analysis agents from multiple perspectives, including empirical diagnosis, generalist agent scaling, process-level supervision, long-horizon evaluation, and unsupervised skill discovery. Together, these works provide a systematic path toward more capable, scalable, and reliable data-analytic agents.

๐Ÿ“ข News

๐Ÿงญ Project Navigation

This repository hosts multiple data analysis projects. The table below provides an overview and links to each project's documentation:

ProjectDescriptionPaperDocumentation
DataMind-AnalysisEmpirical diagnosis and targeted training for understanding why open-source LLMs struggle with data analysisAAAI 2026DataMind-Analysis.md
DataMindScalable data synthesis and agent training recipe for building generalist data-analytic agentsICLR 2026DataMind.md
DataPRMEnvironment-aware process reward model for reliable multi-step data analysisKDD 2026DataPRM.md
LongDS-BenchLong-horizon benchmark for evaluating analytical state management in multi-turn data analysisEMNLP 2026LongDS-Bench.md
DataCOPEUnsupervised verifier-guided skill discovery framework for data-analytic agents. Meanwhile, we also provide a User-Friendly Framework for generating data analysis skills.arXivDataCOPE.md

๐ŸŽ‰Contributors

We deeply appreciate the collaborative efforts of everyone involved. We will continue to enhance and maintain this repository over the long term. If you encounter any issues, feel free to submit them to us!

โœ๏ธ Citation

If you find our work helpful, please use the following citations.

@misc{qiu2026unsupervisedskilldiscoveryagentic,
      title={Unsupervised Skill Discovery for Agentic Data Analysis}, 
      author={Zhisong Qiu and Kangqi Song and Shengwei Tang and Shuofei Qiao and Lei Liang and Huajun Chen and Shumin Deng},
      year={2026},
      eprint={2606.06416},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2606.06416}, 
}

@misc{xu2026longdsbench,
      title={LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis}, 
      author={Kewei Xu and Xiaoben Lu and Shuofei Qiao and Zihan Ding and Haoming Xu and Lei Liang and Ningyu Zhang},
      year={2026},
      eprint={2605.30434},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2605.30434}, 
}

@article{qiu2026rewarding,
  title={Rewarding the scientific process: Process-level reward modeling for agentic data analysis},
  author={Qiu, Zhisong and Qiao, Shuofei and Xu, Kewei and Zhu, Yuqi and Du, Lun and Zhang, Ningyu and Chen, Huajun},
  journal={arXiv preprint arXiv:2604.24198},
  year={2026}
}

@article{qiao2025scaling,
  title={Scaling Generalist Data-Analytic Agents},
  author={Qiao, Shuofei and Zhao, Yanqiu and Qiu, Zhisong and Wang, Xiaobin and Zhang, Jintian and Bin, Zhao and Zhang, Ningyu and Jiang, Yong and Xie, Pengjun and Huang, Fei and others},
  journal={arXiv preprint arXiv:2509.25084},
  year={2025}
}

@article{zhu2025open,
  title={Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study},
  author={Zhu, Yuqi and Zhong, Yi and Zhang, Jintian and Zhang, Ziheng and Qiao, Shuofei and Luo, Yujie and Du, Lun and Zheng, Da and Chen, Huajun and Zhang, Ningyu},
  journal={arXiv preprint arXiv:2506.19794},
  year={2025}
}

Contributors

xkw666

29 commits

consultantQ

21 commits

healer-666

4 commits

LesileZ

3 commits

Languages

Python

94.0%

Shell

4.8%