gair-prox/DCLM-pro

Dataset

13

stars

15

commits

2

linked in READMEs

Feb 15, 2025

updated

common crawl
web
Browse cluster: LLM Pretraining Datasets & Corpora β†’

README

πŸ“š DCLM-pro

ArXiv | Models | Code

DCLM-pro is refined from DCLM using the ProX refining framework. It contains about >500B high quality tokens, ready for general language model pre-training.

License

DCLM-pro is based on DCLM, which is made available under an cc-by-4.0 license.

Citation

@article{zhou2024programming,
  title={Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale},
  author={Zhou, Fan and Wang, Zengzhi and Liu, Qian and Li, Junlong and Liu, Pengfei},
  journal={arXiv preprint arXiv:2409.17115},
  year={2024}
}

Contributors

koalazf99

12 commits

SinclairWang

3 commits

gair-prox/DCLM-pro

Dataset

13

stars

15

commits

2

linked in READMEs

Feb 15, 2025

updated

common crawl
web
Browse cluster: LLM Pretraining Datasets & Corpora β†’

README

πŸ“š DCLM-pro

ArXiv | Models | Code

DCLM-pro is refined from DCLM using the ProX refining framework. It contains about >500B high quality tokens, ready for general language model pre-training.

License

DCLM-pro is based on DCLM, which is made available under an cc-by-4.0 license.

Citation

@article{zhou2024programming,
  title={Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale},
  author={Zhou, Fan and Wang, Zengzhi and Liu, Qian and Li, Junlong and Liu, Pengfei},
  journal={arXiv preprint arXiv:2409.17115},
  year={2024}
}

Contributors

koalazf99

12 commits

SinclairWang

3 commits