aisingapore/SEA-PILE-v1

Dataset

18

stars

13

commits

1

linked in READMEs

Aug 12, 2026

updated

README

SEA-LION-Pile

SEA-LION-Pile is the pretraining data set for SEA-LION, a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. This repository contains the cleaned mC4 portion of the SEA-LION-Pile.

For the remainder of the SEA-LION-Pile dataset, they may be downloaded from the links provided below.

Dataset Details

SEA-LION was trained on 980B tokens of the following data:

Data SourceUnique TokensMultiplierTotal TokensPercentage
RefinedWeb - English571.3B1571.3B58.20%
mC4 - Chinese91.2B191.2B9.29%
mC4 - Indonesian3.68B414.7B1.50%
mC4 - Malay0.72B42.9B0.29%
mC4 - Filipino1.32B45.3B0.54%
mC4 - Burmese1.2B44.9B0.49%
mC4 - Vietnamese63.4B163.4B6.46%
mC4 - Thai5.8B211.6B1.18%
WangChanBERTa - Thai5B210B1.02%
mC4 - Lao0.27B41.1B0.12%
mC4 - Khmer0.97B43.9B0.40%
mC4 - Tamil2.55B410.2B1.04%
the Stack - Python20.9B241.8B4.26%
the Stack - Javascript55.6B155.6B5.66%
the Stack - Shell1.25B22.5B0.26%
the Stack - SQL6.4B212.8B1.31%
the Stack - Markdown26.6B126.6B2.71%
RedPajama - StackExchange21.2B121.2B2.16%
RedPajama - ArXiv30.6B130.6B3.12%

Additional SEA-LION-Pile (non-mC4) Data Sources

This section contains the links to the additional datasets that form the SEA-LION-Pile.

Limitations

  • As toxic or biased data is prevalent on the internet, it is likely our dataset contains such content.
  • Despite our best efforts to filter content that does not qualify as natural language, and to deduplicate documents, our pipeline may let through documents that may be considered as errors or redundant.
  • Note: The public aisingapore/SEA-PILE-v1 repository hosts the open-source mC4 language subset. The dataset table in this model card describes the complete 980B constructed pre-training mixture used during training (which includes RefinedWeb, mC4, WangChanBERTa, Stack, and RedPajama). Displayed sub-row sums (981.6B) reflect exact unrounded component counts.

License

This public extract of mC4 is made available under ODC-By 1.0 license; users should also abide to the CommonCrawl ToU.

For all other licenses, please refer to their individual pages above.

We endeavor to ensure data used is permissible and have chosen datasets from creators who have processes to exclude copyrighted or disputed data. For other new data, we have obtained permission to use and distribute.

References

@misc{lowphansirikul2021wangchanberta,
    title={WangchanBERTa: Pretraining transformer-based Thai Language Models},
    author={Lalita Lowphansirikul and Charin Polpanumas and Nawat Jantrakulchai and Sarana Nutanong},
    year={2021},
    eprint={2101.09635},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

@article{refinedweb,
  title={The {R}efined{W}eb dataset for {F}alcon {LLM}: outperforming curated corpora with web data, and web data only},
  author={Guilherme Penedo and Quentin Malartic and Daniel Hesslow and Ruxandra Cojocaru and Alessandro Cappelli and Hamza Alobeidli and Baptiste Pannier and Ebtesam Almazrouei and Julien Launay},
  journal={arXiv preprint arXiv:2306.01116},
  eprint={2306.01116},
  eprinttype = {arXiv},
  url={https://arxiv.org/abs/2306.01116},
  year={2023}
}

@article{Kocetkov2022TheStack,
  title={The Stack: 3 TB of permissively licensed source code},
  author={Kocetkov, Denis and Li, Raymond and Ben Allal, Loubna and Li, Jia and Mou,Chenghao and Muñoz Ferrandis, Carlos and Jernite, Yacine and Mitchell, Margaret and Hughes, Sean and Wolf, Thomas and Bahdanau, Dzmitry and von Werra, Leandro and de Vries, Harm},
  journal={Preprint},
  year={2022}
}

@software{together2023redpajama,
  author = {Together Computer},
  title = {RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset},
  month = April,
  year = 2023,
  url = {https://github.com/togethercomputer/RedPajama-Data}
}

Contributors

RaymondAISG

6 commits

SAnocha

4 commits

hamsarajan

2 commits

dotw

1 commits

aisingapore/SEA-PILE-v1

Dataset

18

stars

13

commits

1

linked in READMEs

Aug 12, 2026

updated

README

SEA-LION-Pile

SEA-LION-Pile is the pretraining data set for SEA-LION, a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. This repository contains the cleaned mC4 portion of the SEA-LION-Pile.

For the remainder of the SEA-LION-Pile dataset, they may be downloaded from the links provided below.

Dataset Details

SEA-LION was trained on 980B tokens of the following data:

Data SourceUnique TokensMultiplierTotal TokensPercentage
RefinedWeb - English571.3B1571.3B58.20%
mC4 - Chinese91.2B191.2B9.29%
mC4 - Indonesian3.68B414.7B1.50%
mC4 - Malay0.72B42.9B0.29%
mC4 - Filipino1.32B45.3B0.54%
mC4 - Burmese1.2B44.9B0.49%
mC4 - Vietnamese63.4B163.4B6.46%
mC4 - Thai5.8B211.6B1.18%
WangChanBERTa - Thai5B210B1.02%
mC4 - Lao0.27B41.1B0.12%
mC4 - Khmer0.97B43.9B0.40%
mC4 - Tamil2.55B410.2B1.04%
the Stack - Python20.9B241.8B4.26%
the Stack - Javascript55.6B155.6B5.66%
the Stack - Shell1.25B22.5B0.26%
the Stack - SQL6.4B212.8B1.31%
the Stack - Markdown26.6B126.6B2.71%
RedPajama - StackExchange21.2B121.2B2.16%
RedPajama - ArXiv30.6B130.6B3.12%

Additional SEA-LION-Pile (non-mC4) Data Sources

This section contains the links to the additional datasets that form the SEA-LION-Pile.

Limitations

  • As toxic or biased data is prevalent on the internet, it is likely our dataset contains such content.
  • Despite our best efforts to filter content that does not qualify as natural language, and to deduplicate documents, our pipeline may let through documents that may be considered as errors or redundant.
  • Note: The public aisingapore/SEA-PILE-v1 repository hosts the open-source mC4 language subset. The dataset table in this model card describes the complete 980B constructed pre-training mixture used during training (which includes RefinedWeb, mC4, WangChanBERTa, Stack, and RedPajama). Displayed sub-row sums (981.6B) reflect exact unrounded component counts.

License

This public extract of mC4 is made available under ODC-By 1.0 license; users should also abide to the CommonCrawl ToU.

For all other licenses, please refer to their individual pages above.

We endeavor to ensure data used is permissible and have chosen datasets from creators who have processes to exclude copyrighted or disputed data. For other new data, we have obtained permission to use and distribute.

References

@misc{lowphansirikul2021wangchanberta,
    title={WangchanBERTa: Pretraining transformer-based Thai Language Models},
    author={Lalita Lowphansirikul and Charin Polpanumas and Nawat Jantrakulchai and Sarana Nutanong},
    year={2021},
    eprint={2101.09635},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

@article{refinedweb,
  title={The {R}efined{W}eb dataset for {F}alcon {LLM}: outperforming curated corpora with web data, and web data only},
  author={Guilherme Penedo and Quentin Malartic and Daniel Hesslow and Ruxandra Cojocaru and Alessandro Cappelli and Hamza Alobeidli and Baptiste Pannier and Ebtesam Almazrouei and Julien Launay},
  journal={arXiv preprint arXiv:2306.01116},
  eprint={2306.01116},
  eprinttype = {arXiv},
  url={https://arxiv.org/abs/2306.01116},
  year={2023}
}

@article{Kocetkov2022TheStack,
  title={The Stack: 3 TB of permissively licensed source code},
  author={Kocetkov, Denis and Li, Raymond and Ben Allal, Loubna and Li, Jia and Mou,Chenghao and Muñoz Ferrandis, Carlos and Jernite, Yacine and Mitchell, Margaret and Hughes, Sean and Wolf, Thomas and Bahdanau, Dzmitry and von Werra, Leandro and de Vries, Harm},
  journal={Preprint},
  year={2022}
}

@software{together2023redpajama,
  author = {Together Computer},
  title = {RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset},
  month = April,
  year = 2023,
  url = {https://github.com/togethercomputer/RedPajama-Data}
}

Contributors

RaymondAISG

6 commits

SAnocha

4 commits

hamsarajan

2 commits

dotw

1 commits