The following data mix was used to train K2 and achieve results in line with Llama 2 70B.
K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below.
| Dataset | Starting Tokens | Multiplier | Total Tokens | % of Total |
|---|---|---|---|---|
| dm-math | 4.33B | 3x | 13B | 1% |
| pubmed-abstracts (from the Pile) | 4.77B | 3x | 14.3B | 1.1% |
| uspto (from the Pile) | 4.77B | 3x | 14.3B | 1.1% |
| pubmed-central (from the Pile) | 26B | 1x | 26B | 2% |
| redpajama.arxiv | 27.3B | 1x | 27.3B | 2.1% |
| starcoder.spm | 67.6B | 0.5x | 33.8B | 2.6% |
| starcoder.fim | 67.6B | 0.5x | 33.8B | 2.6% |
| redpajama.stackexchange | 61.1B | 1x | 61.1B | 4.7% |
| starcoder | 132.6B | 0.5x | 66.3B | 5.1% |
| pile-of-law | 76.7B | 1x | 76.7B | 5.9% |
| redpajama.book | 80.6B | 1x | 80.6B | 6.2% |
| s2orc | 107.9B | 1x | 107.9B | 8.3% |
| redpajama.wikipedia | 22.1B | 6x | 132.6B | 10.2% |
| refinedweb | 612.3B | 1x | 612.3B | 47.1% |
| Totals | - | - | 1.3T | 100% |
| Dataset | Starting Tokens | Multiplier | Total Tokens | % of Total |
|---|---|---|---|---|
| open-web-math | 14.6B | 1x | 14.6B | 21% |
| redpajama.arxiv | 2B | 1x | 2B | 2.9% |
| simple-wiki | 4.3B | 1x | 4.3B | 6.2% |
| redpajama.book | 2B | 1x | 2B | 2.9% |
| algebraic-stack | 10.9B | 1x | 10.9B | 15.7% |
| pile-of-law | 2B | 0.5x | 33.8B | 2.9% |
| books | 5.8B | 1x | 5.8B | 8.3% |
| pes20 | 1.2B | 1x | 1.2B | 1.8% |
| pubmed-central (from the Pile) | 2B | 1x | 2B | 2.9% |
| redpajama.wikipedia | 2B | 1x | 2B | 2.9% |
| python | 20.5B | 1x | 20.5B | 29.6% |
| s2orc | 2B | 1x | 2B | 2.9% |
| Totals | - | - | 69.4B* | 100% |
| *rounding |
A step-by-step tutorial for reproducing the K2's data preperation can be found in the LLM360 Pretraining Suite here
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
BibTeX:
@misc{
title={LLM360 K2-65B: Scaling Up Open and Transparent Language Models},
author={The LLM360 Team},
year={2024},
}
The following data mix was used to train K2 and achieve results in line with Llama 2 70B.
K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below.
| Dataset | Starting Tokens | Multiplier | Total Tokens | % of Total |
|---|---|---|---|---|
| dm-math | 4.33B | 3x | 13B | 1% |
| pubmed-abstracts (from the Pile) | 4.77B | 3x | 14.3B | 1.1% |
| uspto (from the Pile) | 4.77B | 3x | 14.3B | 1.1% |
| pubmed-central (from the Pile) | 26B | 1x | 26B | 2% |
| redpajama.arxiv | 27.3B | 1x | 27.3B | 2.1% |
| starcoder.spm | 67.6B | 0.5x | 33.8B | 2.6% |
| starcoder.fim | 67.6B | 0.5x | 33.8B | 2.6% |
| redpajama.stackexchange | 61.1B | 1x | 61.1B | 4.7% |
| starcoder | 132.6B | 0.5x | 66.3B | 5.1% |
| pile-of-law | 76.7B | 1x | 76.7B | 5.9% |
| redpajama.book | 80.6B | 1x | 80.6B | 6.2% |
| s2orc | 107.9B | 1x | 107.9B | 8.3% |
| redpajama.wikipedia | 22.1B | 6x | 132.6B | 10.2% |
| refinedweb | 612.3B | 1x | 612.3B | 47.1% |
| Totals | - | - | 1.3T | 100% |
| Dataset | Starting Tokens | Multiplier | Total Tokens | % of Total |
|---|---|---|---|---|
| open-web-math | 14.6B | 1x | 14.6B | 21% |
| redpajama.arxiv | 2B | 1x | 2B | 2.9% |
| simple-wiki | 4.3B | 1x | 4.3B | 6.2% |
| redpajama.book | 2B | 1x | 2B | 2.9% |
| algebraic-stack | 10.9B | 1x | 10.9B | 15.7% |
| pile-of-law | 2B | 0.5x | 33.8B | 2.9% |
| books | 5.8B | 1x | 5.8B | 8.3% |
| pes20 | 1.2B | 1x | 1.2B | 1.8% |
| pubmed-central (from the Pile) | 2B | 1x | 2B | 2.9% |
| redpajama.wikipedia | 2B | 1x | 2B | 2.9% |
| python | 20.5B | 1x | 20.5B | 29.6% |
| s2orc | 2B | 1x | 2B | 2.9% |
| Totals | - | - | 69.4B* | 100% |
| *rounding |
A step-by-step tutorial for reproducing the K2's data preperation can be found in the LLM360 Pretraining Suite here
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
BibTeX:
@misc{
title={LLM360 K2-65B: Scaling Up Open and Transparent Language Models},
author={The LLM360 Team},
year={2024},
}