Source-aware shuffled supervised fine-tuning data formatted for LiquidAI/LFM2.5-1.2B-Instruct.
LiquidAI/LFM2.5-1.2B-Instruct1337token_count includes the LFM BOS token and ChatML turn-end tokens; no extra terminal EOS is appended.chatml: LFM2.5 template text, including <|startoftext|> and ChatML markers.messages: structured messages as [{"content": ..., "role": ...}].token_count: exact LFM token count for the chatml value.source: original dataset identifier; Kyoto-derived rows use their underlying repository name.| Source | Conversations | Tokens | Share of tokens |
|---|---|---|---|
| Yxanul/Mephisto-Knowledge_538k | 465,737 | 252,573,939 | 15.09% |
| HuggingFaceH4/ultrachat_200k | 515,311 | 544,331,752 | 32.52% |
| HuggingFaceTB/smol-smoltalk | 460,341 | 404,763,477 | 24.18% |
| mlabonne/FineTome-100k-dedup | 99,523 | 59,653,187 | 3.56% |
| HuggingFaceTB/smoltalk2 | 640,494 | 193,082,118 | 11.54% |
| allenai/WildChat-1M | 295,413 | 71,554,864 | 4.28% |
| argilla/ifeval-like-data | 20,768 | 4,867,358 | 0.29% |
| cais/mmlu | 85,713 | 29,076,949 | 1.74% |
| openai/gsm8k | 7,473 | 1,281,133 | 0.08% |
| allenai/tulu-3-sft-personas-instruction-following | 20,997 | 5,431,926 | 0.32% |
| allenai/math_qa | 29,837 | 2,581,937 | 0.15% |
| meta-math/MetaMathQA | 372,903 | 82,026,268 | 4.90% |
| WizardLMTeam/WizardLM_evol_instruct_V2_196k | 77,104 | 22,468,345 | 1.34% |
| Total | 3,091,614 | 1,673,693,253 | 100.00% |
Rows tagged as HuggingFaceH4/ultrachat_200k and HuggingFaceTB/smol-smoltalk inside the curated source were excluded because those datasets are included directly.
Additional source contributions came from these public datasets through the Kyoto-Corpus curation:
4 commits
Source-aware shuffled supervised fine-tuning data formatted for LiquidAI/LFM2.5-1.2B-Instruct.
LiquidAI/LFM2.5-1.2B-Instruct1337token_count includes the LFM BOS token and ChatML turn-end tokens; no extra terminal EOS is appended.chatml: LFM2.5 template text, including <|startoftext|> and ChatML markers.messages: structured messages as [{"content": ..., "role": ...}].token_count: exact LFM token count for the chatml value.source: original dataset identifier; Kyoto-derived rows use their underlying repository name.| Source | Conversations | Tokens | Share of tokens |
|---|---|---|---|
| Yxanul/Mephisto-Knowledge_538k | 465,737 | 252,573,939 | 15.09% |
| HuggingFaceH4/ultrachat_200k | 515,311 | 544,331,752 | 32.52% |
| HuggingFaceTB/smol-smoltalk | 460,341 | 404,763,477 | 24.18% |
| mlabonne/FineTome-100k-dedup | 99,523 | 59,653,187 | 3.56% |
| HuggingFaceTB/smoltalk2 | 640,494 | 193,082,118 | 11.54% |
| allenai/WildChat-1M | 295,413 | 71,554,864 | 4.28% |
| argilla/ifeval-like-data | 20,768 | 4,867,358 | 0.29% |
| cais/mmlu | 85,713 | 29,076,949 | 1.74% |
| openai/gsm8k | 7,473 | 1,281,133 | 0.08% |
| allenai/tulu-3-sft-personas-instruction-following | 20,997 | 5,431,926 | 0.32% |
| allenai/math_qa | 29,837 | 2,581,937 | 0.15% |
| meta-math/MetaMathQA | 372,903 | 82,026,268 | 4.90% |
| WizardLMTeam/WizardLM_evol_instruct_V2_196k | 77,104 | 22,468,345 | 1.34% |
| Total | 3,091,614 | 1,673,693,253 | 100.00% |
Rows tagged as HuggingFaceH4/ultrachat_200k and HuggingFaceTB/smol-smoltalk inside the curated source were excluded because those datasets are included directly.
Additional source contributions came from these public datasets through the Kyoto-Corpus curation:
4 commits