AxiomicLabs/SFTset-SLM

Dataset

SFTset-SLM

7

stars

4

commits

Sep 10, 2026

updated

sft
slm

README

SFTset-SLM

Source-aware shuffled supervised fine-tuning data formatted for LiquidAI/LFM2.5-1.2B-Instruct.

Dataset summary

  • Conversations: 3,091,614
  • Tokens: 1,673,693,253
  • Parquet parts: 11
  • Target Parquet file size: 500 MiB
  • Tokenizer: LiquidAI/LFM2.5-1.2B-Instruct
  • Shuffle seed: 1337
  • token_count includes the LFM BOS token and ChatML turn-end tokens; no extra terminal EOS is appended.

Columns

  • chatml: LFM2.5 template text, including <|startoftext|> and ChatML markers.
  • messages: structured messages as [{"content": ..., "role": ...}].
  • token_count: exact LFM token count for the chatml value.
  • source: original dataset identifier; Kyoto-derived rows use their underlying repository name.

Source breakdown

SourceConversationsTokensShare of tokens
Yxanul/Mephisto-Knowledge_538k465,737252,573,93915.09%
HuggingFaceH4/ultrachat_200k515,311544,331,75232.52%
HuggingFaceTB/smol-smoltalk460,341404,763,47724.18%
mlabonne/FineTome-100k-dedup99,52359,653,1873.56%
HuggingFaceTB/smoltalk2640,494193,082,11811.54%
allenai/WildChat-1M295,41371,554,8644.28%
argilla/ifeval-like-data20,7684,867,3580.29%
cais/mmlu85,71329,076,9491.74%
openai/gsm8k7,4731,281,1330.08%
allenai/tulu-3-sft-personas-instruction-following20,9975,431,9260.32%
allenai/math_qa29,8372,581,9370.15%
meta-math/MetaMathQA372,90382,026,2684.90%
WizardLMTeam/WizardLM_evol_instruct_V2_196k77,10422,468,3451.34%
Total3,091,6141,673,693,253100.00%

Rows tagged as HuggingFaceH4/ultrachat_200k and HuggingFaceTB/smol-smoltalk inside the curated source were excluded because those datasets are included directly.

Source credits

Additional source contributions came from these public datasets through the Kyoto-Corpus curation:

Contributors

Datdanboi25

4 commits

AxiomicLabs/SFTset-SLM

Dataset

SFTset-SLM

7

stars

4

commits

Sep 10, 2026

updated

sft
slm

README

SFTset-SLM

Source-aware shuffled supervised fine-tuning data formatted for LiquidAI/LFM2.5-1.2B-Instruct.

Dataset summary

  • Conversations: 3,091,614
  • Tokens: 1,673,693,253
  • Parquet parts: 11
  • Target Parquet file size: 500 MiB
  • Tokenizer: LiquidAI/LFM2.5-1.2B-Instruct
  • Shuffle seed: 1337
  • token_count includes the LFM BOS token and ChatML turn-end tokens; no extra terminal EOS is appended.

Columns

  • chatml: LFM2.5 template text, including <|startoftext|> and ChatML markers.
  • messages: structured messages as [{"content": ..., "role": ...}].
  • token_count: exact LFM token count for the chatml value.
  • source: original dataset identifier; Kyoto-derived rows use their underlying repository name.

Source breakdown

SourceConversationsTokensShare of tokens
Yxanul/Mephisto-Knowledge_538k465,737252,573,93915.09%
HuggingFaceH4/ultrachat_200k515,311544,331,75232.52%
HuggingFaceTB/smol-smoltalk460,341404,763,47724.18%
mlabonne/FineTome-100k-dedup99,52359,653,1873.56%
HuggingFaceTB/smoltalk2640,494193,082,11811.54%
allenai/WildChat-1M295,41371,554,8644.28%
argilla/ifeval-like-data20,7684,867,3580.29%
cais/mmlu85,71329,076,9491.74%
openai/gsm8k7,4731,281,1330.08%
allenai/tulu-3-sft-personas-instruction-following20,9975,431,9260.32%
allenai/math_qa29,8372,581,9370.15%
meta-math/MetaMathQA372,90382,026,2684.90%
WizardLMTeam/WizardLM_evol_instruct_V2_196k77,10422,468,3451.34%
Total3,091,6141,673,693,253100.00%

Rows tagged as HuggingFaceH4/ultrachat_200k and HuggingFaceTB/smol-smoltalk inside the curated source were excluded because those datasets are included directly.

Source credits

Additional source contributions came from these public datasets through the Kyoto-Corpus curation:

Contributors

Datdanboi25

4 commits