IFM/SFT-Reasoning

Dataset

14

stars

86

commits

Sep 2, 2026

updated

instruction-following
k2-horizon
parquet
reasoning
sft
training-data

README

SFT-Reasoning

Dataset Description

Instruction-following and reasoning data prepared for supervised fine-tuning. This repository is part of the K2 Horizon collection.

The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.

K2 Horizon Dataset Series

Dataset repositoryFocusSubsets
IFM/TxT360-v2Web and question-answering text3
IFM/Code-ReasoningCode reasoning and task synthesis7
IFM/Math-ReasoningMathematical reasoning and dialogue5
IFM/SFT-ReasoningInstruction following and SFT-style data2
IFM/Pretrain-BehaviorsBehavior-focused pretraining data7

Dataset Subsets

SubsetData files
instruction-followinginstruction-following/*.parquet
3efforts-pretrain3efforts-pretrain/*.parquet

Repository Structure

README.md
instruction-following/
  <source-file>-<stable-id>-00000.parquet
  <source-file>-<stable-id>-00001.parquet
3efforts-pretrain/
  <source-file>-<stable-id>-00000.parquet
  <source-file>-<stable-id>-00001.parquet

The shard prefix is derived from the source JSONL filename and a stable identifier. Updating one source JSONL file replaces only that file's Parquet shards.

Data Fields

Records originate as JSON objects and are converted to Parquet for release. Field names and nested structures can differ by subset. Inspect features before building a processing pipeline:

from datasets import load_dataset

dataset = load_dataset(
    "IFM/SFT-Reasoning",
    "instruction-following",
    split="train",
    streaming=True,
)
print(dataset.features)
print(next(iter(dataset)))

Data Provenance and Processing

Individual subsets may have undergone source-specific filtering, cleaning, deduplication, quality scoring, or synthetic-data generation. Users should evaluate each subset for their target use case and inspect the available provenance metadata.

Intended Use

This dataset is intended for language-model training and research. The subsets can be streamed independently, combined with user-defined sampling weights, or inspected through the Hugging Face Dataset Viewer.

Limitations and Responsible Use

Large-scale training data can contain factual errors, duplicated material, sensitive topics, stereotypes, unsafe content, and other artifacts. Dataset users are responsible for performing evaluations, risk assessment, and filtering appropriate to their application.

License and Terms of Use

This dataset is licensed under the Apache License 2.0 available at https://www.apache.org/licenses/LICENSE-2.0.

Contributors

zhou94539

84 commits

moonfolk

2 commits

IFM/SFT-Reasoning

Dataset

14

stars

86

commits

Sep 2, 2026

updated

instruction-following
k2-horizon
parquet
reasoning
sft
training-data

README

SFT-Reasoning

Dataset Description

Instruction-following and reasoning data prepared for supervised fine-tuning. This repository is part of the K2 Horizon collection.

The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access.

K2 Horizon Dataset Series

Dataset repositoryFocusSubsets
IFM/TxT360-v2Web and question-answering text3
IFM/Code-ReasoningCode reasoning and task synthesis7
IFM/Math-ReasoningMathematical reasoning and dialogue5
IFM/SFT-ReasoningInstruction following and SFT-style data2
IFM/Pretrain-BehaviorsBehavior-focused pretraining data7

Dataset Subsets

SubsetData files
instruction-followinginstruction-following/*.parquet
3efforts-pretrain3efforts-pretrain/*.parquet

Repository Structure

README.md
instruction-following/
  <source-file>-<stable-id>-00000.parquet
  <source-file>-<stable-id>-00001.parquet
3efforts-pretrain/
  <source-file>-<stable-id>-00000.parquet
  <source-file>-<stable-id>-00001.parquet

The shard prefix is derived from the source JSONL filename and a stable identifier. Updating one source JSONL file replaces only that file's Parquet shards.

Data Fields

Records originate as JSON objects and are converted to Parquet for release. Field names and nested structures can differ by subset. Inspect features before building a processing pipeline:

from datasets import load_dataset

dataset = load_dataset(
    "IFM/SFT-Reasoning",
    "instruction-following",
    split="train",
    streaming=True,
)
print(dataset.features)
print(next(iter(dataset)))

Data Provenance and Processing

Individual subsets may have undergone source-specific filtering, cleaning, deduplication, quality scoring, or synthetic-data generation. Users should evaluate each subset for their target use case and inspect the available provenance metadata.

Intended Use

This dataset is intended for language-model training and research. The subsets can be streamed independently, combined with user-defined sampling weights, or inspected through the Hugging Face Dataset Viewer.

Limitations and Responsible Use

Large-scale training data can contain factual errors, duplicated material, sensitive topics, stereotypes, unsafe content, and other artifacts. Dataset users are responsible for performing evaluations, risk assessment, and filtering appropriate to their application.

License and Terms of Use

This dataset is licensed under the Apache License 2.0 available at https://www.apache.org/licenses/LICENSE-2.0.

Contributors

zhou94539

84 commits

moonfolk

2 commits