OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset for audio-visual and long-video reasoning. It contains two-stage cold-start SFT trajectories with interval selection, zoom-in evidence, and final answers.
data/train.jsonl: HF-ready training JSONL with repo-relative media paths.media/: raw and derived media referenced by train.jsonl.manifests/media_manifest.jsonl: media inventory with repo paths, source
family, media type, file size, and reference counts.manifests/dataset_stats.json: dataset and media statistics.configs/: lmms-engine SFT configs and launch scripts used for training.scripts/materialize_lmms_jsonl.py: rewrites repo-relative paths into local
absolute file:// paths after download.This dataset includes both original-source and derived media:
The final SFT trajectories were produced with Gemini hindsight generation and then filtered with leakage checks, hard rules, judge scoring, and fusion.
The annotations, reasoning traces, metadata, and construction recipes are released by the dataset authors for non-commercial academic research only.
Raw and derived media retain the copyright and usage restrictions of their source datasets, original creators, and original platforms. Access to this repository does not grant additional rights beyond those source licenses and terms. Users must not redistribute the raw or derived media.
If you are a rights holder and want a file removed, please contact the dataset authors through the Hugging Face repository discussion page.
Users should cite this dataset and the applicable upstream datasets, including:
data/train.jsonl stores relative paths such as media/videos/.... To recreate
the absolute file:// format expected by the original lmms-engine training
environment after downloading this repository, run:
python scripts/materialize_lmms_jsonl.py \
--input data/train.jsonl \
--output data/train.local.jsonl \
--root "$(pwd)"
Then point the lmms-engine config to data/train.local.jsonl.
OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset for audio-visual and long-video reasoning. It contains two-stage cold-start SFT trajectories with interval selection, zoom-in evidence, and final answers.
data/train.jsonl: HF-ready training JSONL with repo-relative media paths.media/: raw and derived media referenced by train.jsonl.manifests/media_manifest.jsonl: media inventory with repo paths, source
family, media type, file size, and reference counts.manifests/dataset_stats.json: dataset and media statistics.configs/: lmms-engine SFT configs and launch scripts used for training.scripts/materialize_lmms_jsonl.py: rewrites repo-relative paths into local
absolute file:// paths after download.This dataset includes both original-source and derived media:
The final SFT trajectories were produced with Gemini hindsight generation and then filtered with leakage checks, hard rules, judge scoring, and fusion.
The annotations, reasoning traces, metadata, and construction recipes are released by the dataset authors for non-commercial academic research only.
Raw and derived media retain the copyright and usage restrictions of their source datasets, original creators, and original platforms. Access to this repository does not grant additional rights beyond those source licenses and terms. Users must not redistribute the raw or derived media.
If you are a rights holder and want a file removed, please contact the dataset authors through the Hugging Face repository discussion page.
Users should cite this dataset and the applicable upstream datasets, including:
data/train.jsonl stores relative paths such as media/videos/.... To recreate
the absolute file:// format expected by the original lmms-engine training
environment after downloading this repository, run:
python scripts/materialize_lmms_jsonl.py \
--input data/train.jsonl \
--output data/train.local.jsonl \
--root "$(pwd)"
Then point the lmms-engine config to data/train.local.jsonl.