🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary conditions—and then constructs challenging, self-contained problems from complementary reasoning perspectives. (More meta information is being prepared and will be updated once ready.)
Built on SCI-BASE (~3.36M papers across 10 disciplines), SPARK samples ~370K frontier papers after filtering and discipline-balanced sampling, then proceeds through four stages:
Spark-234K spans 10 scientific disciplines sampled in a discipline-balanced manner, with Physics, Mathematics & Computer Science, and Chemistry as the largest shares.
Spark-234K emphasizes difficult, diverse, and self-contained reasoning:
Spark-234K consistently outperforms substantially larger scientific reasoning datasets when used for SFT on Qwen-series models:
| Training Data | Qwen3-8B | Qwen3-14B |
|---|---|---|
| Best competing dataset | 56.36 | 60.45 |
| Spark-234K | 60.83 | 63.88 |
For more detailed experimental results and analysis, please refer to our paper.
If you find Spark-234K useful, please consider citing:
@misc{li2026sparkskeletonguidedreasoningsynthesis,
title={SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature},
author={Yu Li and Wei Li and Xin Gao and Mengyuan Sun and Xiaoyang Wang and Qizhi Pei and Lijun Wu},
year={2026},
eprint={2608.30214},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.30214},
}
🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary conditions—and then constructs challenging, self-contained problems from complementary reasoning perspectives. (More meta information is being prepared and will be updated once ready.)
Built on SCI-BASE (~3.36M papers across 10 disciplines), SPARK samples ~370K frontier papers after filtering and discipline-balanced sampling, then proceeds through four stages:
Spark-234K spans 10 scientific disciplines sampled in a discipline-balanced manner, with Physics, Mathematics & Computer Science, and Chemistry as the largest shares.
Spark-234K emphasizes difficult, diverse, and self-contained reasoning:
Spark-234K consistently outperforms substantially larger scientific reasoning datasets when used for SFT on Qwen-series models:
| Training Data | Qwen3-8B | Qwen3-14B |
|---|---|---|
| Best competing dataset | 56.36 | 60.45 |
| Spark-234K | 60.83 | 63.88 |
For more detailed experimental results and analysis, please refer to our paper.
If you find Spark-234K useful, please consider citing:
@misc{li2026sparkskeletonguidedreasoningsynthesis,
title={SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature},
author={Yu Li and Wei Li and Xin Gao and Mengyuan Sun and Xiaoyang Wang and Qizhi Pei and Lijun Wu},
year={2026},
eprint={2608.30214},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.30214},
}