MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
1
14 commits
2 linked in READMEs
updated Sep 25, 2026
MolLangData provides large-scale paired data of molecular structures and language descriptions generated by a rule-regularized method.
Our previous work MolLangBench contains human curated and validated datasets for molecular structure recognition, editing, and generation. The generation task is similar to MolLangData but is human curated and can be regarded as a gold-standard. Please refer to MolLangBench.
For the codes to generate this dataset and how to use it, please refer to our GitHub.
This dataset contains two configs:
validated_data
generated_data
The validated_data split uses three independent exact@3 SMILES validators:
GPT-5.2, GLM-5.1-FP8 (high reasoning), and DeepSeek-V4-Pro (medium
reasoning). automated_validator_passed is true only when all three validators
pass. Samples not accepted unanimously use the human-validation labels.
| Validation outcome | Samples |
|---|---|
| All three automated validators passed | 1,256 |
| First human validator passed | 709 |
| Second human validator passed | 7 |
| Final accepted | 1,972 |
| Final rejected | 28 |
The same validation columns are present in generated_data for schema
compatibility, but all validation values in that split are null.
| Generation difficulty | Generation model | Reasoning effort | Generated samples | Validated samples | Validation precision |
|---|---|---|---|---|---|
| Easy | GPT-5.2 | high | 105,085 (65.2%) | 1,317 (65.8%) | 1,300 (98.7%) |
| Medium | GPT-5.2 | xhigh | 40,916 (25.4%) | 496 (24.8%) | 492 (99.2%) |
| Hard | GPT-5.2 | xhigh | 15,110 (9.4%) | 187 (9.4%) | 180 (98.3%) |
| Overall | -- | -- | 161,111 | 2,000 | 1,972 (98.6%) |
This dataset is distributed under the MIT License.
Please cite our paper if you use MolLangData in your research.
@article{MolLangData,
title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
year={2026},
journal={arXiv preprint arXiv:2602.02320},
}
13 commits
1 commits
MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
1
14 commits
2 linked in READMEs
updated Sep 25, 2026
MolLangData provides large-scale paired data of molecular structures and language descriptions generated by a rule-regularized method.
Our previous work MolLangBench contains human curated and validated datasets for molecular structure recognition, editing, and generation. The generation task is similar to MolLangData but is human curated and can be regarded as a gold-standard. Please refer to MolLangBench.
For the codes to generate this dataset and how to use it, please refer to our GitHub.
This dataset contains two configs:
validated_data
generated_data
The validated_data split uses three independent exact@3 SMILES validators:
GPT-5.2, GLM-5.1-FP8 (high reasoning), and DeepSeek-V4-Pro (medium
reasoning). automated_validator_passed is true only when all three validators
pass. Samples not accepted unanimously use the human-validation labels.
| Validation outcome | Samples |
|---|---|
| All three automated validators passed | 1,256 |
| First human validator passed | 709 |
| Second human validator passed | 7 |
| Final accepted | 1,972 |
| Final rejected | 28 |
The same validation columns are present in generated_data for schema
compatibility, but all validation values in that split are null.
| Generation difficulty | Generation model | Reasoning effort | Generated samples | Validated samples | Validation precision |
|---|---|---|---|---|---|
| Easy | GPT-5.2 | high | 105,085 (65.2%) | 1,317 (65.8%) | 1,300 (98.7%) |
| Medium | GPT-5.2 | xhigh | 40,916 (25.4%) | 496 (24.8%) | 492 (99.2%) |
| Hard | GPT-5.2 | xhigh | 15,110 (9.4%) | 187 (9.4%) | 180 (98.3%) |
| Overall | -- | -- | 161,111 | 2,000 | 1,972 (98.6%) |
This dataset is distributed under the MIT License.
Please cite our paper if you use MolLangData in your research.
@article{MolLangData,
title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
year={2026},
journal={arXiv preprint arXiv:2602.02320},
}
13 commits
1 commits