ChemFM/MolLangData

Dataset

MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

1

14 commits

2 linked in READMEs

updated Sep 25, 2026

See the code

README

MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

arXiv GitHub Repo License: MIT Discord

MolLangData provides large-scale paired data of molecular structures and language descriptions generated by a rule-regularized method.

Our previous work MolLangBench contains human curated and validated datasets for molecular structure recognition, editing, and generation. The generation task is similar to MolLangData but is human curated and can be regarded as a gold-standard. Please refer to MolLangBench.

For the codes to generate this dataset and how to use it, please refer to our GitHub.

Dataset Structure

This dataset contains two configs:

  1. validated_data

    • All validated data (2k), including both true descriptions that passed the validation process and the false descriptions.
    • Refer to the validation precision in the dataset statistics table.
  2. generated_data

    • All generated data excluding any validated examples.
    • These examples are not validated.

Validation Labels

The validated_data split uses three independent exact@3 SMILES validators: GPT-5.2, GLM-5.1-FP8 (high reasoning), and DeepSeek-V4-Pro (medium reasoning). automated_validator_passed is true only when all three validators pass. Samples not accepted unanimously use the human-validation labels.

Validation outcomeSamples
All three automated validators passed1,256
First human validator passed709
Second human validator passed7
Final accepted1,972
Final rejected28

The same validation columns are present in generated_data for schema compatibility, but all validation values in that split are null.

Dataset Statistics

Generation difficultyGeneration modelReasoning effortGenerated samplesValidated samplesValidation precision
EasyGPT-5.2high105,085 (65.2%)1,317 (65.8%)1,300 (98.7%)
MediumGPT-5.2xhigh40,916 (25.4%)496 (24.8%)492 (99.2%)
HardGPT-5.2xhigh15,110 (9.4%)187 (9.4%)180 (98.3%)
Overall----161,1112,0001,972 (98.6%)

License

This dataset is distributed under the MIT License.

Citation

Please cite our paper if you use MolLangData in your research.

@article{MolLangData,
  title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
  author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  year={2026},
  journal={arXiv preprint arXiv:2602.02320},
}

Contact

chemistry
molecular_structure_description
molecule_language_alignment

Contributors

feiyang-cai

13 commits

nielsr

1 commits

ChemFM/MolLangData

Dataset

MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

1

14 commits

2 linked in READMEs

updated Sep 25, 2026

See the code

README

MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

arXiv GitHub Repo License: MIT Discord

MolLangData provides large-scale paired data of molecular structures and language descriptions generated by a rule-regularized method.

Our previous work MolLangBench contains human curated and validated datasets for molecular structure recognition, editing, and generation. The generation task is similar to MolLangData but is human curated and can be regarded as a gold-standard. Please refer to MolLangBench.

For the codes to generate this dataset and how to use it, please refer to our GitHub.

Dataset Structure

This dataset contains two configs:

  1. validated_data

    • All validated data (2k), including both true descriptions that passed the validation process and the false descriptions.
    • Refer to the validation precision in the dataset statistics table.
  2. generated_data

    • All generated data excluding any validated examples.
    • These examples are not validated.

Validation Labels

The validated_data split uses three independent exact@3 SMILES validators: GPT-5.2, GLM-5.1-FP8 (high reasoning), and DeepSeek-V4-Pro (medium reasoning). automated_validator_passed is true only when all three validators pass. Samples not accepted unanimously use the human-validation labels.

Validation outcomeSamples
All three automated validators passed1,256
First human validator passed709
Second human validator passed7
Final accepted1,972
Final rejected28

The same validation columns are present in generated_data for schema compatibility, but all validation values in that split are null.

Dataset Statistics

Generation difficultyGeneration modelReasoning effortGenerated samplesValidated samplesValidation precision
EasyGPT-5.2high105,085 (65.2%)1,317 (65.8%)1,300 (98.7%)
MediumGPT-5.2xhigh40,916 (25.4%)496 (24.8%)492 (99.2%)
HardGPT-5.2xhigh15,110 (9.4%)187 (9.4%)180 (98.3%)
Overall----161,1112,0001,972 (98.6%)

License

This dataset is distributed under the MIT License.

Citation

Please cite our paper if you use MolLangData in your research.

@article{MolLangData,
  title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
  author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  year={2026},
  journal={arXiv preprint arXiv:2602.02320},
}

Contact

chemistry
molecular_structure_description
molecule_language_alignment

Contributors

feiyang-cai

13 commits

nielsr

1 commits