ChemFM/MolLangBench

Dataset

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

2

35 commits

3 linked in READMEs

updated Feb 10, 2026

See the code

README

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

This is the official dataset repository for the ICLR 2026 paper: MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation.

The code for using and evaluating the MolLangBench datasets is provided in this GitHub repository.

MolLangBench is a comprehensive benchmark designed to evaluate the fundamental capabilities of AI models in language-prompted molecular structure recognition, editing, and generation.

arXiv GitHub Repo License: MIT Discord


Dataset Structure

MolLangBench consists of three core tasks:

  • Molecular Structure Recognition:
    Evaluate a model's ability to interpret and reason over molecule representations (SMILES strings, images) from natural language prompts.
  • Molecule Editing:
    Assess a model's capability to apply language-based structural edits to molecules.
  • Molecule Generation:
    Evaluate a model's ability to generate new molecule structures based on text prompts.

For complete code, usage instructions, and evaluation pipelines, please visit our GitHub repository.


Benchmark Results

Molecular Structure Recognition

Click to view the full results table
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o4-miniGPT-5Gemini-2.5-ProClaude-Opus-4.1DeepSeek-R1R1-70BLlama-4Qwen3-Maxo3 (image)o4-mini (image)
One-hop neighbors0.355/0.1400.600/0.4250.570/0.3300.735/0.6400.825/0.7200.870/0.8200.935/0.8950.880/0.8450.950/0.9200.905/0.8400.860/0.8050.825/0.7100.585/0.4300.520/0.2900.350/0.1300.890/0.8550.840/0.780
Two-hop neighbors0.215/0.0550.280/0.1000.400/0.2100.465/0.3500.745/0.5600.820/0.7400.935/0.8250.870/0.7900.940/0.9000.885/0.7750.790/0.6700.610/0.4750.305/0.1350.245/0.0650.245/0.0300.770/0.7050.775/0.690
Three-hop neighbors0.165/0.0150.355/0.1650.265/0.1400.400/0.2650.560/0.4000.825/0.7050.925/0.8300.775/0.7100.925/0.8950.820/0.6650.660/0.5250.550/0.3850.300/0.1300.210/0.0550.100/0.0250.695/0.6000.660/0.575
Quaternary carbons0.530/0.2900.690/0.4350.740/0.4400.615/0.4700.865/0.6650.835/0.7400.935/0.8650.845/0.7500.980/0.9100.945/0.8150.875/0.7200.440/0.3300.780/0.6800.345/0.2050.170/0.0700.670/0.6000.720/0.665
Ring junctions0.285/0.0800.495/0.1850.485/0.2100.325/0.1750.575/0.4700.580/0.5200.685/0.6500.590/0.5700.830/0.8100.860/0.7400.655/0.5300.535/0.4200.255/0.1600.385/0.1800.210/0.6500.660/0.5950.615/0.555
Bond connection0.4480.4720.3360.6980.7580.8320.9500.8800.9350.8550.8600.8020.5640.5900.5300.6260.706
Halogen atoms0.845/0.2900.905/0.4200.900/0.3550.920/0.5700.975/0.7400.955/0.7100.965/0.8600.965/0.8201.000/0.9251.000/0.8200.990/0.8600.970/0.7350.740/0.3750.865/0.4550.830/0.2650.855/0.8150.920/0.860
Aldehyde0.855/0.5700.965/0.6100.945/0.7300.855/0.7250.970/0.8250.985/0.9200.990/0.9600.985/0.9451.000/0.9651.000/0.9200.995/0.9500.960/0.8350.715/0.5850.900/0.6450.700/0.5100.925/0.9250.975/0.965
Amide0.505/0.1800.570/0.2050.635/0.3150.585/0.3400.715/0.4400.685/0.5100.765/0.6500.755/0.6100.900/0.7750.920/0.6400.715/0.5850.635/0.4150.495/0.2050.520/0.1950.450/0.1200.565/0.5000.735/0.665
Carboxyl0.760/0.2600.885/0.2350.900/0.4850.840/0.5800.965/0.6750.955/0.7600.985/0.8450.950/0.7250.980/0.8750.990/0.8450.960/0.7150.900/0.6600.820/0.4950.875/0.4350.710/0.2350.785/0.7500.870/0.820
Ester0.600/0.1450.760/0.2850.780/0.3300.675/0.3250.935/0.5000.895/0.6450.955/0.7800.950/0.6400.985/0.8000.975/0.7050.915/0.5900.680/0.4000.615/0.2700.710/0.2200.500/0.1300.720/0.5050.840/0.595
Ketone0.530/0.1550.750/0.2600.870/0.4350.750/0.4650.925/0.6000.985/0.7450.985/0.8650.985/0.7951.000/0.8850.985/0.8250.955/0.7250.880/0.6000.770/0.3700.815/0.3700.575/0.2000.765/0.6750.850/0.775
Benzene0.490/0.1450.540/0.1050.660/0.1550.530/0.2350.720/0.3600.725/0.5650.880/0.6950.730/0.5500.925/0.8400.715/0.4700.695/0.5000.595/0.3850.500/0.1900.590/0.1950.455/0.1050.675/0.4050.680/0.485
Furan0.295/0.2650.820/0.3250.905/0.5150.780/0.5000.920/0.6600.865/0.7450.975/0.8450.940/0.7900.995/0.9150.905/0.7400.940/0.7850.895/0.7100.850/0.4900.935/0.4450.715/0.3250.890/0.8200.870/0.815
Pyridine0.555/0.2250.525/0.2500.730/0.3650.685/0.3750.765/0.5550.860/0.7400.925/0.8250.835/0.7500.915/0.8650.815/0.7200.770/0.6200.685/0.5200.630/0.3400.675/0.2700.485/0.1900.715/0.5850.790/0.665
Thiophene0.860/0.3850.840/0.3250.880/0.4800.840/0.6050.915/0.6900.940/0.7950.970/0.8900.925/0.8201.000/0.9450.925/0.7750.975/0.8200.920/0.7050.850/0.5650.930/0.4550.785/0.3500.960/0.8550.920/0.855
Bond stereo0.3900.3950.6700.4250.3300.3100.4800.3250.6550.2950.5300.3100.3450.5200.4700.5750.640
Chiral stereo0.4400.3950.5300.4650.5100.4350.5450.5200.7000.5450.5100.4400.4950.4200.4650.5100.495
Average0.507/0.2490.625/0.3110.678/0.3910.644/0.4560.776/0.5810.798/0.6800.877/0.7920.817/0.7130.923/0.8620.852/0.7530.814/0.6920.721/0.5660.571/0.3600.614/0.2950.486/0.1860.736/0.6610.772/0.700
  • Each entry reports recognition accuracy / localization accuracy where applicable.
  • Tasks with only recognition evaluation show a single recognition accuracy value.
  • Bold values indicate the best performance among all evaluated language models.
  • "o3 (image)" and "o4-mini (image)" indicate vision-language models evaluated on molecular images.

Molecule Editing and Generation Benchmark Results

Click to view editing and generation results
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o3 (SELFIES)o4-miniGPT-5DeepSeek-R1R1-70BLlama-4Qwen3-MaxGemini-2.5-ProClaude-Opus-4.1GPT-Image-1
Molecule editing0.725/0.591/0.4000.950/0.823/0.5700.835/0.693/0.4650.710/0.589/0.3850.845/0.788/0.6350.805/0.758/0.6500.945/0.903/0.785 (0.900/0.846/0.670)0.960/0.474/0.195 (0.865/0.372/0.140)0.920/0.860/0.6900.945/0.918/0.855 (0.950/0.890/0.820)0.720/0.643/0.4850.675/0.565/0.3750.895/0.772/0.545 (0.890/0.752/0.490)0.690/0.561/0.360 (0.700/0.496/0.230)0.930/0.881/0.745 (0.945/0.876/0.695)0.950/0.884/0.705 (0.965/0.879/0.665)0.135
Molecule generation0.525/0.174/0.0050.800/0.411/0.0550.710/0.344/0.0350.335/0.170/0.0350.385/0.257/0.1000.450/0.349/0.1750.670/0.546/0.290 (0.695/0.569/0.360)0.185/0.005/0.000 (0.080/0.004/0.000)0.600/0.458/0.2600.690/0.596/0.430 (0.820/0.735/0.590)0.400/0.209/0.0450.205/0.077/0.0100.875/0.511/0.115 (0.870/0.557/0.190)0.465/0.104/0.000 (0.550/0.163/0.050)0.865/0.737/0.430 (0.955/0.833/0.555)0.920/0.725/0.330 (0.970/0.790/0.490)0.000
  • Each entry reports Valid SMILES rate / Tanimoto similarity / task accuracy, measuring chemical validity, structural fidelity, and task success respectively.
  • Results are shown for the core set; values in parentheses denote performance on the extended set.
  • For image-generation models, only the final task accuracy is reported.
  • Bold entries indicate the best performance among all evaluated models.

Citation

Please cite our paper if you use MolLangBench in your research:

@inproceedings{MolLangBench,
  title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
  author={Feiyang Cai and Jiahui Bai and Tao Tang and Guijuan He and Joshua Luo and Tianyu Zhu and Srikanth Pilla and Gang Li and Ling Liu and Feng Luo},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
}

License

This dataset is distributed under the MIT License.

molecular_structure_recognition
molecule
molecule_editing
molecule_generation
molecule_graph
molecule_image
multimodal
smiles

Contributors

feiyang-cai

35 commits

ChemFM/MolLangBench

Dataset

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

2

35 commits

3 linked in READMEs

updated Feb 10, 2026

See the code

README

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

This is the official dataset repository for the ICLR 2026 paper: MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation.

The code for using and evaluating the MolLangBench datasets is provided in this GitHub repository.

MolLangBench is a comprehensive benchmark designed to evaluate the fundamental capabilities of AI models in language-prompted molecular structure recognition, editing, and generation.

arXiv GitHub Repo License: MIT Discord


Dataset Structure

MolLangBench consists of three core tasks:

  • Molecular Structure Recognition:
    Evaluate a model's ability to interpret and reason over molecule representations (SMILES strings, images) from natural language prompts.
  • Molecule Editing:
    Assess a model's capability to apply language-based structural edits to molecules.
  • Molecule Generation:
    Evaluate a model's ability to generate new molecule structures based on text prompts.

For complete code, usage instructions, and evaluation pipelines, please visit our GitHub repository.


Benchmark Results

Molecular Structure Recognition

Click to view the full results table
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o4-miniGPT-5Gemini-2.5-ProClaude-Opus-4.1DeepSeek-R1R1-70BLlama-4Qwen3-Maxo3 (image)o4-mini (image)
One-hop neighbors0.355/0.1400.600/0.4250.570/0.3300.735/0.6400.825/0.7200.870/0.8200.935/0.8950.880/0.8450.950/0.9200.905/0.8400.860/0.8050.825/0.7100.585/0.4300.520/0.2900.350/0.1300.890/0.8550.840/0.780
Two-hop neighbors0.215/0.0550.280/0.1000.400/0.2100.465/0.3500.745/0.5600.820/0.7400.935/0.8250.870/0.7900.940/0.9000.885/0.7750.790/0.6700.610/0.4750.305/0.1350.245/0.0650.245/0.0300.770/0.7050.775/0.690
Three-hop neighbors0.165/0.0150.355/0.1650.265/0.1400.400/0.2650.560/0.4000.825/0.7050.925/0.8300.775/0.7100.925/0.8950.820/0.6650.660/0.5250.550/0.3850.300/0.1300.210/0.0550.100/0.0250.695/0.6000.660/0.575
Quaternary carbons0.530/0.2900.690/0.4350.740/0.4400.615/0.4700.865/0.6650.835/0.7400.935/0.8650.845/0.7500.980/0.9100.945/0.8150.875/0.7200.440/0.3300.780/0.6800.345/0.2050.170/0.0700.670/0.6000.720/0.665
Ring junctions0.285/0.0800.495/0.1850.485/0.2100.325/0.1750.575/0.4700.580/0.5200.685/0.6500.590/0.5700.830/0.8100.860/0.7400.655/0.5300.535/0.4200.255/0.1600.385/0.1800.210/0.6500.660/0.5950.615/0.555
Bond connection0.4480.4720.3360.6980.7580.8320.9500.8800.9350.8550.8600.8020.5640.5900.5300.6260.706
Halogen atoms0.845/0.2900.905/0.4200.900/0.3550.920/0.5700.975/0.7400.955/0.7100.965/0.8600.965/0.8201.000/0.9251.000/0.8200.990/0.8600.970/0.7350.740/0.3750.865/0.4550.830/0.2650.855/0.8150.920/0.860
Aldehyde0.855/0.5700.965/0.6100.945/0.7300.855/0.7250.970/0.8250.985/0.9200.990/0.9600.985/0.9451.000/0.9651.000/0.9200.995/0.9500.960/0.8350.715/0.5850.900/0.6450.700/0.5100.925/0.9250.975/0.965
Amide0.505/0.1800.570/0.2050.635/0.3150.585/0.3400.715/0.4400.685/0.5100.765/0.6500.755/0.6100.900/0.7750.920/0.6400.715/0.5850.635/0.4150.495/0.2050.520/0.1950.450/0.1200.565/0.5000.735/0.665
Carboxyl0.760/0.2600.885/0.2350.900/0.4850.840/0.5800.965/0.6750.955/0.7600.985/0.8450.950/0.7250.980/0.8750.990/0.8450.960/0.7150.900/0.6600.820/0.4950.875/0.4350.710/0.2350.785/0.7500.870/0.820
Ester0.600/0.1450.760/0.2850.780/0.3300.675/0.3250.935/0.5000.895/0.6450.955/0.7800.950/0.6400.985/0.8000.975/0.7050.915/0.5900.680/0.4000.615/0.2700.710/0.2200.500/0.1300.720/0.5050.840/0.595
Ketone0.530/0.1550.750/0.2600.870/0.4350.750/0.4650.925/0.6000.985/0.7450.985/0.8650.985/0.7951.000/0.8850.985/0.8250.955/0.7250.880/0.6000.770/0.3700.815/0.3700.575/0.2000.765/0.6750.850/0.775
Benzene0.490/0.1450.540/0.1050.660/0.1550.530/0.2350.720/0.3600.725/0.5650.880/0.6950.730/0.5500.925/0.8400.715/0.4700.695/0.5000.595/0.3850.500/0.1900.590/0.1950.455/0.1050.675/0.4050.680/0.485
Furan0.295/0.2650.820/0.3250.905/0.5150.780/0.5000.920/0.6600.865/0.7450.975/0.8450.940/0.7900.995/0.9150.905/0.7400.940/0.7850.895/0.7100.850/0.4900.935/0.4450.715/0.3250.890/0.8200.870/0.815
Pyridine0.555/0.2250.525/0.2500.730/0.3650.685/0.3750.765/0.5550.860/0.7400.925/0.8250.835/0.7500.915/0.8650.815/0.7200.770/0.6200.685/0.5200.630/0.3400.675/0.2700.485/0.1900.715/0.5850.790/0.665
Thiophene0.860/0.3850.840/0.3250.880/0.4800.840/0.6050.915/0.6900.940/0.7950.970/0.8900.925/0.8201.000/0.9450.925/0.7750.975/0.8200.920/0.7050.850/0.5650.930/0.4550.785/0.3500.960/0.8550.920/0.855
Bond stereo0.3900.3950.6700.4250.3300.3100.4800.3250.6550.2950.5300.3100.3450.5200.4700.5750.640
Chiral stereo0.4400.3950.5300.4650.5100.4350.5450.5200.7000.5450.5100.4400.4950.4200.4650.5100.495
Average0.507/0.2490.625/0.3110.678/0.3910.644/0.4560.776/0.5810.798/0.6800.877/0.7920.817/0.7130.923/0.8620.852/0.7530.814/0.6920.721/0.5660.571/0.3600.614/0.2950.486/0.1860.736/0.6610.772/0.700
  • Each entry reports recognition accuracy / localization accuracy where applicable.
  • Tasks with only recognition evaluation show a single recognition accuracy value.
  • Bold values indicate the best performance among all evaluated language models.
  • "o3 (image)" and "o4-mini (image)" indicate vision-language models evaluated on molecular images.

Molecule Editing and Generation Benchmark Results

Click to view editing and generation results
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o3 (SELFIES)o4-miniGPT-5DeepSeek-R1R1-70BLlama-4Qwen3-MaxGemini-2.5-ProClaude-Opus-4.1GPT-Image-1
Molecule editing0.725/0.591/0.4000.950/0.823/0.5700.835/0.693/0.4650.710/0.589/0.3850.845/0.788/0.6350.805/0.758/0.6500.945/0.903/0.785 (0.900/0.846/0.670)0.960/0.474/0.195 (0.865/0.372/0.140)0.920/0.860/0.6900.945/0.918/0.855 (0.950/0.890/0.820)0.720/0.643/0.4850.675/0.565/0.3750.895/0.772/0.545 (0.890/0.752/0.490)0.690/0.561/0.360 (0.700/0.496/0.230)0.930/0.881/0.745 (0.945/0.876/0.695)0.950/0.884/0.705 (0.965/0.879/0.665)0.135
Molecule generation0.525/0.174/0.0050.800/0.411/0.0550.710/0.344/0.0350.335/0.170/0.0350.385/0.257/0.1000.450/0.349/0.1750.670/0.546/0.290 (0.695/0.569/0.360)0.185/0.005/0.000 (0.080/0.004/0.000)0.600/0.458/0.2600.690/0.596/0.430 (0.820/0.735/0.590)0.400/0.209/0.0450.205/0.077/0.0100.875/0.511/0.115 (0.870/0.557/0.190)0.465/0.104/0.000 (0.550/0.163/0.050)0.865/0.737/0.430 (0.955/0.833/0.555)0.920/0.725/0.330 (0.970/0.790/0.490)0.000
  • Each entry reports Valid SMILES rate / Tanimoto similarity / task accuracy, measuring chemical validity, structural fidelity, and task success respectively.
  • Results are shown for the core set; values in parentheses denote performance on the extended set.
  • For image-generation models, only the final task accuracy is reported.
  • Bold entries indicate the best performance among all evaluated models.

Citation

Please cite our paper if you use MolLangBench in your research:

@inproceedings{MolLangBench,
  title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
  author={Feiyang Cai and Jiahui Bai and Tao Tang and Guijuan He and Joshua Luo and Tianyu Zhu and Srikanth Pilla and Gang Li and Ling Liu and Feng Luo},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
}

License

This dataset is distributed under the MIT License.

molecular_structure_recognition
molecule
molecule_editing
molecule_generation
molecule_graph
molecule_image
multimodal
smiles

Contributors

feiyang-cai

35 commits