TheLuoFengLab/MolLangBench

A comprehensive benchmark for evaluating AI models on language-guided molecular structure recognition and manipulation.

Python

14

26 commits

updated Apr 27, 2026

See the code

README


MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

Stargazers Forks Issues Discord

Hugging Face arXiv License: MIT

Table of Contents
  1. About The Project
  2. Related Project
  3. Getting Started
  4. Quick Start
  5. Usage
  6. Miscellaneous
  7. Benchmark Results
  8. Contact
  9. Acknowledgements
  10. Citation
  11. License

About The Project

MolLangBench is the official repository for the ICLR 2026 paper: MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation.

MolLangBench is a comprehensive benchmark designed to evaluate the fundamental capabilities of AI models in language-prompted molecular structure recognition, editing, and generation.

This repository provides:

  • Code and examples to load and use the dataset directly from the Hugging Face Dataset
  • Evaluation scripts and prompt templates to test OpenAI models (e.g., o1, o3, o4-mini) using either molecular images or SMILES strings as inputs

It is straightforward to extend this repository to evaluate other language or multimodal models by adapting the provided input formatting and evaluation templates.

We also release a companion dataset: MolLangData

MolLangData provides large-scale paired data of molecular structures and natural-language descriptions generated via a rule-regularized pipeline, designed to support training and alignment of molecular language models.

MolLangBench focuses on evaluation, while MolLangData focuses on training data construction. Together they form a unified framework for molecular-language alignment research.

(back to top)

Getting Started

You can easily set up the required environment by following these steps:

  • Clone the repository

    git clone https://github.com/TheLuoFengLab/MolLangBench.git
    cd MolLangBench
    
  • Install dependencies

    pip install -r requirements.txt
    

(back to top)

Quick Start

Load the dataset in Python:

from datasets import load_dataset

# Recognition (train + test)
rec_train = load_dataset("ChemFM/MolLangBench", name="recognition", split="train")
rec_test  = load_dataset("ChemFM/MolLangBench", name="recognition", split="test")

# Filter one specific subtask
subtask = "one_hop_neighbors"
subset  = rec_test.filter(lambda x: x["task"] == subtask)

# Editing (test only)
edit = load_dataset("ChemFM/MolLangBench", name="edit", split="test")

# Generation (test only)
gen  = load_dataset("ChemFM/MolLangBench", name="generation", split="test")

(back to top)

Usage

We provide end-to-end scripts for:

  1. Generating prompt files (.jsonl)
  2. Creating batch job inputs for the OpenAI API
  3. Submitting jobs & retrieving outputs
  4. Computing evaluation metrics

Below is a step-by-step example workflow for the “one hop neighbors” recognition subtask.


1. Prepare Prompts

Prompt templates for all tasks and modalities (SMILES and image) are located in the prompts folder. To generate a .jsonl prompt file, run:

python scripts/create_prompts.py \
    --task_type <recognition|editing|generation> \
    --recognition_subtask <recognition_subtask_name_if_applicable> \
    --modality <smiles|image> \
    --split <train|test> \
    --output_file <output_jsonl_path>

Example: For the "one hop neighbors" recognition subtask (SMILES modality, test split):

python scripts/create_prompts.py \
    --task_type recognition \
    --recognition_subtask one_hop_neighbors \
    --modality smiles \
    --split test \
    --output_file exps/one_hop_neighbors/prompts.jsonl

For image modality, the image is included as a Base64-encoded string in the .jsonl file.


2. Create OpenAI Batch Job File

Generate a batch job file for your prompts and desired model:

python scripts/create_openai_jobs.py \
    --prompt_file <prompts_jsonl_path> \
    --output_file <batch_job_jsonl_path> \
    --model <model_id> \
    --custom_id_prefix <optional_prefix>

Example: Using the o4-mini model for "one hop neighbors":

python scripts/create_openai_jobs.py \
    --prompt_file exps/one_hop_neighbors/prompts.jsonl \
    --output_file exps/one_hop_neighbors/o4-mini/batch_input.jsonl \
    --model o4-mini \
    --custom_id_prefix o4_mini

3. Submit Jobs & Retrieve Outputs

Submit the batch job to the OpenAI API:

python scripts/submit_openai_jobs.py submit \
    --jobs_file <batch_job_jsonl_path> \
    [--api_key YOUR_API_KEY] \
    [--organization YOUR_ORG_ID]

You can also set your API key as an environment variable. The organization ID is optional.

Example:

python scripts/submit_openai_jobs.py submit \
    --jobs_file exps/one_hop_neighbors/o4-mini/batch_input.jsonl

This command will print a batch_id for your job.

To retrieve the results (periodically checks until the job is complete):

python scripts/submit_openai_jobs.py retrieve \
    --batch_id <BATCH_ID> \
    --output_file <results_jsonl_path> \
    [--api_key YOUR_API_KEY] \
    [--organization YOUR_ORG_ID] \
    [--check_interval 60]

Example:

python scripts/submit_openai_jobs.py retrieve \
    --batch_id BATCH_ID \
    --output_file exps/one_hop_neighbors/o4-mini/results.jsonl

4. Evaluate the Results

Evaluate the model outputs with:

python scripts/evaluate_results.py \
    --results_file <results_jsonl_path> \
    --task_type <recognition|editing|generation> \
    --subtask <subtask_name> \
    --modality <smiles|image>

Example:

python scripts/evaluate_results.py \
    --results_file exps/one_hop_neighbors/o4-mini/results.jsonl \
    --task_type recognition \
    --subtask one_hop_neighbors \
    --modality smiles

This will print out the evaluation metrics for your selected task and model.

The default result tags are <count> and <atom_indices>. For certain tasks, you may need to specify custom result tags using the --result_1_tag <result_1_tag> argument.

(back to top)


⚠️ Warning:
For molecule editing and generation tasks using the gpt-image-1 model, the OpenAI API does not support batch job submissions.

  • You must submit each image generation or editing request individually using the OpenAI Images API.
  • For automatic evaluation, you will also need a Mathpix account and API key to convert the generated molecular images back to SMILES strings.

Example scripts for these tasks are provided in the Miscellaneous folder.


Miscellaneous

The Miscellaneous folder contains helpful scripts and utilities, including:

  1. Ground Truth Collection
    Scripts for collecting ground truth information for each recognition task using RDKit.

  2. Image-to-SMILES Conversion
    Scripts to call the Mathpix API for converting molecule images to SMILES strings for automated evaluation.

  3. Per-Image OpenAI API Submission
    Scripts to submit image generation and editing requests (for the image modality) to the OpenAI gpt-image-1 API one-by-one, as batch jobs are not currently supported.

More utilities and improvements will be added in the future.

(back to top)

Benchmark Results

Molecular Structure Recognition

Below are the complete evaluation results for all molecular structure recognition subtasks across a wide range of large language models and vision-language models.

Complete Results for Molecular Structure Recognition Tasks (click to expand)
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o4-miniGPT-5Gemini-2.5-ProClaude-Opus-4.1DeepSeek-R1R1-70BLlama-4Qwen3-Maxo3 (image)o4-mini (image)
One-hop neighbors0.355/0.1400.600/0.4250.570/0.3300.735/0.6400.825/0.7200.870/0.8200.935/0.8950.880/0.8450.950/0.9200.905/0.8400.860/0.8050.825/0.7100.585/0.4300.520/0.2900.350/0.1300.890/0.8550.840/0.780
Two-hop neighbors0.215/0.0550.280/0.1000.400/0.2100.465/0.3500.745/0.5600.820/0.7400.935/0.8250.870/0.7900.940/0.9000.885/0.7750.790/0.6700.610/0.4750.305/0.1350.245/0.0650.245/0.0300.770/0.7050.775/0.690
Three-hop neighbors0.165/0.0150.355/0.1650.265/0.1400.400/0.2650.560/0.4000.825/0.7050.925/0.8300.775/0.7100.925/0.8950.820/0.6650.660/0.5250.550/0.3850.300/0.1300.210/0.0550.100/0.0250.695/0.6000.660/0.575
Quaternary carbons0.530/0.2900.690/0.4350.740/0.4400.615/0.4700.865/0.6650.835/0.7400.935/0.8650.845/0.7500.980/0.9100.945/0.8150.875/0.7200.440/0.3300.780/0.6800.345/0.2050.170/0.0700.670/0.6000.720/0.665
Ring junctions0.285/0.0800.495/0.1850.485/0.2100.325/0.1750.575/0.4700.580/0.5200.685/0.6500.590/0.5700.830/0.8100.860/0.7400.655/0.5300.535/0.4200.255/0.1600.385/0.1800.210/0.6500.660/0.5950.615/0.555
Bond connection0.4480.4720.3360.6980.7580.8320.9500.8800.9350.8550.8600.8020.5640.5900.5300.6260.706
Halogen atoms0.845/0.2900.905/0.4200.900/0.3550.920/0.5700.975/0.7400.955/0.7100.965/0.8600.965/0.8201.000/0.9251.000/0.8200.990/0.8600.970/0.7350.740/0.3750.865/0.4550.830/0.2650.855/0.8150.920/0.860
Aldehyde0.855/0.5700.965/0.6100.945/0.7300.855/0.7250.970/0.8250.985/0.9200.990/0.9600.985/0.9451.000/0.9651.000/0.9200.995/0.9500.960/0.8350.715/0.5850.900/0.6450.700/0.5100.925/0.9250.975/0.965
Amide0.505/0.1800.570/0.2050.635/0.3150.585/0.3400.715/0.4400.685/0.5100.765/0.6500.755/0.6100.900/0.7750.920/0.6400.715/0.5850.635/0.4150.495/0.2050.520/0.1950.450/0.1200.565/0.5000.735/0.665
Carboxyl0.760/0.2600.885/0.2350.900/0.4850.840/0.5800.965/0.6750.955/0.7600.985/0.8450.950/0.7250.980/0.8750.990/0.8450.960/0.7150.900/0.6600.820/0.4950.875/0.4350.710/0.2350.785/0.7500.870/0.820
Ester0.600/0.1450.760/0.2850.780/0.3300.675/0.3250.935/0.5000.895/0.6450.955/0.7800.950/0.6400.985/0.8000.975/0.7050.915/0.5900.680/0.4000.615/0.2700.710/0.2200.500/0.1300.720/0.5050.840/0.595
Ketone0.530/0.1550.750/0.2600.870/0.4350.750/0.4650.925/0.6000.985/0.7450.985/0.8650.985/0.7951.000/0.8850.985/0.8250.955/0.7250.880/0.6000.770/0.3700.815/0.3700.575/0.2000.765/0.6750.850/0.775
Benzene0.490/0.1450.540/0.1050.660/0.1550.530/0.2350.720/0.3600.725/0.5650.880/0.6950.730/0.5500.925/0.8400.715/0.4700.695/0.5000.595/0.3850.500/0.1900.590/0.1950.455/0.1050.675/0.4050.680/0.485
Furan0.295/0.2650.820/0.3250.905/0.5150.780/0.5000.920/0.6600.865/0.7450.975/0.8450.940/0.7900.995/0.9150.905/0.7400.940/0.7850.895/0.7100.850/0.4900.935/0.4450.715/0.3250.890/0.8200.870/0.815
Pyridine0.555/0.2250.525/0.2500.730/0.3650.685/0.3750.765/0.5550.860/0.7400.925/0.8250.835/0.7500.915/0.8650.815/0.7200.770/0.6200.685/0.5200.630/0.3400.675/0.2700.485/0.1900.715/0.5850.790/0.665
Thiophene0.860/0.3850.840/0.3250.880/0.4800.840/0.6050.915/0.6900.940/0.7950.970/0.8900.925/0.8201.000/0.9450.925/0.7750.975/0.8200.920/0.7050.850/0.5650.930/0.4550.785/0.3500.960/0.8550.920/0.855
Bond stereo0.3900.3950.6700.4250.3300.3100.4800.3250.6550.2950.5300.3100.3450.5200.4700.5750.640
Chiral stereo0.4400.3950.5300.4650.5100.4350.5450.5200.7000.5450.5100.4400.4950.4200.4650.5100.495
Average0.507/0.2490.625/0.3110.678/0.3910.644/0.4560.776/0.5810.798/0.6800.877/0.7920.817/0.7130.923/0.8620.852/0.7530.814/0.6920.721/0.5660.571/0.3600.614/0.2950.486/0.1860.736/0.6610.772/0.700

  • Each entry reports recognition accuracy / localization accuracy where applicable.
  • Tasks with only recognition evaluation show a single recognition accuracy value.
  • Bold values indicate the best performance among all evaluated language models.
  • "o3 (image)" and "o4-mini (image)" indicate vision-language models evaluated on molecular images.

Molecule Editing and Generation Benchmark Results

Below are the complete evaluation results for molecule editing and generation tasks across all evaluated language and vision-language models.

Complete Results for Molecular Structure Recognition Tasks (click to expand)
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o3 (SELFIES)o4-miniGPT-5DeepSeek-R1R1-70BLlama-4Qwen3-MaxGemini-2.5-ProClaude-Opus-4.1GPT-Image-1
Molecule editing0.725/0.591/0.4000.950/0.823/0.5700.835/0.693/0.4650.710/0.589/0.3850.845/0.788/0.6350.805/0.758/0.6500.945/0.903/0.785 (0.900/0.846/0.670)0.960/0.474/0.195 (0.865/0.372/0.140)0.920/0.860/0.6900.945/0.918/0.855 (0.950/0.890/0.820)0.720/0.643/0.4850.675/0.565/0.3750.895/0.772/0.545 (0.890/0.752/0.490)0.690/0.561/0.360 (0.700/0.496/0.230)0.930/0.881/0.745 (0.945/0.876/0.695)0.950/0.884/0.705 (0.965/0.879/0.665)0.135
Molecule generation0.525/0.174/0.0050.800/0.411/0.0550.710/0.344/0.0350.335/0.170/0.0350.385/0.257/0.1000.450/0.349/0.1750.670/0.546/0.290 (0.695/0.569/0.360)0.185/0.005/0.000 (0.080/0.004/0.000)0.600/0.458/0.2600.690/0.596/0.430 (0.820/0.735/0.590)0.400/0.209/0.0450.205/0.077/0.0100.875/0.511/0.115 (0.870/0.557/0.190)0.465/0.104/0.000 (0.550/0.163/0.050)0.865/0.737/0.430 (0.955/0.833/0.555)0.920/0.725/0.330 (0.970/0.790/0.490)0.000
  • Each entry reports Valid SMILES rate / Tanimoto similarity / task accuracy, measuring chemical validity, structural fidelity, and task success respectively.
  • Results are shown for the core set; values in parentheses denote performance on the extended set.
  • For image-generation models, only the final task accuracy is reported.
  • Bold entries indicate the best performance among all evaluated models.

(back to top)

Contact

Main Developer: Feiyang Cai - feiyang@clemson.edu
Project Supervisor: Feng Luo - luofeng@clemson.edu

Join our community on Discord to stay updated or ask questions.

(back to top)

Acknowledgements

We sincerely thank Reviewer ccMB from the NeurIPS 2025 Datasets and Benchmarks Track for exceptionally thoughtful and constructive feedback. Although the previous submission was not accepted, the reviewer actively championed the work and provided detailed suggestions and encouragement that meaningfully shaped this project and inspired our continued exploration of molecular-language alignment, including the follow-up work MolLangData.

We also thank Yi Hu for assistance with data annotation and Dr. Yongkai Wu (Clemson University) for providing access to Azure AI Foundry resources that supported large-scale evaluation.

(back to top)

Citation

If you find our work valuable, please consider giving the project a star and citing it in your research:

@inproceedings{MolLangBench,
  title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
  author={Cai, Feiyang and Bai, Jiahui and Tang, Tao and He, Guijuan and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
}

Another project from us: We also release MolLangData, a large-scale dataset of molecular structures paired with natural-language descriptions for training and evaluating molecular structure–language models. If you use MolLangData, please cite:

@article{MolLangData,
  title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
  author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  year={2026},
  journal={arXiv preprint arXiv:2602.02320},
}

Thank you for your support!

(back to top)

License

This project is licensed under the MIT License. You are free to use, modify, and distribute this codebase under the terms of the MIT license.

(back to top)

Contributors

feiyang-cai

26 commits

TheLuoFengLab/MolLangBench

A comprehensive benchmark for evaluating AI models on language-guided molecular structure recognition and manipulation.

Python

14

26 commits

updated Apr 27, 2026

See the code

README


MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

Stargazers Forks Issues Discord

Hugging Face arXiv License: MIT

Table of Contents
  1. About The Project
  2. Related Project
  3. Getting Started
  4. Quick Start
  5. Usage
  6. Miscellaneous
  7. Benchmark Results
  8. Contact
  9. Acknowledgements
  10. Citation
  11. License

About The Project

MolLangBench is the official repository for the ICLR 2026 paper: MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation.

MolLangBench is a comprehensive benchmark designed to evaluate the fundamental capabilities of AI models in language-prompted molecular structure recognition, editing, and generation.

This repository provides:

  • Code and examples to load and use the dataset directly from the Hugging Face Dataset
  • Evaluation scripts and prompt templates to test OpenAI models (e.g., o1, o3, o4-mini) using either molecular images or SMILES strings as inputs

It is straightforward to extend this repository to evaluate other language or multimodal models by adapting the provided input formatting and evaluation templates.

We also release a companion dataset: MolLangData

MolLangData provides large-scale paired data of molecular structures and natural-language descriptions generated via a rule-regularized pipeline, designed to support training and alignment of molecular language models.

MolLangBench focuses on evaluation, while MolLangData focuses on training data construction. Together they form a unified framework for molecular-language alignment research.

(back to top)

Getting Started

You can easily set up the required environment by following these steps:

  • Clone the repository

    git clone https://github.com/TheLuoFengLab/MolLangBench.git
    cd MolLangBench
    
  • Install dependencies

    pip install -r requirements.txt
    

(back to top)

Quick Start

Load the dataset in Python:

from datasets import load_dataset

# Recognition (train + test)
rec_train = load_dataset("ChemFM/MolLangBench", name="recognition", split="train")
rec_test  = load_dataset("ChemFM/MolLangBench", name="recognition", split="test")

# Filter one specific subtask
subtask = "one_hop_neighbors"
subset  = rec_test.filter(lambda x: x["task"] == subtask)

# Editing (test only)
edit = load_dataset("ChemFM/MolLangBench", name="edit", split="test")

# Generation (test only)
gen  = load_dataset("ChemFM/MolLangBench", name="generation", split="test")

(back to top)

Usage

We provide end-to-end scripts for:

  1. Generating prompt files (.jsonl)
  2. Creating batch job inputs for the OpenAI API
  3. Submitting jobs & retrieving outputs
  4. Computing evaluation metrics

Below is a step-by-step example workflow for the “one hop neighbors” recognition subtask.


1. Prepare Prompts

Prompt templates for all tasks and modalities (SMILES and image) are located in the prompts folder. To generate a .jsonl prompt file, run:

python scripts/create_prompts.py \
    --task_type <recognition|editing|generation> \
    --recognition_subtask <recognition_subtask_name_if_applicable> \
    --modality <smiles|image> \
    --split <train|test> \
    --output_file <output_jsonl_path>

Example: For the "one hop neighbors" recognition subtask (SMILES modality, test split):

python scripts/create_prompts.py \
    --task_type recognition \
    --recognition_subtask one_hop_neighbors \
    --modality smiles \
    --split test \
    --output_file exps/one_hop_neighbors/prompts.jsonl

For image modality, the image is included as a Base64-encoded string in the .jsonl file.


2. Create OpenAI Batch Job File

Generate a batch job file for your prompts and desired model:

python scripts/create_openai_jobs.py \
    --prompt_file <prompts_jsonl_path> \
    --output_file <batch_job_jsonl_path> \
    --model <model_id> \
    --custom_id_prefix <optional_prefix>

Example: Using the o4-mini model for "one hop neighbors":

python scripts/create_openai_jobs.py \
    --prompt_file exps/one_hop_neighbors/prompts.jsonl \
    --output_file exps/one_hop_neighbors/o4-mini/batch_input.jsonl \
    --model o4-mini \
    --custom_id_prefix o4_mini

3. Submit Jobs & Retrieve Outputs

Submit the batch job to the OpenAI API:

python scripts/submit_openai_jobs.py submit \
    --jobs_file <batch_job_jsonl_path> \
    [--api_key YOUR_API_KEY] \
    [--organization YOUR_ORG_ID]

You can also set your API key as an environment variable. The organization ID is optional.

Example:

python scripts/submit_openai_jobs.py submit \
    --jobs_file exps/one_hop_neighbors/o4-mini/batch_input.jsonl

This command will print a batch_id for your job.

To retrieve the results (periodically checks until the job is complete):

python scripts/submit_openai_jobs.py retrieve \
    --batch_id <BATCH_ID> \
    --output_file <results_jsonl_path> \
    [--api_key YOUR_API_KEY] \
    [--organization YOUR_ORG_ID] \
    [--check_interval 60]

Example:

python scripts/submit_openai_jobs.py retrieve \
    --batch_id BATCH_ID \
    --output_file exps/one_hop_neighbors/o4-mini/results.jsonl

4. Evaluate the Results

Evaluate the model outputs with:

python scripts/evaluate_results.py \
    --results_file <results_jsonl_path> \
    --task_type <recognition|editing|generation> \
    --subtask <subtask_name> \
    --modality <smiles|image>

Example:

python scripts/evaluate_results.py \
    --results_file exps/one_hop_neighbors/o4-mini/results.jsonl \
    --task_type recognition \
    --subtask one_hop_neighbors \
    --modality smiles

This will print out the evaluation metrics for your selected task and model.

The default result tags are <count> and <atom_indices>. For certain tasks, you may need to specify custom result tags using the --result_1_tag <result_1_tag> argument.

(back to top)


⚠️ Warning:
For molecule editing and generation tasks using the gpt-image-1 model, the OpenAI API does not support batch job submissions.

  • You must submit each image generation or editing request individually using the OpenAI Images API.
  • For automatic evaluation, you will also need a Mathpix account and API key to convert the generated molecular images back to SMILES strings.

Example scripts for these tasks are provided in the Miscellaneous folder.


Miscellaneous

The Miscellaneous folder contains helpful scripts and utilities, including:

  1. Ground Truth Collection
    Scripts for collecting ground truth information for each recognition task using RDKit.

  2. Image-to-SMILES Conversion
    Scripts to call the Mathpix API for converting molecule images to SMILES strings for automated evaluation.

  3. Per-Image OpenAI API Submission
    Scripts to submit image generation and editing requests (for the image modality) to the OpenAI gpt-image-1 API one-by-one, as batch jobs are not currently supported.

More utilities and improvements will be added in the future.

(back to top)

Benchmark Results

Molecular Structure Recognition

Below are the complete evaluation results for all molecular structure recognition subtasks across a wide range of large language models and vision-language models.

Complete Results for Molecular Structure Recognition Tasks (click to expand)
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o4-miniGPT-5Gemini-2.5-ProClaude-Opus-4.1DeepSeek-R1R1-70BLlama-4Qwen3-Maxo3 (image)o4-mini (image)
One-hop neighbors0.355/0.1400.600/0.4250.570/0.3300.735/0.6400.825/0.7200.870/0.8200.935/0.8950.880/0.8450.950/0.9200.905/0.8400.860/0.8050.825/0.7100.585/0.4300.520/0.2900.350/0.1300.890/0.8550.840/0.780
Two-hop neighbors0.215/0.0550.280/0.1000.400/0.2100.465/0.3500.745/0.5600.820/0.7400.935/0.8250.870/0.7900.940/0.9000.885/0.7750.790/0.6700.610/0.4750.305/0.1350.245/0.0650.245/0.0300.770/0.7050.775/0.690
Three-hop neighbors0.165/0.0150.355/0.1650.265/0.1400.400/0.2650.560/0.4000.825/0.7050.925/0.8300.775/0.7100.925/0.8950.820/0.6650.660/0.5250.550/0.3850.300/0.1300.210/0.0550.100/0.0250.695/0.6000.660/0.575
Quaternary carbons0.530/0.2900.690/0.4350.740/0.4400.615/0.4700.865/0.6650.835/0.7400.935/0.8650.845/0.7500.980/0.9100.945/0.8150.875/0.7200.440/0.3300.780/0.6800.345/0.2050.170/0.0700.670/0.6000.720/0.665
Ring junctions0.285/0.0800.495/0.1850.485/0.2100.325/0.1750.575/0.4700.580/0.5200.685/0.6500.590/0.5700.830/0.8100.860/0.7400.655/0.5300.535/0.4200.255/0.1600.385/0.1800.210/0.6500.660/0.5950.615/0.555
Bond connection0.4480.4720.3360.6980.7580.8320.9500.8800.9350.8550.8600.8020.5640.5900.5300.6260.706
Halogen atoms0.845/0.2900.905/0.4200.900/0.3550.920/0.5700.975/0.7400.955/0.7100.965/0.8600.965/0.8201.000/0.9251.000/0.8200.990/0.8600.970/0.7350.740/0.3750.865/0.4550.830/0.2650.855/0.8150.920/0.860
Aldehyde0.855/0.5700.965/0.6100.945/0.7300.855/0.7250.970/0.8250.985/0.9200.990/0.9600.985/0.9451.000/0.9651.000/0.9200.995/0.9500.960/0.8350.715/0.5850.900/0.6450.700/0.5100.925/0.9250.975/0.965
Amide0.505/0.1800.570/0.2050.635/0.3150.585/0.3400.715/0.4400.685/0.5100.765/0.6500.755/0.6100.900/0.7750.920/0.6400.715/0.5850.635/0.4150.495/0.2050.520/0.1950.450/0.1200.565/0.5000.735/0.665
Carboxyl0.760/0.2600.885/0.2350.900/0.4850.840/0.5800.965/0.6750.955/0.7600.985/0.8450.950/0.7250.980/0.8750.990/0.8450.960/0.7150.900/0.6600.820/0.4950.875/0.4350.710/0.2350.785/0.7500.870/0.820
Ester0.600/0.1450.760/0.2850.780/0.3300.675/0.3250.935/0.5000.895/0.6450.955/0.7800.950/0.6400.985/0.8000.975/0.7050.915/0.5900.680/0.4000.615/0.2700.710/0.2200.500/0.1300.720/0.5050.840/0.595
Ketone0.530/0.1550.750/0.2600.870/0.4350.750/0.4650.925/0.6000.985/0.7450.985/0.8650.985/0.7951.000/0.8850.985/0.8250.955/0.7250.880/0.6000.770/0.3700.815/0.3700.575/0.2000.765/0.6750.850/0.775
Benzene0.490/0.1450.540/0.1050.660/0.1550.530/0.2350.720/0.3600.725/0.5650.880/0.6950.730/0.5500.925/0.8400.715/0.4700.695/0.5000.595/0.3850.500/0.1900.590/0.1950.455/0.1050.675/0.4050.680/0.485
Furan0.295/0.2650.820/0.3250.905/0.5150.780/0.5000.920/0.6600.865/0.7450.975/0.8450.940/0.7900.995/0.9150.905/0.7400.940/0.7850.895/0.7100.850/0.4900.935/0.4450.715/0.3250.890/0.8200.870/0.815
Pyridine0.555/0.2250.525/0.2500.730/0.3650.685/0.3750.765/0.5550.860/0.7400.925/0.8250.835/0.7500.915/0.8650.815/0.7200.770/0.6200.685/0.5200.630/0.3400.675/0.2700.485/0.1900.715/0.5850.790/0.665
Thiophene0.860/0.3850.840/0.3250.880/0.4800.840/0.6050.915/0.6900.940/0.7950.970/0.8900.925/0.8201.000/0.9450.925/0.7750.975/0.8200.920/0.7050.850/0.5650.930/0.4550.785/0.3500.960/0.8550.920/0.855
Bond stereo0.3900.3950.6700.4250.3300.3100.4800.3250.6550.2950.5300.3100.3450.5200.4700.5750.640
Chiral stereo0.4400.3950.5300.4650.5100.4350.5450.5200.7000.5450.5100.4400.4950.4200.4650.5100.495
Average0.507/0.2490.625/0.3110.678/0.3910.644/0.4560.776/0.5810.798/0.6800.877/0.7920.817/0.7130.923/0.8620.852/0.7530.814/0.6920.721/0.5660.571/0.3600.614/0.2950.486/0.1860.736/0.6610.772/0.700

  • Each entry reports recognition accuracy / localization accuracy where applicable.
  • Tasks with only recognition evaluation show a single recognition accuracy value.
  • Bold values indicate the best performance among all evaluated language models.
  • "o3 (image)" and "o4-mini (image)" indicate vision-language models evaluated on molecular images.

Molecule Editing and Generation Benchmark Results

Below are the complete evaluation results for molecule editing and generation tasks across all evaluated language and vision-language models.

Complete Results for Molecular Structure Recognition Tasks (click to expand)
TaskGPT-4oGPT-4.5GPT-4.1o1-minio1o3-minio3o3 (SELFIES)o4-miniGPT-5DeepSeek-R1R1-70BLlama-4Qwen3-MaxGemini-2.5-ProClaude-Opus-4.1GPT-Image-1
Molecule editing0.725/0.591/0.4000.950/0.823/0.5700.835/0.693/0.4650.710/0.589/0.3850.845/0.788/0.6350.805/0.758/0.6500.945/0.903/0.785 (0.900/0.846/0.670)0.960/0.474/0.195 (0.865/0.372/0.140)0.920/0.860/0.6900.945/0.918/0.855 (0.950/0.890/0.820)0.720/0.643/0.4850.675/0.565/0.3750.895/0.772/0.545 (0.890/0.752/0.490)0.690/0.561/0.360 (0.700/0.496/0.230)0.930/0.881/0.745 (0.945/0.876/0.695)0.950/0.884/0.705 (0.965/0.879/0.665)0.135
Molecule generation0.525/0.174/0.0050.800/0.411/0.0550.710/0.344/0.0350.335/0.170/0.0350.385/0.257/0.1000.450/0.349/0.1750.670/0.546/0.290 (0.695/0.569/0.360)0.185/0.005/0.000 (0.080/0.004/0.000)0.600/0.458/0.2600.690/0.596/0.430 (0.820/0.735/0.590)0.400/0.209/0.0450.205/0.077/0.0100.875/0.511/0.115 (0.870/0.557/0.190)0.465/0.104/0.000 (0.550/0.163/0.050)0.865/0.737/0.430 (0.955/0.833/0.555)0.920/0.725/0.330 (0.970/0.790/0.490)0.000
  • Each entry reports Valid SMILES rate / Tanimoto similarity / task accuracy, measuring chemical validity, structural fidelity, and task success respectively.
  • Results are shown for the core set; values in parentheses denote performance on the extended set.
  • For image-generation models, only the final task accuracy is reported.
  • Bold entries indicate the best performance among all evaluated models.

(back to top)

Contact

Main Developer: Feiyang Cai - feiyang@clemson.edu
Project Supervisor: Feng Luo - luofeng@clemson.edu

Join our community on Discord to stay updated or ask questions.

(back to top)

Acknowledgements

We sincerely thank Reviewer ccMB from the NeurIPS 2025 Datasets and Benchmarks Track for exceptionally thoughtful and constructive feedback. Although the previous submission was not accepted, the reviewer actively championed the work and provided detailed suggestions and encouragement that meaningfully shaped this project and inspired our continued exploration of molecular-language alignment, including the follow-up work MolLangData.

We also thank Yi Hu for assistance with data annotation and Dr. Yongkai Wu (Clemson University) for providing access to Azure AI Foundry resources that supported large-scale evaluation.

(back to top)

Citation

If you find our work valuable, please consider giving the project a star and citing it in your research:

@inproceedings{MolLangBench,
  title={MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation},
  author={Cai, Feiyang and Bai, Jiahui and Tang, Tao and He, Guijuan and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
}

Another project from us: We also release MolLangData, a large-scale dataset of molecular structures paired with natural-language descriptions for training and evaluating molecular structure–language models. If you use MolLangData, please cite:

@article{MolLangData,
  title={A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method},
  author={Cai, Feiyang and He, Guijuan and Hu, Yi and Wang, Jingjing and Luo, Joshua and Zhu, Tianyu and Pilla, Srikanth and Li, Gang and Liu, Ling and Luo, Feng},
  year={2026},
  journal={arXiv preprint arXiv:2602.02320},
}

Thank you for your support!

(back to top)

License

This project is licensed under the MIT License. You are free to use, modify, and distribute this codebase under the terms of the MIT license.

(back to top)

Contributors

feiyang-cai

26 commits

Languages

Python

100.0%