A Self-feedback Knowledge Elicitation Approach for Chemical Reaction Predictions
Jupyter Notebook
11
19 commits
updated Aug 15, 2025
🎉🎉🎉 Our article was published in Engineering Applications of Artificial Intelligence (EAAI), May 2025 🥳
The task of chemical reaction predictions (CRPs) plays a pivotal role in advancing drug discovery and material science. However, its effectiveness is constrained by the vast and uncertain chemical reaction space and challenges in capturing reaction selectivity, particularly due to existing methods' limitations in exploiting the data's inherent knowledge. To address these challenges, we introduce a data-curated self-feedback knowledge elicitation approach. This method starts from iterative optimization of molecular representations and facilitates the extraction of knowledge on chemical reaction types (RTs). Then, we employ adaptive prompt learning to infuse the prior knowledge into the large language model (LLM). As a result, we achieve significant enhancements: a 14.2% increase in retrosynthesis prediction accuracy, a 74.2% rise in reagent prediction accuracy, and an expansion in the model's capability for handling multi-task chemical reactions. This research offers a novel paradigm for knowledge elicitation in scientific research and showcases the untapped potential of LLMs in CRPs.

Paradigm of the Review:
Note: The dataset and model sections detail respective directories. Some data and checkpoints might not be available due to size constraints and permissions.
The SLM4CRP_with_RTs dataset is a CRPs dataset featuring RT labels, developed from the Mol-Instruction. We introduce a novel knowledge elicitation approach integrating a self-feedback mechanism with data curation using LLMs. This dataset embodies domain-specific knowledge by combining reactants and products of chemical reactions with annotated RTs, demonstrating that domain-integrated data can enhance the capabilities of LLMs.
forward_reaction_prediction.json: Contains data for forward reaction prediction tasks.retrosynthesis.json: Includes data for retrosynthesis tasks.reagent_prediction.json: Features data for predicting reagents in chemical reactions.reactions.json: Serves multiple tasks involving different types of chemical reactions.
ckptsDirectory containing checkpoints used for different purposes:
datasets/MolThis directory includes materials for working with molecular data:
srcSource code directory housing the implementation details:
datasets: Code for constructing datasets.
dataset_manager.py: Creation of datasets for adaptive knowledge injection.dataset_manager_label.py: Creation of datasets for knowledge elicitation.evaluations: Scripts for computing various evaluation metrics.models: Core models for the tasks.
init.py: Initializes model parameters and settings.model_manager.py: Manages the loading and handling of models for knowledge injection.model_manager_label.py: Manages the loading and handling of models for knowledge elicitation.chemT5.py: Text+Chem T5.utils: Utility functions and initializations.
init.py: General utility tool initialization.xutils.py: Advanced and specialized utility tool initialization.task_manager.py: Function to execute tasks related to adaptive knowledge injection.task_manager_label.py: Function to execute tasks related to knowledge elicitation.mode: Select the operation mode. Options include data_check, encoder_check, train, and eval.N: Number of clusters.reaction_type: Specifies whether to include RT during training.task: Task type. Options include forward, retro, reagent, and reactions.batch_size: Set the batch size for operations.

To validate the practical significance of RT annotation, we analyze samples filtered through the concat(input, output)_{vec} vector with N=10 labeled results, focusing on samples with an RT label of 0. These instances typically involve simple atomic substitutions, verifying the predominance of substitution reactions in these cases. This analysis highlights the real-world relevance of our RT annotation method.
@article{LIU2025111112,
title = {A self-feedback knowledge elicitation approach for chemical reaction predictions},
journal = {Engineering Applications of Artificial Intelligence},
volume = {156},
pages = {111112},
year = {2025},
issn = {0952-1976},
doi = {https://doi.org/10.1016/j.engappai.2025.111112},
author = {Pengfei Liu and Jun Tao and Zhixiang Ren}
}
The development of the SLM4CRP_with_RTs dataset was greatly inspired by the Mol-Instruction approach to CRPs. We are also thankful to Hugging Face for providing the initial model weights that facilitated our research.
19 commits
Jupyter Notebook
97.3%
Python
2.7%
A Self-feedback Knowledge Elicitation Approach for Chemical Reaction Predictions
Jupyter Notebook
11
19 commits
updated Aug 15, 2025
🎉🎉🎉 Our article was published in Engineering Applications of Artificial Intelligence (EAAI), May 2025 🥳
The task of chemical reaction predictions (CRPs) plays a pivotal role in advancing drug discovery and material science. However, its effectiveness is constrained by the vast and uncertain chemical reaction space and challenges in capturing reaction selectivity, particularly due to existing methods' limitations in exploiting the data's inherent knowledge. To address these challenges, we introduce a data-curated self-feedback knowledge elicitation approach. This method starts from iterative optimization of molecular representations and facilitates the extraction of knowledge on chemical reaction types (RTs). Then, we employ adaptive prompt learning to infuse the prior knowledge into the large language model (LLM). As a result, we achieve significant enhancements: a 14.2% increase in retrosynthesis prediction accuracy, a 74.2% rise in reagent prediction accuracy, and an expansion in the model's capability for handling multi-task chemical reactions. This research offers a novel paradigm for knowledge elicitation in scientific research and showcases the untapped potential of LLMs in CRPs.

Paradigm of the Review:
Note: The dataset and model sections detail respective directories. Some data and checkpoints might not be available due to size constraints and permissions.
The SLM4CRP_with_RTs dataset is a CRPs dataset featuring RT labels, developed from the Mol-Instruction. We introduce a novel knowledge elicitation approach integrating a self-feedback mechanism with data curation using LLMs. This dataset embodies domain-specific knowledge by combining reactants and products of chemical reactions with annotated RTs, demonstrating that domain-integrated data can enhance the capabilities of LLMs.
forward_reaction_prediction.json: Contains data for forward reaction prediction tasks.retrosynthesis.json: Includes data for retrosynthesis tasks.reagent_prediction.json: Features data for predicting reagents in chemical reactions.reactions.json: Serves multiple tasks involving different types of chemical reactions.
ckptsDirectory containing checkpoints used for different purposes:
datasets/MolThis directory includes materials for working with molecular data:
srcSource code directory housing the implementation details:
datasets: Code for constructing datasets.
dataset_manager.py: Creation of datasets for adaptive knowledge injection.dataset_manager_label.py: Creation of datasets for knowledge elicitation.evaluations: Scripts for computing various evaluation metrics.models: Core models for the tasks.
init.py: Initializes model parameters and settings.model_manager.py: Manages the loading and handling of models for knowledge injection.model_manager_label.py: Manages the loading and handling of models for knowledge elicitation.chemT5.py: Text+Chem T5.utils: Utility functions and initializations.
init.py: General utility tool initialization.xutils.py: Advanced and specialized utility tool initialization.task_manager.py: Function to execute tasks related to adaptive knowledge injection.task_manager_label.py: Function to execute tasks related to knowledge elicitation.mode: Select the operation mode. Options include data_check, encoder_check, train, and eval.N: Number of clusters.reaction_type: Specifies whether to include RT during training.task: Task type. Options include forward, retro, reagent, and reactions.batch_size: Set the batch size for operations.

To validate the practical significance of RT annotation, we analyze samples filtered through the concat(input, output)_{vec} vector with N=10 labeled results, focusing on samples with an RT label of 0. These instances typically involve simple atomic substitutions, verifying the predominance of substitution reactions in these cases. This analysis highlights the real-world relevance of our RT annotation method.
@article{LIU2025111112,
title = {A self-feedback knowledge elicitation approach for chemical reaction predictions},
journal = {Engineering Applications of Artificial Intelligence},
volume = {156},
pages = {111112},
year = {2025},
issn = {0952-1976},
doi = {https://doi.org/10.1016/j.engappai.2025.111112},
author = {Pengfei Liu and Jun Tao and Zhixiang Ren}
}
The development of the SLM4CRP_with_RTs dataset was greatly inspired by the Mol-Instruction approach to CRPs. We are also thankful to Hugging Face for providing the initial model weights that facilitated our research.
19 commits
Jupyter Notebook
97.3%
Python
2.7%