This repo contains the datasets for reproducing the results of our ICML 2025 paper: *Hierarchical Graph Tokenization for Molecule-Language Alignment*, which has also been presented at ICML 2024 workshop on Foundation Models in the Wild. πππ
2
5 commits
1 linked in READMEs
updated Aug 18, 2025
This repo contains the datasets for reproducing the results of our ICML 2025 paper: Hierarchical Graph Tokenization for Molecule-Language Alignment, which has also been presented at ICML 2024 workshop on Foundation Models in the Wild. πππ
This is the dataset, stored in file hi_data_dict_lap_fgprompt.pkl, we curated from PubChem to perform the stage 1 pretraining, i.e., SFT, our graph-language model.
In contrast, data_dict.pkl contains the vanilla stage 1 pretraining data.
This is the dataset we use to evaluate the motif hallucination of different models. The specific dataset we used in the paper is stored in hight_smiles100.jsonl.
If you find our data, paper and repo useful, please cite our paper:
@inproceedings{chen2025hierarchical,
title={Hierarchical Graph Tokenization for Molecule-Language Alignment},
author={Yongqiang Chen and Quanming Yao and Juzheng Zhang and James Cheng and Yatao Bian},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=wpbNczwAwV}
}
5 commits
This repo contains the datasets for reproducing the results of our ICML 2025 paper: *Hierarchical Graph Tokenization for Molecule-Language Alignment*, which has also been presented at ICML 2024 workshop on Foundation Models in the Wild. πππ
2
5 commits
1 linked in READMEs
updated Aug 18, 2025
This repo contains the datasets for reproducing the results of our ICML 2025 paper: Hierarchical Graph Tokenization for Molecule-Language Alignment, which has also been presented at ICML 2024 workshop on Foundation Models in the Wild. πππ
This is the dataset, stored in file hi_data_dict_lap_fgprompt.pkl, we curated from PubChem to perform the stage 1 pretraining, i.e., SFT, our graph-language model.
In contrast, data_dict.pkl contains the vanilla stage 1 pretraining data.
This is the dataset we use to evaluate the motif hallucination of different models. The specific dataset we used in the paper is stored in hight_smiles100.jsonl.
If you find our data, paper and repo useful, please cite our paper:
@inproceedings{chen2025hierarchical,
title={Hierarchical Graph Tokenization for Molecule-Language Alignment},
author={Yongqiang Chen and Quanming Yao and Juzheng Zhang and James Cheng and Yatao Bian},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=wpbNczwAwV}
}
5 commits