Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text
0
7 commits
2 linked in READMEs
updated May 18, 2025
[NAACL 2025]
The Protein2Text model is a multimodal transformer-based model designed to generate human-interpretable text from protein sequences. It combines a protein sequence encoder (ESM2) and a large language model (LLaMA 3.1-8B Instruct), leveraging resampling mechanisms to improve text generation. The model was trained and fine-tuned on the Protein2Text-QA dataset, which contains question-answer (QA) pairs generated from biomedical literature.
The model was fine-tuned on the Protein2Text-QA dataset, which includes:
| Phase | Global Batch Size | Learning Rate | Epochs | Max Length | Weight Decay | Precision | Optimizer | Gradient Accumulation Steps | Warmup Ratio |
|---|---|---|---|---|---|---|---|---|---|
| Pretraining | 256 | 2 × 10⁻³ | 1 | 2048 | 0 | bf16 (Mixed Precision) | AdamW | 1 step | 0.03 |
| Fine-tuning | 128 | 8 × 10⁻⁶ | 5 | 2048 | 0 | bf16 (Mixed Precision) | AdamW | 1 step | 0.03 |
BibTeX:
@inproceedings{jararweh2025protein2text,
title={Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text},
author={Jararweh, Ala and Macaulay, Oladimeji and Arredondo, David and Hu, Yue and Tafoya, Luis E and Virupakshappa, Kushal and Sahu, Avinash},
booktitle={Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)},
pages={918--937},
year={2025}
}
7 commits
Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text
0
7 commits
2 linked in READMEs
updated May 18, 2025
[NAACL 2025]
The Protein2Text model is a multimodal transformer-based model designed to generate human-interpretable text from protein sequences. It combines a protein sequence encoder (ESM2) and a large language model (LLaMA 3.1-8B Instruct), leveraging resampling mechanisms to improve text generation. The model was trained and fine-tuned on the Protein2Text-QA dataset, which contains question-answer (QA) pairs generated from biomedical literature.
The model was fine-tuned on the Protein2Text-QA dataset, which includes:
| Phase | Global Batch Size | Learning Rate | Epochs | Max Length | Weight Decay | Precision | Optimizer | Gradient Accumulation Steps | Warmup Ratio |
|---|---|---|---|---|---|---|---|---|---|
| Pretraining | 256 | 2 × 10⁻³ | 1 | 2048 | 0 | bf16 (Mixed Precision) | AdamW | 1 step | 0.03 |
| Fine-tuning | 128 | 8 × 10⁻⁶ | 5 | 2048 | 0 | bf16 (Mixed Precision) | AdamW | 1 step | 0.03 |
BibTeX:
@inproceedings{jararweh2025protein2text,
title={Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text},
author={Jararweh, Ala and Macaulay, Oladimeji and Arredondo, David and Hu, Yue and Tafoya, Luis E and Virupakshappa, Kushal and Sahu, Avinash},
booktitle={Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)},
pages={918--937},
year={2025}
}
7 commits