Curated list of 1,000+ scientific foundation models spanning life sciences, chemistry, physics, medicine, and more.
37
26 commits
updated Jun 17, 2026
One map, every frontier — your starting point for 1,000+ AI-for-Science models.
A comprehensive, bilingual (EN/ZH) guide to foundation models driving the next wave of scientific breakthroughs — 1,000+ models across nine domains, from protein design to weather prediction. Cross-disciplinary models are listed wherever they apply.
Language: English | Chinese
Large-scale pretrained models learning protein representations from amino acid sequences.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| ESM-1b | Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences | Large-scale unsupervised protein language model trained on 250 million sequences to capture evolutionary diversity and biological properties. | PNAS |
| ESM-2 | Evolutionary-scale prediction of atomic-level protein structure with a language model | 15-billion-parameter protein language model enabling atomic-level structure prediction from single sequences (ESMFold). | Science |
| ESM3 | Simulating 500 million years of evolution with a language model | Multimodal protein generative model that jointly processes sequence, structure, and function to simulate 500 million years of evolution. | Science |
| ESMFold | Evolutionary-scale prediction of atomic-level protein structure with a language model | Single-sequence protein structure prediction method based on ESM-2, approximately 60× faster than AlphaFold2. | Science |
| ESM Cambrian (ESMC) | ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning | Next-generation ESM protein language model that reveals intrinsic biological principles of proteins through unsupervised learning. | EvolutionaryScale |
| ProtTrans | ProtTrans: Toward understanding the language of life through self-supervised learning | Suite of large-scale protein pretrained models (ProtBERT, ProtXLNet, ProtT5, etc.) trained on 393 billion amino acids. | IEEE TPAMI |
| ProteinBERT | ProteinBERT: a universal deep-learning model of protein sequence and function | Universal protein model jointly pretrained on sequences and GO annotations for diverse protein property prediction. | Bioinformatics |
| ProtGPT2 | ProtGPT2 is a deep unsupervised language model for protein design | GPT-2-based protein sequence generation model that produces novel sequences resembling natural proteins. | Nature Communications |
| ProGen | Large language models generate functional protein sequences across diverse families | Large-scale protein language model from Salesforce trained on 280 million sequences to generate functional artificial proteins. | Nature Biotechnology |
| ProGen2 | ProGen2: Exploring the boundaries of protein language models | 6.4-billion-parameter protein language model trained on genomic, metagenomic, and protein family data. | Cell Systems |
| ProGen3 | Scaling unlocks broader generation and deeper functional understanding of proteins | Sparse mixture-of-experts protein generative model trained on 1.5 trillion tokens for enhanced protein design. | bioRxiv |
| SaProt | SaProt: Protein language modeling with structure-aware vocabulary | Structure-aware protein language model that encodes 3D structures as discrete tokens via Foldseek for joint training with sequences. | ICLR 2024 |
| xTrimoPGLM | xTrimoPGLM: Unified 100B-scale pre-trained transformer for deciphering the language of protein | 100-billion-parameter unified protein language model supporting both protein understanding and generation tasks. | Nature Methods |
| Ankh | Ankh: Optimized protein language model unlocks general-purpose modelling | Optimized protein language model emphasizing training efficiency to achieve competitive performance with fewer resources. | arXiv |
| ProCyon | ProCyon: A multimodal foundation model for protein phenotypes | Multimodal protein foundation model integrating sequence, structure, and natural language data to predict protein phenotypes. | bioRxiv |
| ProSST | ProSST: Protein language modeling with quantized structure and disentangled attention | Protein language model combining amino acid sequences with quantized 3D structural information. | NeurIPS 2024 |
| InstructPLM | InstructPLM: Aligning protein language models to follow protein design instructions | Instruction-tuned ESM2 model that surpasses ESM3 on protein design tasks through alignment training. | arXiv |
| UniRep | Unified rational protein engineering with sequence-based deep representation learning | RNN-based protein language model for unified representation learning and rational protein engineering. | Nature Methods |
| MSA Transformer | MSA Transformer | Transformer model that operates over multiple sequence alignments for improved protein modeling. | ICML |
| EVE | Disease variant prediction with deep generative models of evolutionary data | Evolutionary variational autoencoder for predicting disease-causing genetic variants from protein family data. | Nature |
| Tranception | Protein fitness prediction with autoregressive transformers and inference-time retrieval | Autoregressive language model with retrieval-time alignment context for protein fitness prediction. | ICML |
| RITA | RITA: a Study on Scaling Up Generative Protein Sequence Models | Scaling study of autoregressive protein generative models up to 1.2 billion parameters. | arXiv |
| ProtSSN | Semantical and Geometrical Protein Encoding for Zero-Shot Engineering | Structure-plus-sequence denoising pretraining framework for zero-shot protein engineering. | eLife |
| ProLLaMA | ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing | Multi-task protein LLM based on the LLaMA architecture for diverse protein language processing tasks. | TAI |
| ProTrek | ProTrek: Navigating the Protein Universe through Tri-Modal Contrastive Learning | Tri-modal protein model learning joint representations of sequence, structure, and function via contrastive learning. | Nature Biotechnology |
| ProtWord | ProtWord: A Discrete Protein Language Model for Functional Discovery and De Novo Design | 150M-parameter discrete protein language model that translates sequences into an 8,192-token vocabulary for functional discovery. | bioRxiv |
| ProtLLM | ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training | Interleaved protein-language large language model with a dynamic protein mounting mechanism. | ACL |
| ProtHyena | Hyena architecture enables fast and efficient protein language modeling | Fast Hyena-based protein language model leveraging implicit convolution for efficient sequence processing. | iMetaOmics |
| DPLM | Diffusion Language Models Are Versatile Protein Learners | Diffusion-based language model for versatile protein sequence generation and understanding. | ICML |
| PTM-Mamba | PTM-Mamba: a PTM-aware protein language model with bidirectional gated Mamba blocks | State-space model with bidirectional gated Mamba blocks for post-translational-modification-aware protein representation. | Nature Methods |
| DeepSequence | Deep generative models of genetic variation capture the effects of mutations | Variational autoencoder over aligned protein families for predicting the effects of mutations. | Nature Methods |
| PoET | PoET: A generative model of protein families as sequences-of-sequences | Sequences-of-sequences family modeling framework with retrieval for protein fitness prediction. | Advances in Neural Information Processing Systems 36 |
| Prot2Text | Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers | Multimodal model generating natural language descriptions of protein functions from structure and sequence. | AAAI 2024 |
| PAIR | Boosting the predictive power of protein representations with a corpus of text annotations | Method that boosts protein representations by incorporating text annotation corpora. | Nature Machine Intelligence |
| MULAN | MULAN: Multimodal protein language model for sequence and structure encoding | Multimodal protein encoder that jointly models sequence and structure information. | Bioinformatics Advances |
| ProteinSage | ProteinSage: From implicit learning to explicit structural constraints for efficient protein language modeling | Protein language model combining implicit learning with explicit structural constraints for improved efficiency. | bioRxiv |
| OneProt | OneProt: Towards Multi-Modal Protein Foundation Models | Multi-modal protein foundation model integrating structure, sequence, text, and binding site data. | arXiv |
| ECNet | ECNet is an evolutionary context-integrated deep learning framework for protein engineering | Deep learning framework integrating evolutionary context for protein engineering and fitness prediction. | Nature Communications |
| PoET-2 | Understanding protein function with a multimodal retrieval-augmented foundation model | Next-generation retrieval-augmented multimodal protein family model improving fitness prediction over PoET. | arXiv |
| FlexRibbon | FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling | 3-billion-parameter model jointly pretrained on amino acid sequences and 3D structures capturing flexible conformations. | bioRxiv |
| ProteinAligner | ProteinAligner: A Tri-Modal Contrastive Learning Framework for Protein Representation Learning | Tri-modal contrastive learning framework integrating protein sequences, structures, and scientific literature. | OpenReview |
| ProteinTalks | ProteinTalks: Multi-Modal Protein Language Model with Natural Language Interaction | Multi-modal protein language model supporting natural language interaction for protein knowledge retrieval. | bioRxiv |
| GearNet | Protein Representation Learning by Geometric Structure Pretraining | Relational graph neural network learning protein structure representations via geometric pretraining and multi-view contrastive learning. | ICLR |
| MIF | Masked inverse folding with sequence transfer for protein representation learning | Self-supervised masked inverse folding pretraining for learning protein representations from structures. | Protein Engineering, Design and Selection |
| PPLM | A paired sequence language model for protein-protein interaction | Paired sequence language model predicting protein-protein interactions from paired amino acid sequences. | Nature Communications |
| AIDO.Protein | Mixture of experts enable efficient and effective protein understanding and design | 16B-parameter mixture-of-experts protein module trained on 1.2 trillion amino acids within the AIDO ecosystem. | bioRxiv |
| BioReason-Pro | BioReason-Pro: Advancing Protein Function Prediction with Multimodal Biological Reasoning | First multimodal reasoning LLM for protein function prediction integrating ESM3 embeddings with GO-GPT ontology modeling; achieves 73.6% F_max on GO term prediction, preferred over UniProt annotations by human experts 79% of the time. | bioRxiv |
Methods for predicting 3D protein structures from sequences.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AlphaFold2 | Highly accurate protein structure prediction with AlphaFold | Revolutionary protein structure prediction model achieving atomic-level accuracy at CASP14. | Nature |
| AlphaFold3 | Accurate structure prediction of biomolecular interactions with AlphaFold 3 | Diffusion-based model predicting biomolecular interaction structures for proteins, nucleic acids, small molecules, and ions. | Nature |
| RoseTTAFold | Accurate prediction of protein structures and interactions using a three-track neural network | Three-track neural network for protein structure prediction as an open-source alternative to AlphaFold2. | Science |
| RoseTTAFold2 | Efficient and accurate prediction of protein structure using RoseTTAFold2 | Upgraded RoseTTAFold combining key features from AlphaFold2 and the original RoseTTAFold. | bioRxiv |
| RoseTTAFold All-Atom | Generalized biomolecular modeling and design with RoseTTAFold All-Atom | All-atom biomolecular modeling framework supporting complex prediction of proteins, nucleic acids, small molecules, and metal ions. | Science |
| OmegaFold | High-resolution de novo structure prediction from primary sequence | MSA-free single-sequence protein structure prediction leveraging pretrained protein language models. | bioRxiv |
Generative models for de novo protein backbone and sequence design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| ProteinMPNN | Robust deep learning-based protein sequence design using ProteinMPNN | Message-passing neural network for protein sequence design with experimentally validated high success rates. | Science |
| RFdiffusion | De novo design of protein structure and function with RFdiffusion | Diffusion model built on RoseTTAFold for de novo protein backbone design. | Nature |
| RFdiffusion3 | RFdiffusion3: All-atom biomolecular design | Upgraded RFdiffusion supporting all-atom-level protein and biomolecular design. | bioRxiv |
| Chroma | Illuminating protein space with a programmable generative model | Programmable protein diffusion generative model from Generate:Biomedicines. | Nature |
| FrameDiff | SE(3) diffusion model with application to protein backbone generation | SE(3)-equivariant diffusion model for protein backbone generation without pretrained structure prediction networks. | ICLR |
| FrameFlow | SE(3) stochastic flow matching for protein backbone generation | SE(3) flow matching model for protein backbone generation. | ICLR |
| FoldingDiff | Protein structure generation via folding diffusion | Diffusion model based on protein folding angle representations that simulates natural folding processes. | Nature Communications |
| Genie | Genie: SE(3)-equivariant generative model for protein backbone design | SE(3)-equivariant DDPM model for protein backbone design. | ICML (Workshop) |
| Genie 2 | Out of many, one: Designing and scaffolding proteins at the scale of the structural universe with Genie 2 | Upgraded Genie capturing a broader and more diverse protein structure space. | arXiv |
| Proteus | Proteus: Exploring protein structure generation for enhanced designability and efficiency | Efficient protein backbone generation model without requiring pretrained structure prediction networks. | bioRxiv |
| FoldFlow | FoldFlow: SE(3) stochastic flow matching for protein backbone generation | Family of stochastic flow matching models for protein backbone generation. | ICLR |
| ProteinGenerator (PG) | Multistate and functional protein design using RoseTTAFold | RoseTTAFold-based diffusion model that simultaneously generates protein sequences and structures. | Nature Biotechnology |
| SeedProteo | SeedProteo: All-atom protein design | Diffusion-based all-atom protein design model integrating structure and sequence information for binder design. | arXiv |
| ESM-IF (ESM-IF1) | Language models generalize beyond natural proteins | Inverse folding model conditioned on backbone structures for generating protein sequences. | bioRxiv |
| EvoDiff | Protein generation with evolutionary diffusion | Sequence-based diffusion model for protein generation leveraging evolutionary data. | bioRxiv/Nature Biotechnology |
| ZymCTRL | ZymCTRL: a conditional language model for the generation of artificial enzymes | Conditional enzyme language model trained on 37 million BRENDA enzyme sequences. | bioRxiv |
| PiFold | PiFold: Toward effective and efficient protein inverse folding | Efficient inverse folding model with a novel PiGNN architecture. | ICLR |
| LM-Design | Structure-informed language models are protein designers | Structure-informed language model for protein design combining PLM and structural context. | ICML |
| ProteinDT | A text-guided protein design framework | Text-guided protein generation framework using multimodal learning. | Nature Machine Intelligence |
| Protpardelle | An all-atom protein generative model | All-atom generative model for producing complete protein structures including side chains. | PNAS |
| Multiflow | Generative Flows on Discrete State-Spaces for Protein Co-Design | Discrete flow matching framework for joint sequence-structure protein co-design. | ICML |
| La-Proteina | La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching | Atomistic protein generation method via partially latent flow matching. | arXiv |
| EvoFlows | Evolutionary Edit-Based Flow-Matching for Protein Engineering | Evolutionary edit-based flow matching approach for directed protein engineering. | arXiv |
| Fold2Seq | Fold2Seq: A joint sequence(1D)-fold(3D) embedding-based generative model for protein design | Joint sequence-fold embedding generative model learning from both 3D structures and 1D sequences. | ICML |
| TERMinator | TERMinator: A neural framework for structure-based protein design using tertiary repeating motifs | Neural network framework for protein design based on tertiary repeating motifs. | Nature Communications |
| AlphaDesign | AlphaDesign: A graph protein design method and benchmark on AlphaFoldDB | Graph-based protein design method benchmarked on the AlphaFold Database. | arXiv |
| Latent-X | Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design | Atom-level frontier model generating all-atom structures and sequences for de novo protein binder design. | arXiv |
| Latent-X2 | Drug-like antibodies with low immunogenicity in human panels designed with Latent-X2 | Generative model designing drug-like antibodies with strong binding affinity and low immunogenicity. | arXiv |
| ProDiT | Generating functional proteins with a multimodal diffusion transformer | Multimodal diffusion Transformer protein design model trained on 214 million proteins. | bioRxiv |
| SimpleDesign | SimpleDesign: Joint Model for Protein Sequence and Structure Codesign | End-to-end joint model for protein sequence-structure co-design without tokenizers. | ICLR |
| PXDesign | PXDesign: Fast De Novo Design of Protein Binders | ByteDance Protenix fast and modular pipeline for de novo protein binder design with 20–73% hit rates. | bioRxiv |
Foundation models for peptide design, antimicrobial peptides, and cyclic peptides.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| PepMLM | Target Sequence-Conditioned Generation of Therapeutic Peptide Binders via Span Masked Language Modeling | ESM-2 fine-tuned span masked LM generating linear peptide binders conditioned on target protein sequences. | Nature Biotechnology / ICLR |
| PepBERT | PepBERT: Lightweight language models for peptide representation | Lightweight dedicated peptide language model for bioactive peptide discovery and representation learning. | bioRxiv |
| PepDoRA | PepDoRA: A Unified Peptide Language Model via Weight-Decomposed Low-Rank Adaptation | Unified peptide language model predicting multiple peptide properties via weight-decomposed low-rank adaptation. | arXiv |
| AMP-Designer | A foundation model approach to guide antimicrobial peptide design in the era of AI-driven scientific discovery | LLM-based foundation model for de novo design of antimicrobial peptides with significant antibacterial activity. | arXiv |
| AMP-Diffusion | AMP-Diffusion: Integrating Latent Diffusion with Protein Language Models for Antimicrobial Peptide Generation | Latent diffusion model integrated with ESM-2 for generating novel antimicrobial peptides. | bioRxiv |
| deepAMP | A Foundation Model Identifies Broad-Spectrum Antimicrobial Peptides against Drug-Resistant Bacterial Infection | Deep generative framework based on peptide language models identifying broad-spectrum antimicrobial peptides. | Nature Communications |
| RFpeptides | Accurate de novo design of high-affinity protein-binding macrocycles | RoseTTAFold-based denoising diffusion model for de novo macrocyclic peptide design. | Nature Chemical Biology |
| AfCycDesign | Cyclic peptide structure prediction and design using AlphaFold2 | AlphaFold2-based method for cyclic peptide structure prediction, redesign, and de novo generation. | Nature Communications |
| CP-Composer | Zero-Shot Cyclic Peptide Design via Composable Geometric Constraints | Framework for zero-shot cyclic peptide design using composable geometric constraints. | ICML |
| CpSDE | Designing Cyclic Peptides via Harmonic SDE with Atom-Bond Modeling | Cyclic peptide design method using harmonic stochastic differential equations with atom-bond modeling. | ICML |
| PDeepPP | A general language model for peptide identification | General deep learning framework for peptide function prediction combining pretrained PLMs with Transformer-CNN architecture. | arXiv |
Models predicting protein–protein interactions and complex structures.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| PLM-interact | PLM-interact: extending protein language models to predict protein-protein interactions | Extension of protein language models for PPI prediction by jointly encoding protein pairs from sequences alone. | Nature Communications |
| IntFold | IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction | Controllable biomolecular structure prediction foundation model matching AlphaFold3 accuracy with support for PPI complexes and allosteric states. | arXiv |
Models capturing protein thermodynamics, conformational dynamics, and molecular dynamics.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| ProTDyn | ProTDyn: a foundation Protein language model for Thermodynamics and Dynamics | Protein thermodynamics and dynamics foundation model unifying conformational ensemble generation and multi-timescale dynamics modeling. | NeurIPS |
| SeqDance | Learning Biophysical Dynamics with Protein Language Models | Protein language model incorporating biophysical dynamics, trained on MD simulations and normal mode analyses of 64,000+ proteins. | bioRxiv |
| ESMDance | Learning Biophysical Dynamics with Protein Language Models | ESM-2 fine-tuned variant for protein conformational dynamics prediction. | bioRxiv |
| MD-LLM-1 | MD-LLM-1: A Large Language Model for Molecular Dynamics | First molecular dynamics LLM, fine-tuning Mistral 7B for protein conformational dynamics prediction. | arXiv |
| VibeGen | Agentic End-to-End De Novo Protein Design for Tailored Dynamics Using a Language Diffusion Model | Language diffusion model framework for end-to-end protein design targeting specific vibrational dynamics. | arXiv |
| DynamicsPLM | Learning Protein Representations with Conformational Dynamics | Protein language model learning from computationally generated conformational dynamics ensembles. | bioRxiv |
| DPLM-2 | DPLM-2: A Multimodal Diffusion Protein Language Model | Multimodal discrete diffusion protein language model jointly modeling amino acid sequences and 3D structures. | NeurIPS 2025 |
| METL | Biophysics-based protein language models for protein engineering | Mutation effect transfer learning framework integrating biophysical modeling and machine learning for protein engineering. | Nature Methods |
Pretrained models for RNA sequence representation, structure inference, and function prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| RNA-FM | Interpretable RNA foundation model from unannotated data for highly accurate RNA structure and function predictions | RNA foundation model pretrained on 23 million non-coding RNA sequences for structure and function prediction. | Nature Methods |
| RiNALMo | RiNALMo: general-purpose RNA language models can generalize well on structure prediction tasks | 650-million-parameter general-purpose RNA language model trained on 36 million non-coding RNA sequences with strong structure prediction generalization. | Nature Communications |
| ERNIE-RNA | ERNIE-RNA: an RNA language model with structure-enhanced representations | Modified BERT-based RNA language model incorporating base-pairing structure information for enhanced representations. | Nature Communications |
| RNA-MSM | Multiple sequence alignment-based RNA language model and its application to structural inference | MSA-based RNA language model leveraging co-evolutionary information from homologous RNAs for structure inference. | Nucleic Acids Research |
| UNI-RNA | UNI-RNA: Universal pre-trained models revolutionize RNA research | Universal pretrained RNA model trained on the largest RNA sequence dataset supporting multiple downstream tasks. | bioRxiv |
| RNAErnie | Multi-purpose RNA language modelling with motif-aware pretraining and type-guided fine-tuning | Multi-purpose RNA language model combining motif-aware pretraining with RNA-type-guided fine-tuning. | Nature Machine Intelligence |
| AIDO.RNA | A large-scale foundation model for RNA function and structure prediction | 1.6-billion-parameter RNA foundation model trained on 42 million non-coding RNA sequences at single-nucleotide resolution. | bioRxiv |
| BiRNA-BERT | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization | Adaptive tokenization RNA language model overcoming limitations in sequence length and diversity. | Communications Biology |
| mRNABERT | mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset | Language model specifically designed for mRNA sequence engineering with a dual-tokenization scheme. | Nature Communications |
| CodonBERT | CodonBERT: Large language models for mRNA design and optimization | Codon-level tokenized mRNA language model for mRNA design and optimization. | BEACON Benchmark |
| NucleicBERT | NucleicBERT: A large language model for RNA structure prediction | BERT-based self-supervised masked language model for RNA sequence analysis and structure prediction. | bioRxiv |
| MP-RNA | MP-RNA: Unleashing multi-species RNA foundation model via calibrated secondary structure prediction | Multi-species RNA foundation model emphasizing calibrated secondary structure prediction. | EMNLP 2024 |
| OmniGenome | Bridging sequence-structure alignment in RNA foundation models | RNA foundation model precisely aligning RNA sequences with secondary structures for bidirectional mapping. | arXiv |
| structRFM | A fully open structure-guided RNA foundation model for robust structural and functional inference | Fully open-source structure-guided RNA foundation model integrating sequence and secondary structure for robust inference. | bioRxiv |
| SpliceBERT | SpliceBERT: a pre-trained RNA language model for analyzing vertebrate splicing | Pretrained RNA language model specialized for splicing analysis in vertebrates. | Genome Biology |
| PlantRNA-FM | An interpretable RNA foundation model for exploring functional RNA motifs in plants | Interpretable RNA foundation model for plant biology trained on 1,124+ plant species. | Nature Machine Intelligence |
| AllSplice | Perturbation-aware predictive modeling of RNA splicing using bidirectional transformers | Perturbation-aware bidirectional transformer for RNA splicing prediction. | bioRxiv |
| HydraRNA | HydraRNA: A hybrid architecture based full-length RNA language model | Hybrid-architecture full-length RNA language model combining multiple architecture strengths for processing long RNA sequences. | Genome Biology |
| EVA-RNA | EVA-RNA: A Scaling Cross-Species Transcriptomic Foundation Model for Immunology & Inflammation | Cross-species transcriptomic foundation model trained on 500K+ human and mouse samples for immunology and inflammation research. | OpenReview |
| GRNFormer | GRNFormer: A Biologically-Guided Framework for Integrating Gene Regulatory Networks into RNA Foundation Models | Biologically-guided framework integrating gene regulatory networks into RNA foundation model training. | Findings of the Association for Computational Linguistics: ACL 2025 |
| BMFM-RNA | BMFM-RNA: Whole-cell expression decoding improves transcriptomic foundation models | IBM biomedical foundation model RNA module using whole-cell expression decoding to enhance transcriptomic models. | arXiv (IBM) |
| CodonFM | CodonFM: Foundation Models for Codons | Codon foundation model trained on 130 million protein-coding sequences across 20,000+ species by NVIDIA and Arc Institute. | GitHub |
| Orthrus | Orthrus: Evolutionary and Functional RNA Foundation Models | Mamba-based RNA foundation model pretrained with biologically-augmented contrastive learning. | bioRxiv |
Deep learning methods for RNA 2D and 3D structure prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| RhoFold+ | Accurate RNA 3D structure prediction using a language model-based deep learning approach | Pretrained RNA language model-based approach for accurate RNA 3D structure prediction trained on 23.7 million sequences. | Nature Methods |
| RNA-FrameFlow | Flow Matching for de novo 3D RNA Backbone Design | SE(3) flow matching method for de novo 3D RNA backbone structure generation. | arXiv |
| DRfold2 | Ab initio RNA structure prediction with composite language model | Deep learning framework for ab initio RNA 3D structure prediction using a composite language model. | bioRxiv |
| 3DRNALM | Accurate RNA 3D structure prediction using a language model-based framework | Language model-based framework for accurate RNA 3D structure prediction. | Nature Communications |
| NuFold | NuFold: end-to-end approach for RNA tertiary structure prediction | End-to-end deep learning model for RNA tertiary structure prediction. | Nature Communications |
Generative models for RNA therapeutic sequence design and structure co-design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| RNAGenesis | RNAGenesis: A Generalist Foundation Model for Functional RNA Therapeutics | Generalist RNA therapeutics foundation model unifying sequence representation, structure prediction, and functional de novo design. | bioRxiv |
| GEMORNA | Deep generative models design mRNA sequences with enhanced translational capacity and stability | Deep generative model designing mRNA sequences with enhanced translational capacity and stability. | Science |
| EVA | A Long-Context Generative Foundation Model Deciphers RNA Design Principles | 1.4-billion-parameter MoE generative foundation model trained on 114 million full-length RNA sequences for long-context RNA design. | bioRxiv |
| RNACG | RNACG: A Universal RNA Sequence Conditional Generation model based on Flow-Matching | Universal RNA sequence conditional generation model based on flow matching. | arXiv |
| RiboFlow | RiboFlow: Conditional De Novo RNA Co-Design via Synergistic Flow Matching | Synergistic flow matching framework for RNA sequence-structure co-design targeting ligand binding. | NeurIPS |
| RiboGen | RiboGen: RNA Sequence and Structure Co-Generation with Equivariant MultiFlow | Equivariant multi-flow model simultaneously generating RNA sequences and full-atom 3D structures. | ICLR |
| SANDSTORM | Generative and predictive neural networks for the design of functional RNA molecules | Neural network for RNA function prediction integrating sequence and secondary structure data. | Nature Communications |
| GARDN | Generative and predictive neural networks for the design of functional RNA molecules | Generative neural network for functional RNA design, paired with SANDSTORM for prediction. | Nature Communications |
| mRNA-GPT | Large generative mRNA language foundation model for efficient coding sequence generation and design | 302M-parameter GPT-2-based mRNA generative language model for coding sequence generation across three biological domains. | bioRxiv |
| codonGPT | codonGPT: reinforcement learning on a generative language model enables scalable mRNA design | Reinforcement learning plus generative language model for scalable mRNA codon optimization design. | Nucleic Acids Research |
| RiboDecode | Deep generative optimization of mRNA codon sequences for enhanced mRNA translation and therapeutic efficacy | Deep generative optimization framework for mRNA codon sequences enhancing translation efficiency and therapeutic efficacy. | Nature Communications |
Foundation models for DNA sequence understanding, variant effect prediction, and gene regulation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Evo | Sequence modeling and design from molecular to genome scale with Evo | Arc Institute DNA foundation model processing molecular-to-genome-scale sequences (>650K tokens). | Science |
| Evo 2 | Genome modelling and design across all domains of life with Evo 2 | DNA foundation model trained on 9 trillion base pairs covering all domains of life with 1-million-token context window. | Nature |
| DNABERT | DNABERT: Pre-trained bidirectional encoder representations from transformers model for DNA-language in genome | First BERT-based pretrained model for genomic DNA sequences using k-mer tokenization. | Bioinformatics |
| DNABERT-2 | DNABERT-2: Efficient foundation model and benchmark for multi-species genome | Upgraded DNABERT using BPE tokenization instead of k-mer for multi-species genome analysis. | ICLR 2024 |
| Nucleotide Transformer | Nucleotide Transformer: Building and evaluating robust foundation models for human genomics | Large-scale genomic foundation model (50M–2.5B parameters) from InstaDeep trained on 3,200+ human genomes. | Nature Methods |
| HyenaDNA | HyenaDNA: Long-range genomic sequence modeling at single nucleotide resolution | Hyena implicit convolution-based genomic model for single-nucleotide-resolution long-range modeling up to 1 million bp. | NeurIPS 2023 |
| Enformer | Effective gene expression prediction from sequence by integrating long-range interactions | Transformer model from DeepMind/Calico predicting gene expression and chromatin states from DNA sequences. | Nature Methods |
| Caduceus | Caduceus: Bi-directional equivariant long-range DNA sequence modeling | Bidirectional Mamba-based DNA language model supporting reverse complement equivariance. | ICML |
| GenSLMs | GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics | Genome-scale language model pretrained on 110 million prokaryotic gene sequences analyzing SARS-CoV-2 evolution (Gordon Bell Prize). | IJHPCA |
| GROVER | DNA language model GROVER learns sequence context in the human genome | DNA language model trained on the human genome using BPE to define DNA vocabulary and capture CpG methylation features. | Nature Machine Intelligence |
| Sei | A sequence-based global map of regulatory activity for deciphering human genetics | Deep learning framework predicting 21,900+ chromatin features and mapping sequences to 40 regulatory activity classes. | Nature Genetics |
| GPN | DNA language models are powerful predictors of genome-wide variant effects | Unsupervised DNA language model predicting genome-wide variant effects from genomic sequences. | PNAS |
| Borzoi | Borzoi decodes the complex DNA signals governing gene regulation | Deep learning model predicting RNA-seq coverage from DNA sequences to decode gene regulatory signals. | Nature Genetics |
| PlantCaduceus | Cross-species modeling of plant genomes at single-nucleotide resolution using a pretrained DNA language model | Plant-specific DNA language model trained on 16 angiosperm genomes supporting cross-species analysis. | PNAS |
| HybriDNA | HybriDNA: A hybrid Transformer-Mamba2 DNA language model | Hybrid Transformer-Mamba2 DNA language model supporting ultra-long sequences (131kb) at single-nucleotide resolution. | arXiv |
| Nucleotide Transformer v3 (NTv3) | A foundational model for joint sequence-function multi-species prediction | Multi-species long-range genomic prediction and functional annotation foundation model from InstaDeep. | bioRxiv |
| Basenji | Sequential regulatory activity prediction across chromosomes with convolutional neural networks | CNN for predicting gene expression and regulatory activity from DNA sequences. | Genome Research |
| Basenji2 | Cross-species regulatory sequence activity prediction | Cross-species DNA regulatory activity prediction model. | PLoS Computational Biology |
| ChromBPNet | ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility | Base-resolution deep learning model for chromatin accessibility prediction with bias factorization. | bioRxiv |
| scBasset | scBasset: Sequence-based modeling of single-cell ATAC-seq using convolutional neural networks | CNN for modeling single-cell ATAC-seq chromatin accessibility from DNA sequences. | Nature Methods |
| EpiGePT | EpiGePT: a Pretrained Transformer model for epigenomics | Pretrained transformer model for epigenomics data analysis and prediction. | bioRxiv |
| AgroNT | A foundational large language model for edible plant genomes | Crop-specific genomic foundation model for plant genomics and breeding applications. | Communications Biology |
| AlphaGenome | Advancing regulatory variant effect prediction with AlphaGenome | Megabase-scale DNA model from Google DeepMind predicting gene expression and chromatin signals for variant effect analysis. | Nature |
| GENA-LM | GENA-LM: A Family of Open-Source Foundational DNA Language Models for Long Sequences | Open-source family of transformer DNA language models supporting up to 36k bp sequences. | bioRxiv/Bioinformatics |
| DNAGPT | DNAGPT: A Generalized Pre-trained Tool for Multiple DNA Sequence Analysis Tasks | Generative pretrained model for multiple DNA analysis tasks. | bioRxiv/PLoS ONE |
| MoDNA | MoDNA: Motif-Oriented Pre-training For DNA Language Model | Motif-oriented pretrained DNA language model capturing regulatory motif patterns. | ACM BCB |
| GPN-MSA | GPN-MSA: An alignment-based DNA language model for genome-wide variant effect prediction | DNA language model using multi-species alignment for genome-wide variant effect prediction. | Nature Biotechnology |
| MergeDNA | MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging | Hierarchical dynamic tokenization approach for context-aware genome modeling. | AAAI |
| BMFM-DNA | BMFM-DNA: A SNP-aware DNA foundation model to capture variant effects | IBM SNP-aware DNA foundation model for capturing variant effects in genomic sequences. | arXiv |
| EpiAgent | EpiAgent: foundation model for single-cell epigenomics | Foundation model for single-cell ATAC-seq epigenomic data analysis. | Nature Methods |
| Gene42 | Gene42: Long-Range Genomic Foundation Model With Dense Attention | Decoder-only long-range genomic foundation model processing up to 192,000 bp at single-nucleotide resolution. | arXiv |
| dnaHNet | dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning | Scalable hierarchical foundation model for genomic sequence learning. | arXiv |
| JEPA-DNA | JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures | Genomic foundation model based on joint-embedding predictive architecture combining generative and discriminative objectives. | arXiv |
| GeneZip | GeneZip: Region-Aware Compression for Long Context DNA Modeling | Region-aware compression method for long-context DNA sequence modeling. | arXiv |
| ModernGENA | ModernGENA: A modernized BERT-style DNA foundation model | Modernized BERT architecture adapted for DNA foundation modeling (ModernBERT for genomics). | OpenReview |
| OpticalDNA | OpticalDNA: Reimagining DNA sequence analysis as an OCR task | Novel framework reimagining DNA sequence analysis as an optical character recognition task. | arXiv |
| OmniReg-GPT | OmniReg-GPT: A Generative Pre-trained Model for Universal Gene Regulation Prediction | Generative pretrained model for universal gene regulation prediction across species and tissues. | Nature Communications |
| BOTANIC-0 | BOTANIC-0: A Plant Genomic Foundation Model | Family of plant genomic foundation models (100M–1B parameters) pretrained on 1,600+ curated plant genomes. | bioRxiv |
| Species-aware DNA LM | Species-aware DNA Language Modeling | DNA language model incorporating species-specific information during pretraining. | bioRxiv |
| Genos | Genos: A Large Human-Centric Genomic Foundation Model | Large-scale (up to 10B parameters) human-centric genomic foundation model with MoE-Transformer architecture from BGI. | GigaScience |
| GENERator | GENERator: A Long-Context Generative Genomic Foundation Model | Long-context generative genomic foundation model pretrained on 386 billion nucleotides with 98k context length. | arXiv |
| BioReason | BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model | First deep integration of DNA foundation models (Nucleotide Transformer/Evo2) with LLMs for multi-step biological reasoning; raises KEGG disease pathway prediction from 86% to 98% with interpretable reasoning traces. | NeurIPS 2025 |
Foundation models for single-cell transcriptomics, perturbation prediction, and virtual cell modeling.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| scGPT | scGPT: toward building a foundation model for single-cell multi-omics using generative AI | Generative pretrained transformer for single-cell multi-omics, enabling cell annotation, perturbation prediction, and gene network inference from 33M+ cells. | Nature Methods |
| scBERT | scBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data | Large-scale BERT-based pretrained model for automated cell type annotation from scRNA-seq data. | Nature Machine Intelligence |
| Geneformer | Transfer learning enables predictions in network biology | Transformer model pretrained on ~30 million single-cell transcriptomes for transfer learning in gene network biology. | Nature |
| scFoundation | Large-scale foundation model on single-cell transcriptomics | 100M-parameter single-cell transcriptomics foundation model (xTrimoGene) trained on 50M+ human cells. | Nature Methods |
| UCE | Universal Cell Embeddings: A foundation model for cell biology | Universal cell embedding model that creates a unified representation space across species and tissues. | bioRxiv |
| SCimilarity | A cell atlas foundation model for scalable search of similar human cells | Deep metric-learning foundation model for single-cell profiles, enabling rapid similarity search and annotation across a 23.4M-cell human atlas from 412 scRNA-seq studies. | Nature |
| GeneCompass | GeneCompass: Deciphering universal gene regulatory mechanisms with a knowledge-informed cross-species foundation model | Knowledge-enhanced cross-species foundation model trained on 100M+ human and mouse cells for deciphering gene regulatory mechanisms. | Cell Research |
| CellPLM | CellPLM: Pre-training of cell language model beyond single cells | Cell language model that integrates gene-gene and cell-cell interactions, going beyond single-cell level pretraining. | ICLR 2024 |
| tGPT | Generative pretraining from large-scale transcriptomes for single-cell deciphering | Generative pretrained model on 22.3 million single-cell transcriptomes for cell deciphering and clinical translation. | iScience |
| CellFM | CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells | 800M-parameter foundation model pretrained on 100 million human cell transcriptomes. | Nature Communications |
| Nicheformer | Nicheformer: a foundation model for single-cell and spatial omics | Single-cell and spatial omics foundation model trained on SpatialCorpus-110M (110M+ cells), capturing spatial microenvironments. | Nature Methods |
| scMulan | scMulan: A multitask generative pre-trained language model for single-cell analysis | Multitask generative pretrained language model that encodes cells as structured "c-sentences" for single-cell analysis. | RECOMB 2024 |
| Cell2Sentence | Cell2Sentence: Teaching large language models the language of biology | Converts gene expression profiles into natural language sentences, enabling GPT-2 adaptation for single-cell transcriptomics. | ICML 2024 |
| C2S-Scale | Scaling large language models for next-generation single-cell analysis | Scales the Cell2Sentence framework with larger models and broader data for improved single-cell RNA-seq analysis. | bioRxiv |
| GET | GET: A foundation model of transcription across human cell types | Universal expression transformer predicting gene expression from chromatin accessibility across 213 human cell types. | Nature |
| GenePT | GenePT: A simple but effective foundation model for genes and cells built from ChatGPT | Simple and effective gene/cell foundation model using ChatGPT embeddings for gene and cell representation. | bioRxiv |
| LangCell | LangCell: Language-cell pre-training for cell identity understanding | Joint pretraining of natural language and single-cell transcriptomics to enhance cell identity understanding. | arXiv |
| CellVQ | Illuminating cell states by a comprehensive and interpretable single cell foundation model | Comprehensive and interpretable single-cell foundation model trained on 68 million cells for illuminating cell states. | Nature Communications |
| scKGBERT | scKGBERT: A knowledge-enhanced foundation model for single-cell transcriptomics | Knowledge graph-enhanced foundation model for single-cell transcriptomics. | Genome Biology |
| SATURN | Toward universal cell embeddings: integrating scRNA-seq datasets across species | Cross-species universal cell embedding framework that integrates scRNA-seq datasets across organisms. | Nature Methods |
| scTab | scTab: Scaling cross-tissue single-cell annotation models | Scalable deep learning model for cross-tissue cell type annotation trained on 22 million cells. | Nature Communications |
| scPRINT | scPRINT: pre-training on 50 million cells allows robust gene network predictions | Transformer foundation model trained on 50M cells for robust gene network inference. | Nature Communications |
| scPRINT-2 | scPRINT-2: Towards the next-generation of cell foundation models | Next-generation cell foundation model trained on 350M cells across 16 organisms. | bioRxiv |
| scELMo | scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis | Language model embeddings applied to single-cell data analysis. | bioRxiv |
| scVI | scVI: Variational Inference for Single-Cell Gene Expression | Deep generative model providing probabilistic framework for single-cell transcriptomics analysis. | Nature Methods |
| scANVI | scANVI: semi-supervised integration of single-cell multi-omic data | Semi-supervised deep generative model for integrating single-cell multi-omic data. | Molecular Systems Biology |
| totalVI | Joint probabilistic modeling of single-cell multi-omic data with totalVI | Joint probabilistic model for simultaneous RNA and protein single-cell data analysis. | Nature Methods |
| scPoli | Population-level integration of single-cell datasets enables multi-scale analysis | Population-level single-cell dataset integration enabling multi-scale biological analysis. | Nature Methods |
| scHyena | scHyena: Foundation Model for Full-Length Single-Cell RNA-Seq Analysis in Brain | Hyena architecture-based foundation model for full-length scRNA-seq analysis in brain tissue. | arXiv |
| TOSICA | Transformer for one stop interpretable cell type annotation | Transfer learning framework for single-cell omics analysis across datasets and modalities. | Nature Communications |
| xTrimoGene | xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data | Efficient and scalable representation learner for scRNA-seq data using asymmetric encoder-decoder architecture. | NeurIPS 2023 |
| CancerFoundation | A single-cell RNA sequencing foundation model to decipher drug resistance in cancer | Cancer-specific scRNA-seq foundation model for deciphering drug resistance mechanisms. | bioRxiv |
| Cell-GraphCompass | Cell-GraphCompass: Modeling Single Cells with Graph Structure Foundation Model | Graph structure foundation model for single-cell analysis using graph-based cell representations. | National Science Review |
| scLong | scLong: A billion-parameter foundation model for capturing long-range gene context | Billion-parameter foundation model attending to all 28,000 genes simultaneously for long-range context. | Nature Communications |
| Tahoe-x1 | Tahoe-x1: Scaling Perturbation-Trained Single-Cell Foundation Models to 3 Billion Parameters | 3B-parameter perturbation-trained single-cell foundation model for predicting cellular responses. | bioRxiv |
| PULSAR | PULSAR: a Foundation Model for Multi-scale and Multicellular Biology | Multi-scale foundation model integrating 36M+ cells for multicellular biology analysis. | bioRxiv |
| TranscriptFormer | A Cross-Species Generative Cell Atlas Across 1.5 Billion Years of Evolution | Cross-species generative cell atlas foundation model spanning 1.5 billion years of evolution. | bioRxiv |
| TCRfoundation | TCRfoundation: A multimodal foundation model for single-cell immune profiling | Multimodal foundation model integrating gene expression with TCR sequences for immune profiling. | GitHub |
| CELLama | CELLama: Foundation Model for Single Cell and Spatial Transcriptomics | Cell embedding model leveraging language model capabilities for single-cell and spatial transcriptomics. | bioRxiv |
| scPROTEIN | scPROTEIN: versatile deep graph contrastive learning framework for single-cell proteomics | Graph contrastive learning framework for single-cell proteomics embedding and analysis. | Nature Methods |
| TEDDY | TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology | Family of foundation models designed for comprehensive single-cell biology understanding. | ICML Workshop |
| Tabula | Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell Transcriptomics | Tabular self-supervised foundation model tailored for single-cell transcriptomics data. | NeurIPS 2025 |
| ChromFound | ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibility Data | Universal foundation model for single-cell chromatin accessibility (scATAC-seq) data analysis. | NeurIPS |
| EpiFoundation | EpiFoundation: A Foundation Model for Single-Cell ATAC-seq via Peak-to-Gene Alignment | Foundation model for scATAC-seq using peak-to-gene alignment for epigenomic analysis. | bioRxiv |
| SCARF | SCARF: Single Cell ATAC-seq and RNA-seq Foundation model | Multi-modal foundation model jointly modeling scATAC-seq and scRNA-seq data. | bioRxiv |
| CAPTAIN | CAPTAIN: A multimodal foundation model pretrained on co-assayed single-cell RNA and protein | Multimodal foundation model for co-assayed scRNA and protein data integration. | bioRxiv |
| OKR-Cell | OKR-Cell: Open world knowledge aided single-cell foundation model with cross-modal pre-training | Open-world knowledge-enhanced single-cell foundation model with cross-modal pretraining. | arXiv |
| GeneJepa | GeneJepa: A predictive world model of the transcriptome | Predictive world model for transcriptomics based on JEPA architecture. | arXiv |
| VCWorld | VCWorld: Biological world model for virtual cell simulation | Biological world model for simulating virtual cell dynamics and perturbation responses. | arXiv |
| CellHermes | CellHermes: Harmonizing multimodal data for omics understanding | Multimodal data harmonization model for unified omics understanding. | arXiv |
| CellTok | CellTok: Early-fusion multimodal LLM for single-cell transcriptomics via tokenization | Early-fusion multimodal LLM for single-cell transcriptomics using gene expression tokenization. | arXiv |
| sciLaMA | sciLaMA: Single-cell representation learning leveraging prior knowledge from LLMs | Single-cell representation learning leveraging prior knowledge from large language models. | arXiv |
| scConcept | scConcept: Contrastive pretraining for technology-agnostic single-cell representations | Contrastive pretraining framework for technology-agnostic single-cell representations. | arXiv |
| scLinguist | scLinguist: Hyena-based foundation model for cross-modality translation in single-cell multi-omics | Hyena-based foundation model for cross-modality translation in single-cell multi-omics. | arXiv |
| scNET | scNET: Context-specific gene and cell embeddings by integrating scRNA with PPI | Context-specific gene and cell embeddings integrating scRNA-seq with protein-protein interaction networks. | arXiv |
| scLAMBDA | scLAMBDA: Modeling single-cell multi-gene perturbation responses | Model for predicting single-cell responses to multi-gene combinatorial perturbations. | arXiv |
| GeneMamba | GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data | Mamba architecture-based efficient single-cell foundation model with scalable computation. | arXiv |
| Atacformer | Atacformer: Transformer-based foundation model for ATAC-seq data analysis | Transformer-based foundation model for ATAC-seq chromatin accessibility data analysis. | arXiv |
| CLM-X | CLM-X: Cross-Modal Language Model for Single-Cell Multi-Omics | Cross-modal language model for unified single-cell multi-omics representation learning. | arXiv |
| CellOracle | CellOracle: Dissecting cell identity via network inference and in silico gene perturbation | Computational framework using gene regulatory networks to simulate gene perturbation effects on cell identity. | Nature |
| Stack | Stack: In-Context Learning of Single-Cell Biology | Arc Institute single-cell foundation model trained on 149M human cells enabling zero-shot prediction via in-context learning. | bioRxiv |
| Lingshu-Cell | Lingshu-Cell: cellular world model for transcriptome modeling | Masked discrete diffusion cellular world model for transcriptome modeling from Alibaba DAMO. | arXiv |
| OmniCell | OmniCell: Unified Foundation Modeling of Single-Cell and Spatial Transcriptomics | Unified foundation model for both single-cell and spatial transcriptomics analysis. | bioRxiv |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AlphaCell | Towards building a World Model to simulate perturbation-induced cellular dynamics | Virtual cell world model that simulates perturbation-induced cellular dynamics. | bioRxiv |
| CellFluxV2 | CellFluxV2: An Image Generative Foundation Model for Virtual Cell Modeling | Flow matching-based generative foundation model for virtual cell image modeling. | bioRxiv |
| X-Cell | X-Cell: Scaling Causal Perturbation Prediction Across Diverse Cellular Contexts | Large-scale diffusion language model predicting genome-wide transcriptional responses across diverse cellular contexts. | bioRxiv |
Foundation models that integrate molecular, cellular, and tissue-level information across biological scales.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Xpressor | Towards foundation models that learn across biological scales | Cross-scale learning framework integrating molecular, cellular, and tissue-level gene expression via cross-attention. | bioRxiv |
| AIDO | Toward AI-driven digital organism: Multiscale foundation models for predicting, simulating and programming biology at all levels | AI-driven digital organism system integrating DNA→RNA→protein→cell multi-scale foundation models. | arXiv |
| CDT (Central Dogma Transformer) | Central Dogma Transformer: Towards Mechanism-Oriented AI for Cellular Understanding | Architecture integrating pretrained DNA (Enformer), RNA (scGPT), and protein (ProteomeLM) models via directional cross-attention mirroring the central dogma information flow, producing unified Virtual Cell Embeddings. | arXiv |
| CDT-II | Central Dogma Transformer II: An AI Microscope for Understanding Cellular Regulatory Mechanisms | AI microscope with DNA/RNA self-attention and cross-attention for transcriptional control; achieves per-gene mean r=0.84 on K562 CRISPRi data, recovers GFI1B regulatory network (6.6× enrichment), and predicts therapeutic target consequences via gradient attribution. | arXiv |
| CDT-III | Central Dogma Transformer III: Interpretable AI Across DNA, RNA, and Protein | Two-stage Virtual Cell Embedder (VCE-N for nuclear transcription, VCE-C for cytosolic translation) extending to full central dogma with protein prediction; achieves RNA r=0.843 and protein r=0.969, rediscovers 5/7 known Alemtuzumab side effects without clinical data. | arXiv |
Foundation models for antibody engineering, structure prediction, and immune receptor analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| IgBERT | Large scale paired antibody language models | BERT-based antibody language model trained on 2B+ unpaired and 2M paired antibody sequences from OAS. | PLOS Computational Biology |
| IgT5 | Large scale paired antibody language models | T5-based paired antibody language model companion to IgBERT for antibody design and engineering. | PLOS Computational Biology |
| AntiBERTy | Deciphering antibody affinity maturation with language models and weakly supervised learning | BERT-based model trained on 558M antibody sequences for affinity maturation analysis. | arXiv |
| AbLang | AbLang: an antibody language model for completing antibody sequences | Antibody-specific language model trained on OAS for residue prediction and antibody representation. | Bioinformatics Advances |
| AbLang2 | Addressing the antibody germline bias and its effect on language models for improved antibody design | Improved antibody language model for paired heavy-light chains with reduced germline bias. | Bioinformatics |
| IGLOO | Tokenizing Loops of Antibodies | Multimodal antibody loop tokenizer enhancing protein language models for antibody research. | NeurIPS 2025 Workshop |
| Ab-RoBERTa | Antibody Foundational Model : Ab-RoBERTa | RoBERTa-based antibody language model for paratope prediction and antibody design. | arXiv |
| BALM | Accurate Prediction of Antibody Function and Structure Using Bio-Inspired Antibody Language Model | Bio-inspired antibody language model for predicting antibody structure and function. | Science Advances |
| DASM | Separating selection from mutation in antibody language models | Deep amino acid selection model that separates selection from mutation in antibody sequence modeling. | eLife |
| nanoBERT | nanoBERT: a deep learning model for gene agnostic navigation of the nanobody mutational space | Nanobody-specific transformer model for predicting amino acid substitutions in VHH sequences. | Bioinformatics Advances |
| FAbCon | A generative foundation model for antibody sequence understanding | 2.4B-parameter generative foundation model for antibody sequence understanding. | bioRxiv |
| S2ALM | S2ALM: Sequence-Structure Pre-trained Large Language Model for Antibody | Sequence-structure pretrained large language model for comprehensive antibody understanding. | Research |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| IgFold | Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies | Fast antibody structure prediction using pretrained LM on 558M sequences with graph neural networks. | Nature Communications |
| DeepAb | Antibody structure prediction using interpretable deep learning | Interpretable deep learning model for antibody Fv structure prediction specializing in CDR loop modeling. | Patterns |
| ABlooper | ABlooper: fast accurate antibody CDR loop structure prediction with accuracy estimation | Rapid equivariant neural network for antibody CDR loop structure prediction with accuracy estimation. | Bioinformatics |
| AntiFold | AntiFold: Improved structure-based antibody design using inverse folding | Antibody-specific inverse folding model fine-tuned from ESM-IF1 for CDR sequence generation from structures. | Bioinformatics Advances |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| IgGM | A generative foundation model for antibody design | Generative foundation model for comprehensive antibody design. | bioRxiv |
| DiffAb | Antigen-Specific Antibody Design and Optimization with Diffusion-Based Generative Models | Diffusion-based generative model for antigen-specific antibody CDR-H3 design, jointly modeling sequence, structure, and orientation. | NeurIPS 2022 |
| dyMEAN | Full-Atom Antibody Design via dyMEAN | End-to-end full-atom antibody design using dynamic multi-channel equivariant graph network. | ICML |
| MEAN | Conditional Antibody Design as 3D Equivariant Graph Translation | 3D equivariant graph neural network for conditional antibody CDR sequence-structure co-design. | ICLR 2023 |
| RefineGNN | Iterative Refinement Graph Neural Network for Antibody Sequence-Structure Co-design | Iterative refinement GNN for antibody CDR co-design of sequence and 3D structure via autoregressive generation. | ICLR |
| Ophiuchus-Ab | Ophiuchus-Ab: A Versatile Generative Foundation Model for Advanced Antibody-Based Immunotherapy | Diffusion language model for antibody immunotherapy and paired antibody repertoire generation. | bioRxiv |
| NanoAbLLaMA | NanoAbLLaMA: construction of nanobody libraries with protein large language models | LLaMA2-based language model fine-tuned for nanobody (VHH) library construction and design. | Frontiers in Chemistry |
| CoSiNE | CoSiNE: Conditionally Site-Independent Neural Evolution of Antibody Sequences | Conditionally site-independent neural evolution model explicitly modeling antibody affinity maturation. | arXiv |
| AbBFN2 | AbBFN2: A flexible antibody foundation model based on Bayesian Flow Networks | Flexible antibody foundation model based on Bayesian flow networks for multi-objective unified modeling. | bioRxiv |
| AntibodyDesignBFN | AntibodyDesignBFN: High-Fidelity Fixed-Backbone Antibody Design via Discrete Bayesian Flow Networks | High-fidelity fixed-backbone antibody design using discrete Bayesian flow networks. | arXiv |
| AbAffinity | AbAffinity: A Large Language Model for Predicting Antibody Binding Affinity | Large language model for predicting antibody-antigen binding affinity. | arXiv |
| CALM | CALM: Cross-attention Adaptive Immune Receptor–Antigen Language Model | Cross-attention language model for antibody-antigen specificity prediction. | bioRxiv |
| JAM-2 | JAM-2: Fully computational design of drug-like antibodies | Fully computational model for designing drug-like antibodies developed by Nabla Bio. | Technical Report |
| Chai-2 | Chai-2: Zero-shot antibody discovery | Zero-shot antibody discovery model developed by Chai Discovery. | bioRxiv |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| TCR-BERT | TCR-BERT: learning the grammar of T-cell receptors for flexible antigen-binding analyses | Modified BERT trained on TCR sequences via self-supervised learning for antigen-specificity prediction. | PMLR v240 |
| tcrLM | tcrLM: a lightweight protein language model for predicting T cell receptor and epitope binding specificity | Lightweight BERT-based LM pretrained on 100M+ TCR CDR3 sequences for TCR-epitope binding prediction. | arXiv |
| TCR-GPT | TCR-GPT: Integrating Autoregressive Model and Reinforcement Learning for T-Cell Receptor Repertoires Generation | Decoder-only transformer for TCR sequence generation using autoregressive modeling with reinforcement learning. | arXiv |
| SCEPTR | Contrastive learning of T cell receptor representations | Lightweight BERT-like transformer for TCR analysis using autocontrastive and masked-language pretraining. | Cell Systems |
| ERGO-II | Prediction of Specific TCR-Peptide Binding From Large Dictionaries of TCR-Peptide Pairs | Deep learning model (LSTM + autoencoder) for TCR-peptide binding prediction using NLP techniques. | Frontiers in Immunology |
| mvTCR | Multi-modal generative modeling for joint analysis of single-cell T cell receptor and gene expression data | Multimodal variational autoencoder integrating single-cell TCR sequences with gene expression data. | Nature Communications |
| NetTCR-2.0 | NetTCR-2.0 enables accurate prediction of TCR-peptide binding | Deep learning model for TCR-peptide-MHC binding prediction using paired TCRα and β sequences. | Communications Biology |
| TCR-TRANSLATE | Conditional generation of real antigen-specific T cell receptor sequences | ML framework for generating antigen-specific TCR sequences including for unseen epitopes. | Nature Machine Intelligence |
Foundation models for enzyme function prediction, kinetics modeling, and de novo enzyme design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| EnzyGen | Generative Enzyme Design Guided by Functionally Important Sites and Small-Molecule Substrates | Generative enzyme design model leveraging functional sites and substrate information. | arXiv |
| RFdiffusion2 | Atom-level enzyme active site scaffolding using RFdiffusion2 | Atom-level enzyme active site scaffolding tool based on RFdiffusion architecture. | Nature Methods |
| CLEAN | Enzyme function prediction using contrastive learning | Contrastive learning model for enzyme EC number prediction, outperforming BLAST and traditional methods. | Science |
| CLEAN-Contact | Improved enzyme functional annotation prediction using contrastive learning with structural inference | Extension of CLEAN integrating protein contact maps for improved enzyme function annotation. | Communications Biology |
| EnzBERT | Predicting enzymatic function of protein sequences with attention | BERT-based model for predicting enzyme EC numbers from protein sequences using attention mechanisms. | Bioinformatics |
| EnzymeFlow | EnzymeFlow: Generating Reaction-specific Enzyme Catalytic Pockets through Flow Matching and Co-Evolutionary Dynamics | Generative model using flow matching to design reaction-specific enzyme catalytic pockets. | NeurIPS |
| EnzymeCAGE | EnzymeCAGE: A Geometric Foundation Model for Enzyme Retrieval with Evolutionary Insights | Geometric foundation model trained on ~1M enzyme-reaction pairs for enzyme retrieval and function prediction. | bioRxiv |
| CatPred | CatPred: a comprehensive framework for deep learning in vitro enzyme kinetic parameters | Deep learning framework for predicting enzyme kinetic parameters (kcat, Km, Ki) from sequences. | Nature Communications |
| UniKP | UniKP: a unified framework for the prediction of enzyme kinetic parameters | Unified deep learning framework using pretrained protein LMs to predict kcat, Km, and catalytic efficiency. | Nature Communications |
| TurNuP | Turnover number predictions for kinetically uncharacterized enzymes using machine and deep learning | Deep learning model for predicting enzyme turnover numbers (kcat) for uncharacterized enzymes. | Nature Communications |
| EnzyControl | EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation | Framework for substrate-specific enzyme backbone generation with functional control. | arXiv |
| ProtDETR | Interpretable Enzyme Function Prediction via Residue-Level Detection | Attention-based framework for residue-level enzyme EC number prediction inspired by object detection. | arXiv |
| EZpred | EZpred: improving deep learning-based enzyme function prediction using unlabeled sequence homologs | Deep learning framework leveraging unlabeled homolog sequences for improved enzyme EC prediction. | bioRxiv |
| BEC-Pred | A general model for predicting enzyme functions based on enzymatic reactions | BERT-based model predicting enzyme EC numbers from SMILES representations of substrates and products. | Journal of Cheminformatics |
| HIT-EC | HIT-EC: Trustworthy prediction of enzyme commission numbers using a hierarchical interpretable transformer | Hierarchical interpretable transformer for trustworthy enzyme EC number prediction. | Nature Communications |
| ENZYME-UNIFIED | ENZYME-UNIFIED: Learning Holistic Representations of Enzyme Function with a Hybrid Interaction Model | Holistic enzyme function representation learning via hybrid interaction modeling. | OpenReview |
Foundation models for spatially resolved gene expression, tissue architecture, and histology-omics integration.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Novae | Novae: a graph-based foundation model for spatial transcriptomics data | Graph-based foundation model for spatial transcriptomics trained on 30 million cells. | Nature Methods |
| SpaFoundation | SpaFoundation: a visual foundation model for spatial transcriptomics | Visual foundation model for spatial transcriptomics using 1.84M histological images. | bioRxiv |
| STPath | STPath: a generative foundation model for integrating spatial transcriptomics and WSIs | Generative foundation model integrating spatial transcriptomics with whole-slide histology images. | npj Digital Medicine |
| OmiCLIP | A visual-omics foundation model to bridge histopathology with spatial transcriptomics | Visual-omics foundation model bridging histopathology images and spatial transcriptomics data. | Nature Methods |
| SpatialFusion | SpatialFusion: A lightweight multimodal foundation model for spatial transcriptomics | Lightweight multimodal foundation model integrating gene expression, histopathology, and pathway data. | bioRxiv |
| SAGE-FM | SAGE-FM: A lightweight and interpretable foundation model for spatial transcriptomics | GCN-based lightweight and interpretable foundation model for spatial transcriptomics. | arXiv |
| scGPT-spatial | scGPT-spatial: Continual Pretraining of Single-Cell FM for Spatial Transcriptomics | Extension of scGPT for spatial transcriptomics via continual pretraining. | bioRxiv |
| SpatialScope | SpatialScope: integrating spatial and single-cell transcriptomics data using deep generative models | Deep generative model for integrating spatial transcriptomics with scRNA-seq data. | Nature Communications |
| stFormer | stFormer: a foundation model for spatial transcriptomics | Transformer foundation model integrating ligand-receptor interactions into spatial gene representations. | bioRxiv |
| STAGE | STAGE: A Foundation Model for Spatial Transcriptomics Analysis via Graph Embeddings | Foundation model using graph embeddings and hierarchical prototypes for spatial transcriptomics. | OpenReview |
| STORM | STORM: A multimodal foundation model of spatial transcriptomics and histology | Multimodal spatial transcriptomics and histology foundation model trained on 1.2M spatially-resolved profiles across 18 organs. | arXiv |
| SEAL | SEAL: Spatial Expression-Aligned Learning for pathology foundation models | Spatial expression-aligned learning framework enhancing pathology foundation models with spatial transcriptomics data. | arXiv |
| MINT | MINT: Molecularly Informed Training with Spatial Transcriptomics Supervision for Pathology Foundation Models | Molecularly informed training with spatial transcriptomics supervision for pathology foundation models. | arXiv |
| HINGE | HINGE: Adapting Pre-trained Single-Cell Foundation Models to Spatial Gene Expression | Adapts pretrained single-cell foundation models to spatial gene expression using histological image conditioning. | arXiv |
| HEIST | HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data | Graph foundation model for both spatial transcriptomics and proteomics data analysis. | arXiv |
| TISSUENARRATOR | TISSUENARRATOR: Generative modeling of spatial transcriptomics with LLMs | LLM-based generative modeling framework for spatial transcriptomics data. | arXiv |
| SpaTranslator | SpaTranslator: Deep generative framework for universal spatial multi-omics cross-modality translation | Deep generative framework for universal cross-modality translation of spatial multi-omics data. | arXiv |
| SpatialProp | SpatialProp: Tissue perturbation modeling with spatially resolved single-cell transcriptomics | Tissue perturbation modeling using spatially resolved single-cell transcriptomics. | arXiv |
| SWITCH | SWITCH: Integrative deep learning of spatial multi-omics | Integrative deep learning framework for spatial multi-omics data analysis. | arXiv |
| CancerSTFormer | CancerSTFormer enables multi-scale analysis of spot-resolution spatial transcriptomes | Multi-scale spatial transcriptomics foundation model for cancer at 50µm and 250µm resolution. | bioRxiv |
| STAGATE | Deciphering spatial domains from spatially resolved transcriptomics with adaptive graph attention auto-encoder | Graph attention auto-encoder for spatial domain identification integrating gene expression and spatial location. | Nature Communications |
| CellViT | CellViT: Vision Transformers for precise cell segmentation and classification | Vision Transformer for precise cell/nuclei segmentation in H&E whole-slide images. | Medical Image Analysis |
| STAMP | Interpretable spatially aware dimension reduction of spatial transcriptomics with STAMP | Deep generative model for spatially-aware interpretable dimension reduction of spatial transcriptomics. | Nature Methods |
| SpaGT | Spatially informed graph transformers for spatially resolved transcriptomics | Graph transformer integrating spatial coordinates and gene expression for spatial domain identification. | Communications Biology |
| BrainBeacon | BrainBeacon: A Cross-Species Foundation Model for Single-cell Spatial Transcriptomics of Brain | Cross-species brain spatial transcriptomics foundation model integrating multi-species data for digital twin brain. | bioRxiv |
| SToFM | SToFM: A multi-scale foundation model for spatial transcriptomics | Multi-scale spatial transcriptomics foundation model integrating macroscopic tissue morphology and microscopic cellular environments. | ICML 2025 |
| OmniCell | OmniCell: Unified Foundation Modeling of Single-Cell and Spatial Transcriptomics | Unified foundation model for both single-cell and spatial transcriptomics analysis. | bioRxiv |
| PAST | PAST: A multimodal single-cell foundation model for histopathology and spatial transcriptomics in cancer | Multimodal single-cell foundation model integrating histopathology images and spatial transcriptomics data for cancer analysis. | arXiv |
Foundation models for glycan structure representation, protein-glycan interactions, and carbohydrate analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| GlycanGT | GlycanGT: A Foundation Model for Glycan Graphs with Pretrained Representation and Generative Learning | First glycan graph foundation model using graph transformer for glycan representation and generative learning. | bioRxiv |
| SweetBERT | Exploring BERT-based models for IUPAC glycan nomenclature | BERT-based glycan sequence language model encoding IUPAC glycan nomenclature and branching structures. | ICLR 2025 Workshop |
| GlycanAA | Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-training | All-atom glycan modeling framework using hierarchical message passing and multi-scale pretraining. | ICML |
| MCNet | Atom-level machine learning of protein-glycan interactions and cross-chiral recognition | Atom-level machine learning model for protein-glycan interaction prediction including mirror-image glycan recognition. | Science Advances |
| DeepGlycanSite | Highly accurate carbohydrate-binding site prediction with DeepGlycanSite | High-accuracy deep learning model for predicting carbohydrate-binding sites on proteins. | Nature Communications |
| SweetNet | Using graph convolutional neural networks to learn a representation for glycans | Graph convolutional neural network for glycan representation learning handling complex branching structures. | Cell Reports |
| GlycoBERT | Transformer-based Deep Learning for Glycan Structure Inference from MS/MS | BERT-based transformer for inferring glycan structures from tandem mass spectrometry data. | bioRxiv |
Foundation models for metabolomic profiling, spectral analysis, and multi-disease prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MetaboLM | MetaboLM: a metabolomic language model for multi-disease early prediction and risk stratification | Transformer-based metabolomics language model trained on ~84,000 healthy plasma metabolomes for multi-disease early prediction. | Nature Communications |
| DSCF | Deep spectral component filtering as a foundation model for spectral analysis demonstrated in metabolic profiling | Self-supervised deep spectral component filtering foundation model for metabolic profiling analysis. | Nature Machine Intelligence |
Foundation models for cryo-electron microscopy image processing, density map analysis, and structure refinement.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CryoFM | CryoFM: A Flow-based Foundation Model for Cryo-EM Densities | Flow matching-based foundation model learning high-quality biomolecular density map distributions. | NeurIPS 2024 / bioRxiv |
| Cryo-IEF | A comprehensive foundation model for cryo-EM image processing | Comprehensive cryo-EM image processing foundation model pretrained via contrastive learning on 65M particle images. | Nature Methods |
| CryoLVM | CryoLVM: Self-supervised Learning from Cryo-EM Density Maps | Self-supervised cryo-EM density map foundation model using JEPA architecture. | arXiv |
| CryoNet.Refine | CryoNet.Refine: A One-step Diffusion Model for Rapid Refinement of Structural Models with Cryo-EM Density Map Restraints | One-step diffusion model for rapid structural model refinement with cryo-EM density map constraints. | arXiv |
| CryoDRGN-AI | CryoDRGN-AI: neural ab initio reconstruction for cryo-EM | Neural ab initio reconstruction method for heterogeneous cryo-EM data. | Nature Methods |
Foundation models for metagenomic sequencing, microbiome analysis, and pathogen monitoring.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| METAGENE-1 | Metagenomic Foundation Model for Pandemic Monitoring | 7B-parameter autoregressive transformer trained on >1.5 trillion bases of metagenomic DNA/RNA for pathogen monitoring. | arXiv |
| MGM | MGM as a large-scale pretrained foundation model for microbiome analyses in diverse contexts | Large-scale microbiome foundation model trained on >263,000 microbiome samples for diverse contexts. | Advanced Science |
| BiomeGPT | BiomeGPT: A foundation model for the human gut microbiome | Transformer-based human gut microbiome foundation model trained on >13,300 metagenomic samples. | bioRxiv |
| GenomeOcean | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies | Efficient 4B-parameter genome foundation model trained on large-scale metagenomic assemblies (>600 Gbp). | bioRxiv |
| MicroGenomer | MicroGenomer: A Foundation Model for Transferable Microbial Genome Representations | Transferable microbial genome representation model trained on >234.5B base pairs for multi-scale analysis. | bioRxiv (BGI Research) |
| Generanno | Generanno: A Genomic Foundation Model for Metagenomic Annotation | Genomic foundation model for metagenomic annotation trained on 715B prokaryotic base pairs. | bioRxiv |
| FGBERT | FGBERT: Function-Driven Pre-trained Gene Language Model for Metagenomics | Function-driven pretrained gene language model using protein-level context-aware tokenizer for metagenomics. | arXiv |
| MetagenBERT | MetagenBERT: a Transformer Architecture using Foundational DNA Read Embedding Models for novel Metagenome Representation | Transformer framework using DNABERT-2/DNABERT-S for metagenome representation from raw DNA reads. | arXiv |
| Darwin-7B | Darwin-7B: A Multi-Omic Foundation Model for the Human Gut Microbiome via Sparsified Quality-Aware Tokenization | 7B-parameter multi-omic foundation model for the human gut microbiome, trained with sparsified quality-aware tokenization. | ICLR 2026 Workshop |
| ViraLM | ViraLM: virus discovery through genome foundation model | Virus genome foundation model for virus discovery from metagenomic sequences. | Bioinformatics |
Foundation models for phylogenetic tree inference and evolutionary genomic modeling.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Phyla | Evolutionary Reasoning Does Not Arise in Standard Usage of Protein Language Models | Phylogenetic inference foundation model (Phyla, originally proposed in v1 of the same preprint) using hybrid state-space transformer with tree loss function. | bioRxiv |
| PhyloGPN | A Phylogenetic Approach to Genomic Language Modeling | Phylogenetic tree-based genomic language model using multi-species whole-genome alignments and evolutionary models. | Lecture Notes in Computer Science |
Foundation models for integrating DNA, RNA, protein, and other multi-omic data modalities.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| OmniBioTE | Large-Scale Multi-omic Biosequence Transformers for Modeling Protein-Nucleic Acid Interactions | Large-scale multi-omic biosequence transformer trained on >250B tokens of protein and nucleic acid sequences. | PLOS ONE |
| Omni-DNA | Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning | Unified genomic foundation model supporting DNA/RNA/protein cross-modal multi-task learning from Microsoft. | NeurIPS 2025 (Microsoft) |
| OmniNA | OmniNA: A foundation model for nucleotide sequences | Nucleotide sequence foundation model pretrained on >91.7M sequences (>1 trillion bases) for cross-species understanding. | bioRxiv |
| spEMO | Leveraging multi-modal foundation models for analysing spatial multi-omic and histopathology data | Multi-modal foundation model framework integrating spatial multi-omics with histopathology image data. | Nature Biomedical Engineering |
| scMamba | scMamba: A Scalable Foundation Model for Single-Cell Multi-Omics Integration Beyond Highly Variable Feature Selection | Scalable Mamba-based foundation model for single-cell multi-omics integration without feature selection. | arXiv |
Foundation models for DNA methylation, chromatin modifications, and epigenetic regulation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CpGPT | CpGPT: a Foundation Model for DNA Methylation | DNA methylation foundation model predicting CpG site methylation states for aging and disease research. | bioRxiv |
| MethylGPT | MethylGPT: A Foundation Model for the DNA Methylome | DNA methylome foundation model pretrained on large-scale methylation data for epigenetic age prediction and cancer classification. | bioRxiv |
| scDNAm-GPT | scDNAm-GPT: A Foundation Model for Single-Cell DNA Methylation Analysis | Single-cell DNA methylation analysis foundation model for resolving epigenetic heterogeneity at single-cell resolution. | bioRxiv |
Foundation models for mass spectrometry-based proteomics, metabolomics, and compound identification.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| DIA-BERT | DIA-BERT: pre-trained end-to-end transformer models for enhanced DIA proteomics data analysis | Transformer-based foundation model for data-independent acquisition proteomics, improving peptide identification and quantification. | Nature Communications |
| DreaMS | DreaMS: Deep Representations Empowering the Annotation of Mass Spectra | Deep representation learning foundation model for mass spectra annotation and metabolite identification. | Nature Biotechnology |
| LSM-MS2 | LSM-MS2: Large-Scale Mass Spectrometry Foundation Model | Large-scale tandem mass spectrometry foundation model pretrained on millions of MS2 spectra for compound identification. | ChemRxiv |
| OmniNovo | OmniNovo: A Universal Foundation Model for De Novo Peptide Sequencing | Universal foundation model for de novo peptide sequencing directly from mass spectrometry data. | arXiv |
| MS-FM | Foundation model for mass spectrometry proteomics | Unified mass spectrometry proteomics foundation model pretrained on de novo sequencing data. | arXiv |
| InstaNovo | InstaNovo: diffusion-powered de novo peptide sequencing | Diffusion-powered model for de novo peptide sequencing from mass spectrometry data. | Nature Machine Intelligence |
Foundation models for neural activity prediction, brain imaging, and computational neuroscience.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| VFAM | Foundation model of neural activity predicts response to new stimulus types | Neural activity foundation model trained on large-scale mouse visual cortex data, predicting responses to novel stimulus types. | Nature |
Foundation models for designing synthetic regulatory elements and engineering biological systems.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| DNA-Diffusion | Designing synthetic regulatory elements using DNA-Diffusion | Generative diffusion model for designing synthetic DNA regulatory elements. | Nature Genetics |
Foundation models for molecular property prediction, representation learning, and chemical language modeling on SMILES and molecular graphs.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MoLFormer | Large-Scale Chemical Language Representations Capture Molecular Structure and Properties | Transformer-based chemical language model pretrained on 1.1B SMILES with linear attention and rotary embeddings for molecular property prediction. | Nature Machine Intelligence |
| GP-MoLFormer | GP-MoLFormer: A Foundation Model For Molecular Generation | Transformer-based generative foundation model with 46.8M parameters trained on 1.1B SMILES for molecular generation tasks. | Digital Discovery |
| ChemBERTa | ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction | RoBERTa-based chemical language model pretrained on 10M PubChem SMILES for molecular property prediction. | NeurIPS ML4Molecules Workshop |
| ChemBERTa-2 | ChemBERTa-2: Towards Chemical Foundation Models | Evolved ChemBERTa pretrained on 77M PubChem SMILES with optimized pretraining strategies including multi-task regression. | arXiv |
| ChemFM | ChemFM as a Scaling Law Guided Foundation Model Pre-trained on Informative Chemicals | 3B-parameter chemical foundation model trained on 178M UniChem molecules using self-supervised causal language modeling guided by scaling laws. | Communications Chemistry |
| ChemDFM | Developing ChemDFM as a Large Language Foundation Model for Chemistry | LLaMA-13B-based chemistry LLM trained on 34B tokens of chemical literature and fine-tuned with 2.7M instruction pairs. | Cell Reports Physical Science |
| Uni-Mol | Uni-Mol: A Universal 3D Molecular Representation Learning Framework | Universal 3D molecular representation learning framework that directly leverages molecular 3D structures for pretraining and achieves SOTA on property prediction. | ICLR |
| Uni-Mol2 | Uni-Mol2: Exploring Molecular Pretraining Model at Scale | Largest 3D molecular foundation model (1.1B parameters) with dual-track Transformer integrating atomic, graph, and 3D geometric features trained on 884M molecules. | NeurIPS |
| MolBERT | MolBERT: Molecular Representation Learning with Language Models and Domain-Relevant Auxiliary Tasks | BERT-based molecular representation model using SMILES with self-supervised auxiliary tasks for meaningful molecular embeddings. | arXiv |
| GROVER | Self-Supervised Graph Transformer on Large-Scale Molecular Data | Self-supervised graph Transformer combining GNN message passing with Transformer attention, pretrained on large-scale molecular data. | NeurIPS |
| GEM | Geometry-Enhanced Molecular Representation Learning for Property Prediction | Geometry-enhanced molecular representation learning framework that exploits 3D spatial structure information for improved property prediction. | Nature Machine Intelligence |
| Graphormer | Do Transformers Really Perform Bad for Graph Representation? | Graph Transformer framework from Microsoft that won 1st place on OGB-LSC molecular tasks, introducing spatial and edge encodings for graph structure modeling. | NeurIPS |
| MoleculeSTM | Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing | Multi-modal model jointly learning molecular structures and text descriptions for text-driven molecular retrieval and editing. | Nature Machine Intelligence |
| 3D-MoLM | Towards 3D Molecule-Text Interpretation in Language Models | Pioneering framework integrating a 3D molecular encoder with language models via a 3D molecule-text projector for LLM-based molecular understanding. | ICLR |
| MolE | MolE: A Foundation Model for Molecular Graphs Using Disentangled Attention | Molecular graph foundation model from Recursion using disentangled-attention Transformers, pretrained in two self-supervised stages on ~842M molecules. | Nature Communications |
| MiniMol | MiniMol: A Parameter-Efficient Foundation Model for Molecular Learning | Parameter-efficient molecular foundation model (only 10M parameters) pretrained on 3,300+ diverse bioactivity datasets from 6M molecules. | ICML |
| SMILES-Mamba | SMILES-Mamba: Chemical Mamba Foundation Models for Drug ADMET Prediction | Mamba-architecture chemical foundation model with two-stage training (self-supervised pretraining + supervised fine-tuning) for drug ADMET prediction. | NeurIPS 2024 Workshop |
| SMI-TED | SMI-TED: Large-Scale Foundation Model for Materials and Chemistry | Large-scale SMILES encoder-decoder foundation model from IBM, self-supervised on 91M PubChem SMILES for chemistry and materials science. | ICLR 2024 Workshop |
| KPGT | A Knowledge-Guided Pre-training Framework for Improving Molecular Representation | Knowledge-guided graph Transformer pretraining framework integrating chemical knowledge to enhance molecular representation learning. | Nature Communications |
| GIN (Pretrained) | Strategies for Pre-Training Graph Neural Networks | Pioneering work proposing GNN pretraining strategies (node-level + graph-level) with GIN pretrained on 2M molecules for property prediction. | ICLR |
| MIST | Foundation Models for Discovery and Exploration in Chemical Space | Family of large-scale molecular foundation models (Molecular Insight SMILES Transformers) trained on vast unlabeled molecules, predicting 400+ structure-property relationships. | arXiv |
| M2UMol | Multi-to-Uni Modal Knowledge Transfer Pre-training for Molecular Representation Learning | Multi-modal to uni-modal knowledge transfer pretraining framework that distills diverse molecular modality knowledge into a 2D encoder. | Nature Communications |
| Omni-Mol | Exploring Universal Convergent Space for Omni-Molecular Tasks | Unified language model enabling any-to-any modality molecular tasks in a universal convergent space. | NeurIPS |
| TamGen | TamGen: drug design with target-aware molecule generation through a chemical language model | GPT-style chemical language model for target-aware molecule generation and drug design. | Nature Communications |
| MoleculeGPT | MoleculeGPT: Instruction Following LLMs for Molecular Property Prediction | LLM fine-tuned with molecular instruction data for natural-language-driven molecular property prediction. | NeurIPS 2024 Workshop |
| SAFE-GPT | SAFE: A Molecular-Centric Foundation Model with SAFE Representation | Foundation model using Sequential Attachment-based Fragment Embedding (SAFE) molecular representation for generative chemistry. | Digital Discovery |
| DrugGPT | DrugGPT: A GPT-based Strategy for Designing Potential Ligands Targeting Specific Proteins | GPT-based drug design model that generates drug-like molecules targeting specific protein binding pockets. | bioRxiv |
| MultiPUFFIN | Multimodal domain-constrained foundation model | Multi-modal domain-constrained foundation model integrating SMILES, molecular graphs, and 3D geometry for molecular understanding. | arXiv |
| FragCLM | Foundation chemical language model for fragment-based drug discovery | Foundation chemical language model trained on the ZINC-22 fragment dataset for comprehensive fragment-based drug discovery. | arXiv |
Foundation models for chemical reaction prediction, retrosynthetic planning, and synthesis route design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Molecular Transformer | Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction | Pioneering seq2seq Transformer that frames chemical reaction prediction as SMILES translation with uncertainty calibration. | ACS Central Science |
| Chemformer | Chemformer: a pre-trained transformer for computational chemistry | BART-based pretrained Transformer for reaction prediction, retrosynthesis, and other computational chemistry tasks over molecular SMILES. | Machine Learning: Science and Technology |
| RXNFP | Mapping the space of chemical reactions using attention-based neural networks | Transformer model from IBM that learns chemical reaction fingerprints for reaction classification and reaction space mapping. | Nature Machine Intelligence |
| T5Chem | Unified Deep Learning Model for Multitask Reaction Predictions with Explanation | T5-based unified Transformer supporting multi-task chemical reaction predictions including forward synthesis, retrosynthesis, and yield prediction. | Journal of Chemical Information and Modeling |
| Llamole | Multimodal Large Language Models for Inverse Molecular Design with Retrosynthetic Planning | Multi-modal LLM integrating graph diffusion Transformer and GNN for inverse molecular design with retrosynthetic route planning. | NeurIPS |
| RetroSynFormer | Retrosynformer: Planning Multi-step Chemical Synthesis Routes via a Decision Transformer | Decision Transformer for multi-step retrosynthetic planning that models retrosynthesis as a sequence prediction problem. | Digital Discovery |
| SynLlama | SynLlama: Generating Synthesizable Molecules and Their Analogs with Large Language Models | LLM fine-tuned from Meta Llama 3 for generating synthesizable molecules along with complete synthesis routes. | ACS Central Science |
| SynFormer | Generative Artificial Intelligence for Navigating Synthesizable Chemical Space | Generative model framework for exploring synthesizable chemical space, ensuring generated molecules are synthetically accessible. | PNAS |
| ReactionT5 | ReactionT5: A Pre-trained Transformer Model for Accurate Chemical Reaction Prediction with Limited Data | T5-based pretrained Transformer for chemical reaction prediction, pretrained on the Open Reaction Database and excelling with limited data. | Journal of Cheminformatics |
| DeepRetro | DeepRetro Discovers Retrosynthetic Pathways Through Iterative Large Language Model Reasoning | Advanced retrosynthesis framework combining LLM reasoning, reaction templates, and expert feedback for iterative pathway discovery. | Scientific Reports |
| RXNGraphormer | A unified pre-trained deep learning framework for cross-task reaction performance prediction | Unified pretrained reaction graph Transformer integrating GNN and Transformer to learn bond formation/breaking mechanisms across tasks. | Nature Machine Intelligence |
| RSGPT | RSGPT: a generative transformer for retrosynthesis planning pre-trained on ten billion datapoints | Generative Transformer for retrosynthesis planning pretrained on 10 billion datapoints for large-scale synthetic route prediction. | Nature Communications |
| Chem-R | Chem-R: Learning to Reason as a Chemist | Chemical reasoning model that emulates chemists' deep thinking processes through a three-phase training framework. | NeurIPS |
Foundation models for molecular docking, binding affinity prediction, and protein-ligand complex structure prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Pearl | Pearl: A Foundation Model for Placing Every Atom in the Right Location | Protein-ligand structure prediction foundation model from Genesis Molecular AI using large-scale synthetic data and SO(3)-equivariant architecture, surpassing AlphaFold 3. | NeurIPS |
| DiffDock | DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking | Diffusion-based molecular docking method that models protein-ligand docking as a generative problem on SE(3) without requiring prior binding site knowledge. | ICLR |
| DiffDock-L | Fine-Tuning DiffDock-L for Allosteric Kinase Docking | Large-scale DiffDock variant fine-tuned for allosteric kinase binding site docking. | J. Chem. Inf. Model. |
| NeuralPLexer | State-specific Protein-Ligand Complex Structure Prediction with a Multiscale Deep Generative Model | Multi-scale deep generative model that predicts protein-ligand complex 3D structures directly from protein sequence and ligand graph, including conformational changes. | Nature Machine Intelligence |
| Uni-Mol Docking V2 | Uni-Mol Docking V2: Towards Realistic and Accurate Binding Pose Prediction | Molecular docking method in the Uni-Mol series using pretrained molecular and pocket encoders to predict protein-ligand binding poses with >77% success rate. | Lecture Notes in Computer Science |
| Umol | Structure Prediction of Protein-Ligand Complexes from Sequence Information with Umol | AI system predicting full-flexibility, all-atom protein-ligand complex structures solely from amino acid sequence and SMILES. | Nature Communications |
| LigUnity | A Foundation Model for Protein-Ligand Affinity Prediction Through Unified Representation | Unified representation learning foundation model for protein-ligand affinity prediction supporting both virtual screening and lead optimization. | bioRxiv preprint |
| PhysDock | PhysDock: A Physics-Guided All-Atom Diffusion Model for Protein-Ligand Complex Prediction | Physics-guided all-atom diffusion model for protein-ligand complex prediction integrating detailed atomic-level flexibility modeling. | bioRxiv |
| Boltz-2 | Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction | Advanced model for accurate and efficient protein-ligand binding affinity prediction building on AlphaFold3 and Boltz-1 architectures. | bioRxiv |
Equivariant and invariant neural network architectures for learning 3D molecular representations, energy prediction, and force fields.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| SchNet | SchNet: A Continuous-filter Convolutional Neural Network for Modeling Quantum Interactions | Continuous-filter convolutional neural network learning rotationally invariant representations of quantum interactions for molecular energy and force prediction. | NeurIPS 2017 |
| ViSNet | ViSNet: An Equivariant Geometry-Enhanced Graph Neural Network with Vector-Scalar Interactive Message Passing | Equivariant geometry-enhanced GNN with vector-scalar interactive message passing that avoids expensive higher-order tensor operations via runtime geometric computation. | Nature Communications |
| EPT | An equivariant pretrained transformer for unified 3D molecular representation learning | E(3)-equivariant all-atom pretrained Transformer for unified 3D molecular representation learning across diverse scientific domains. | Nature Communications |
Generative models for de novo molecular design, 3D conformation generation, and structure-based molecule generation using diffusion, VAEs, autoregressive, and flow-based approaches.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MolGPT | MolGPT: Molecular Generation Using a Transformer-Decoder Model | GPT-based molecular generation model using a Transformer decoder to autoregressively generate SMILES satisfying specific property constraints. | J. Chem. Inf. Model. |
| cMolGPT | cMolGPT: A Conditional Generative Pre-Trained Transformer for Target-Specific de novo Molecular Generation | Conditional molecular GPT extending MolGPT with target-specific controls for de novo molecular generation. | Molecules |
| GenMol | GenMol: A Drug Discovery Generalist with Discrete Diffusion | General-purpose molecular generation model from NVIDIA using masked discrete diffusion over SAFE representations for multi-stage drug discovery. | ICLR |
| NExT-Mol | NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generation | Foundation model integrating 1D SELFIES language modeling with 3D diffusion for 3D molecule generation. | ICLR 2025 |
| DiTMC | Sampling 3D Molecular Conformers with Diffusion Transformers | Diffusion Transformer framework for sampling accurate 3D molecular conformers integrating discrete molecular graphs with continuous coordinates. | NeurIPS |
| SynCoGen | Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling | Framework for synthesizable 3D molecule generation that jointly models molecular building blocks, chemical reactions, and atomic coordinates. | ICLR |
Foundation models for interpreting and predicting molecular spectra including NMR, IR, Raman, and mass spectrometry.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MolSpectLLM | MolSpectLLM: A Large Language Model for Molecular Spectroscopy Interpretation | Large language model for molecular spectroscopy interpretation, linking NMR, IR, and MS spectral data to molecular structures for spectrum-to-structure reasoning. | arXiv |
Chemical language models applied to food-related molecular property prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FART | A chemical language model for molecular taste prediction | Chemical language model predicting molecular taste properties from SMILES representations. | npj Science of Food |
Foundation models for electrochemical applications including battery electrolyte design.
| Model | Paper Title | Description | Link |
|---|
Machine-learned interatomic potentials for molecular dynamics simulations.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MACE | MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields | Higher-order equivariant message passing framework using atomic cluster expansion for accurate and efficient force field computation. | NeurIPS 2022 |
| MACE-MP-0 | A Foundation Model for Atomistic Materials Chemistry | Pre-trained universal force field covering 89 elements, trained on Materials Project data for general materials chemistry simulation. | J. Chem. Phys. |
| CHGNet | CHGNet as a Pretrained Universal Neural Network Potential for Charge-Informed Atomistic Modelling | Pre-trained universal graph neural network potential with charge information, trained on Materials Project DFT data. | Nature Machine Intelligence |
| M3GNet | A Universal Graph Deep Learning Interatomic Potential for the Periodic Table | Universal graph deep learning interatomic potential trained on Materials Project relaxation data, covering all periodic table elements. | Nature Computational Science |
| SevenNet | SevenNet: Scalable Graph Neural Network Interatomic Potential | Scalable GNN interatomic potential based on NequIP architecture with LAMMPS parallel MD support. | J. Chem. Theory Comput. |
| Orb | Orb: A Fast, Scalable Neural Network Potential | Fast and scalable neural network potential by Orbital Materials, 3–6× faster than existing universal potentials while maintaining SOTA accuracy. | arXiv |
| NequIP | E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials | E(3)-equivariant GNN using equivariant convolutions instead of invariant descriptors, achieving high accuracy with minimal training data. | Nature Communications |
| Allegro | Learning local equivariant representations for large-scale atomistic dynamics | Highly scalable E(3)-equivariant architecture using local equivariant representations to support large-scale molecular dynamics. | Nature Communications |
| Allegro-FM | Allegro-FM: Toward an Equivariant Foundation Model for Exascale Molecular Dynamics Simulations | Equivariant foundation model targeting exascale molecular dynamics simulations based on the Allegro architecture. | J. Phys. Chem. Lett. |
| DPA-2 | DPA-2: a large atomic model as a multi-task learner | Large-scale Deep Potential multi-task atomic model pre-trained across diverse chemical and materials systems with fine-tuning support. | npj Computational Materials |
| ANI-1 | ANI-1: an extensible neural network potential with DFT accuracy at force field computational cost | Pioneering extensible neural network potential achieving DFT accuracy at force-field computational cost for H/C/N/O organic molecules. | Chemical Science |
| ANI-2x | Extending the Applicability of the ANI Deep Learning Molecular Potential to Sulfur and Halogens | Extension of ANI to sulfur and halogens (F/Cl), broadening coverage to a wider organic molecular space. | J. Chem. Theory Comput. |
| AIMNet2 | AIMNet2: A Neural Network Potential to Meet Your Neutral, Charged, Organic, and Elemental-Organic Needs | Highly transferable neural network potential supporting neutral and charged organic molecules across 14 elements. | Chemical Science |
| GRACE | Graph Atomic Cluster Expansion | Universal MLIP framework based on graph atomic cluster expansion, covering 97 elements. | npj Comp. Mater. |
| Orb-v3 | Orb-v3: Atomistic Simulation at Scale | Major upgrade of Orb with improved accuracy and efficiency for large-scale atomistic simulation. | arXiv |
| PET-MAD | Lightweight universal interatomic potential for advanced materials | Lightweight universal interatomic potential covering the full periodic table for advanced materials simulations. | Nature Communications |
| Grappa | Machine-learned molecular mechanics via E(3)-equivariant neural networks | E(3)-equivariant neural network approach to machine-learned molecular mechanics force fields. | Chemical Science |
Predicting physical, electronic, and structural properties of crystalline and molecular materials.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CGCNN | Crystal Graph Convolutional Neural Networks for an Accurate and Interpretable Prediction of Material Properties | Pioneering crystal graph convolutional neural network for direct, interpretable prediction of material properties from crystal structures. | Physical Review Letters |
| MEGNet | Graph Networks as a Universal Machine Learning Framework for Molecules and Crystals | Universal materials graph network supporting property prediction for both molecules and crystals with global state features. | Chemistry of Materials |
| ALIGNN | Atomistic Line Graph Neural Network for Improved Materials Property Predictions | NIST atomistic line graph neural network that explicitly models bond angles, outperforming CGCNN and MEGNet. | npj Computational Materials |
| MultiMat | Multimodal Foundation Models for Material Property Prediction and Discovery | Multimodal foundation model integrating crystal structure, density of states, charge density, and text for comprehensive materials property prediction. | Newton |
| SMI-TED | SMI-TED: Large-Scale Foundation Model for Materials and Chemistry | IBM large-scale SMILES encoder-decoder model pre-trained on 91M PubChem SMILES for materials and chemistry applications. | ICLR 2024 Workshop |
| DARWIN 1.5 | DARWIN 1.5: Large Language Models as Materials Science Adapted Learners | Open-source materials science LLM that predicts material properties and facilitates discovery from natural language input. | arXiv |
| MatBERT | Quantifying the Advantage of Domain-Specific Pre-training on Named Entity Recognition Tasks in Materials Science | BERT model pre-trained on materials science literature (LBNL), outperforming general models on materials NLP tasks. | Patterns |
| MatSciBERT | MatSciBERT: A Materials Domain Language Model for Text Mining and Information Extraction | Domain-specific BERT trained on materials science literature for enhanced text mining and information extraction. | npj Computational Materials |
| MOFTransformer | A Multi-modal Pre-training Transformer for Universal Transfer Learning in Metal-Organic Frameworks | Multi-modal pre-trained Transformer for MOF property prediction, trained on 1M hypothetical MOFs with atomic graph and energy grid embeddings. | Nature Machine Intelligence |
| CrystalFormer | Space Group Informed Transformer for Crystalline Materials Generation | Autoregressive Transformer guided by space group symmetry and Wyckoff positions for crystalline materials generation. | Science Bulletin |
| MatInFormer | Materials Informatics Transformer: A Language Model for Interpretable Materials Properties Prediction | Materials informatics Transformer leveraging LLM techniques for interpretable materials property prediction. | arXiv |
| KPGT | A Knowledge-Guided Pre-training Framework for Improving Molecular Representation | Knowledge-guided graph Transformer pre-training framework using chemical knowledge to enhance molecular representation learning. | Nature Communications |
| Matformer | Periodic Graph Transformers for Crystal Material Property Prediction | Periodic graph Transformer with periodicity-aware multi-graph attention for crystal material property prediction. | NeurIPS 2022 |
| PotNet | Complete and Efficient Graph Transformers for Crystal Material Property Prediction | Complete and efficient crystal graph Transformer achieving full graph representation via interatomic potential information. | ICLR |
| LLM-Prop | LLM-Prop: Predicting Physical And Electronic Properties of Crystalline Solids From Their Text Descriptions | Uses large language models to predict physical and electronic properties of crystals from text descriptions. | arXiv |
| EScAIP | EScAIP: Efficiently Scaled Attention Interatomic Potential | Efficiently scaled attention-based interatomic potential achieving high accuracy and scalability for materials property prediction. | ICLR |
| AlloyGPT | End-to-end prediction and design of additively manufacturable alloys | Autoregressive language model for end-to-end alloy design and property prediction. | npj Computational Materials |
| aLLoyM | aLLoyM: a large language model for alloy phase diagram prediction | Large language model for predicting alloy phase diagrams. | npj Computational Materials |
| MaskTerial | MaskTerial: a foundation model for automated 2D material flake detection | Foundation model for automated detection of 2D material flakes. | Digital Discovery |
| CLOUD | CLOUD: A Scalable and Physics-Informed Foundation Model for Crystal Representation Learning | Scalable physics-informed crystal representation foundation model trained on 6M+ crystal structures. | Nature Communications |
| LLaMat | A family of large language models for materials research with insights into model adaptability in continued pretraining | Family of large language models adapted for materials science tasks. | Nature Machine Intelligence |
Large-scale pre-trained models spanning molecules, materials, and catalysts across the periodic table.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| UMA | UMA: A Family of Universal Models for Atoms | Meta FAIR universal atomic model family trained on 500M+ 3D atomic structures spanning molecules, materials, and catalysts. | NeurIPS |
| Zatom-1 | Zatom-1: A Multimodal Flow Foundation Model for 3D Molecules and Materials | Open-source multimodal flow foundation model unifying generation and prediction for 3D molecules and materials. | arXiv |
| MIST | Foundation Models for Discovery and Exploration in Chemical Space | Large-scale molecular foundation model family (Molecular Insight SMILES Transformers) predicting 400+ structure-property relationships. | arXiv |
| MatterSim | MatterSim: A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures | Microsoft deep learning atomic model covering all elements at 0–5000 K and 0–1000 GPa. | arXiv |
| GNoME | Scaling Deep Learning for Materials Discovery | Google DeepMind GNN materials explorer discovering 2.2M new stable inorganic crystal structures. | Nature |
| JMP | From Molecules to Materials: Pre-training Large Generalizable Models for Atomic Property Prediction | Meta FAIR joint multi-domain pre-training on ~120M atomic systems spanning molecules and materials. | ICLR |
| ATOMICA | Learning Universal Representations of Intermolecular Interactions with ATOMICA | Geometric deep learning model learning universal atomic-level representations of intermolecular interactions. | bioRxiv |
| eSEN | Efficient Scalable Equivariant Networks | Scalable equivariant architecture forming the backbone of UMA, achieving SOTA on molecular and materials benchmarks. | arXiv |
Generative models for discovering and designing novel crystal structures.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CDVAE | Crystal Diffusion Variational Autoencoder for Periodic Material Generation | Crystal diffusion VAE combining diffusion processes with VAE for end-to-end stable periodic crystal structure generation. | ICLR |
| DiffCSP | Crystal Structure Prediction by Joint Equivariant Diffusion | Joint equivariant diffusion model simultaneously diffusing atom coordinates and lattice parameters for crystal structure prediction. | NeurIPS 2023 |
| SyMat | Towards Symmetry-Aware Generation of Periodic Materials | Symmetry-aware periodic material generation model explicitly leveraging space group symmetry constraints. | NeurIPS 2023 |
| MatterGen | MatterGen: A Generative Model for Inorganic Materials Design | Microsoft diffusion model for inverse design of inorganic crystals conditioned on chemistry, symmetry, and property constraints. | Nature |
| Crystal-GFN | Crystal-GFN: Sampling Crystals with Desirable Properties and Constraints | GFlowNet-based crystal sampling framework that efficiently explores crystal space under property and composition constraints. | arXiv |
| FlowMM | FlowMM: Generating Materials with Riemannian Flow Matching | Riemannian flow matching on crystal manifolds for geometry-aware materials structure generation. | ICML |
| FlowLLM | FlowLLM: Flow Matching for Material Generation with Large Language Models as Base Distributions | Combines LLM base distributions with flow matching, leveraging chemical priors for improved crystal generation. | NeurIPS 2024 |
| CrystalFlow | CrystalFlow: A Flow-Based Generative Model for Crystalline Materials | Flow-based generative model achieving high-fidelity crystal structure generation via normalizing flows. | Nature Communications |
| WyckoffDiff | WyckoffDiff: Diffusion in the Wyckoff Space for Crystal Structure Generation | Diffusion model operating in Wyckoff position space with symmetry-aware representations for improved structural validity. | ICML |
| MatterGPT | MatterGPT: A Generative Transformer for Multi-Property Inverse Design of Solid-State Materials | Autoregressive Transformer supporting multi-property conditioned inverse design of solid-state materials. | arXiv |
| CrystaLLM | CrystaLLM: Large Language Model for Crystallography | LLM that generates crystal structures directly from CIF text without explicit geometric encoding. | Nature Communications |
| UniMat | Scalable Diffusion for Materials Generation | Scalable diffusion model for crystal materials generation with a unified representation across varying crystal sizes. | ICLR |
| DAO-G / DAO-P | Siamese Foundation Models for Crystal Structure Prediction | Siamese pre-training framework: DAO-G for crystal generation and DAO-P for property prediction. | arXiv |
| MOFGPT | Transformer-based generative model for de novo MOF design | Transformer generative model for de novo design of metal-organic frameworks. | arXiv |
| Matra-Genoa | Autoregressive generative material Transformer | Autoregressive Transformer for generative materials design. | npj Computational Materials |
Models for catalytic reaction prediction, adsorption energies, and surface chemistry.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AdsorbML | AdsorbML: A Leap in Efficiency for Adsorption Energy Calculations using Generalizable Machine Learning Potentials | Generalizable ML potentials for efficient adsorption energy calculation, accelerating catalyst screening with OC20 pre-trained models. | npj Computational Materials |
| CatBERTa | CatBERTa: A RoBERTa-based Catalyst Property Prediction Model | RoBERTa-based model predicting catalyst adsorption energies and activities from textual descriptions. | arXiv |
| eSCN | Reducing SO(3) Convolutions to SO(2) for Efficient Equivariant GNNs | Efficient equivariant spherical channel network reducing SO(3) to SO(2) convolutions for major computational speedup. | ICML |
| SCN | Spherical Channels for Modeling Atomic Interactions | Spherical channel network using spherical harmonics for atomic interaction modeling, excelling on OC20 catalyst tasks. | NeurIPS 2022 |
| EquiformerV2 | EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations | Improved equivariant Transformer supporting higher-degree representations, achieving SOTA on OC20/OC22 benchmarks. | ICLR |
| eqV2 | Improved EquiformerV2 for OC20/OC22 | Improved EquiformerV2 variant for general atomic property prediction on Open Catalyst datasets. | arXiv |
| CatDRX | Reaction-conditioned generative model for catalyst design and optimization with CatDRX | Reaction-conditioned generative model for designing catalysts tailored to specific reactions. | Communications Chemistry |
Deep learning models for predicting DFT Hamiltonians and electronic properties.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| DeepH | Deep-learning density functional theory Hamiltonian for efficient ab initio electronic-structure calculation | Deep learning model that directly predicts DFT Hamiltonian matrices to accelerate ab initio electronic structure calculations. | Nature Computational Science |
| DeepH-E3 | DeepH-E3: E(3)-Equivariant Deep Learning for Efficient ab initio Electronic Structure | E(3)-equivariant version of DeepH for more accurate and efficient Hamiltonian matrix element prediction. | Nature Communications |
| HamGNN | HamGNN: Graph Neural Networks for Predicting Hamiltonian Matrix | Graph neural network for Hamiltonian matrix prediction via equivariant message passing at DFT-level accuracy. | arXiv |
| NextHAM | NextHAM: Next-Generation Hamiltonian Prediction with Equivariant Graph Neural Networks | Next-generation equivariant GNN for electronic structure prediction of larger-scale materials systems. | arXiv |
| MACE-H | Equivariant electronic Hamiltonian prediction with many-body message passing | MACE-based equivariant GNN for predicting electronic Hamiltonians. | npj Computational Materials |
Language models and graph networks for polymer informatics and design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| polyBERT | polyBERT: a chemical language model to enable fully machine-driven ultrafast polymer informatics | BERT-based chemical language model trained on polymer SMILES for ultrafast polymer property prediction. | Nature Communications |
| polyGNN | polyGNN: Multitask Graph Neural Networks for Polymer Informatics | Multitask graph neural network for simultaneous polymer property prediction and structure-property learning. | Chemistry of Materials |
| polyBART | polyBART: A Generative Transformer for Polymer Design | BART-based generative Transformer for conditional polymer generation and property-guided inverse design. | arXiv |
| POLYT5 | POLYT5: an encoder-decoder foundation chemical language model for generative polymer design | T5 encoder-decoder foundation chemical language model for polymer design. | npj Artificial Intelligence |
Pre-trained models for battery research, electrode materials, and energy storage.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| BatteryBERT | BatteryBERT: A Pretrained Language Model for Battery Research | BERT fine-tuned on battery literature for text mining and information extraction in battery research. | J. Chem. Inf. Model. |
| BatteryFormer | BatteryFormer: Graph Transformer for Battery Material Property Prediction | Graph Transformer predicting battery electrode capacity, voltage, and cycle life properties. | arXiv |
Foundational graph neural network architectures underlying many materials science models.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| SchNet | SchNet: A Continuous-Filter Convolutional Neural Network for Modeling Quantum Interactions | Pioneering continuous-filter convolutional network encoding interatomic distances as continuous representations for quantum interactions. | NeurIPS |
Foundation models for metamaterial structure-property relationships.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MetaFO | Toward a robust and generalizable metamaterial foundation model | Bayesian Transformer metamaterial foundation model for zero-shot structure-property prediction. | npj Computational Materials |
Models for predicting superconducting properties and critical temperatures.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| BEE-NET | Developing a complete AI-accelerated workflow for superconductor discovery | Equivariant GNN predicting Eliashberg spectral functions and critical temperatures for superconductor discovery. | npj Computational Materials |
| DeeperBand | A deep learning approach to search for superconductors from electronic bands | Symmetry-aware 3D Vision Transformer predicting superconductivity from electronic band structures. | IOPscience |
Foundation models and deep learning architectures for jet tagging, particle tracking, and collider event analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FM4NPP | A Scaling Foundation Model for Nuclear and Particle Physics | Large-scale self-supervised foundation model for sparse detector data, trained on 11M+ collision events achieving SOTA on sPHENIX experiments. | ICLR |
| OmniLearn | OmniLearn: A Method to Simultaneously Facilitate All Jet Physics Tasks | Multi-task jet physics foundation model learning universal representations via multi-class classification pre-training. | arXiv |
| OmniLearned | Foundation Model Framework for All Tasks Involving Jet Physics | Upgraded OmniLearn framework trained on 1B+ jet events with Transformer architecture for jet classification, regression, and generation. | Physical Review D |
| OmniJet-α | OmniJet-α: The first cross-task foundation model for particle physics | First cross-task particle physics foundation model supporting both jet generation and jet tagging. | Machine Learning: Science and Technology |
| Bumblebee | Bumblebee: Foundation Model for Particle Physics Discovery | BERT-inspired particle physics foundation model embedding four-momentum vectors without positional encoding to capture generative and reconstruction-level information. | NeurIPS 2024 Workshop |
| EveNet | EveNet: A Foundation Model for Particle Collision Data Analysis | Event-level collision data foundation model pre-trained on 500M simulated events with hybrid self-supervised learning for multi-task analysis. | arXiv |
| HEP-JEPA | HEP-JEPA: A foundation model for collider physics using joint embedding predictive architecture | Collider physics foundation model using joint embedding predictive architecture (JEPA) for self-supervised jet tagging. | arXiv |
| JetCLR | Symmetries, Safety, and Self-Supervision | Contrastive self-supervised jet representation learning framework using permutation-invariant Transformer encoder with symmetry augmentation. | SciPost Phys. |
| CaloFM | Foundation Model for Calorimetry via MoE | Mixture-of-experts foundation model for calorimeter simulation. | arXiv |
| JetFormer | Scalable Transformer for Jet Tagging | Scalable Transformer architecture for jet tagging. | arXiv |
| PanopTag | PanopTag: Simultaneously Tagging All Jets in a Particle Collision Event | First method to simultaneously tag all jets in a collision event using encoder-decoder Transformer with event-level context. | arXiv |
| TrackingBERT | A Language Model for Particle Tracking | BERT-based foundation model for particle track reconstruction by tokenizing detector data for LHC tracking. | arXiv |
Neural operators and foundation models for solving partial differential equations and fluid simulations.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FNO | Fourier Neural Operator for Parametric Partial Differential Equations | Pioneering neural operator learning function mappings in Fourier space, resolution-independent and efficient for parametric PDEs. | arXiv |
| DeepONet | Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators | Deep operator network with branch/trunk architecture learning continuous nonlinear operators for PDE solving. | Nature Machine Intelligence |
| FourCastNet | FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators | NVIDIA high-resolution global weather model based on adaptive Fourier neural operators at 0.25° resolution. | arXiv |
| Poseidon | Poseidon: Efficient Foundation Models for PDEs | Efficient PDE foundation model using multi-scale operator Transformer with temporal conditional layer normalization. | NeurIPS 2024 |
| MPP | Multiple Physics Pretraining for Physical Surrogate Models | Task-agnostic Transformer pre-trained autoregressively on multiple spatiotemporal physical systems for enhanced generalization. | NeurIPS |
| ICON / ICON-LM | In-context operator learning with data prompts for differential equation problems | In-context operator learning network solving multiple PDE families via data prompts without retraining. | PNAS 2023 |
| VICON | VICON: Vision In-Context Operator Networks for Multi-Physics Fluid Dynamics | Vision in-context operator network applying Vision Transformer to multi-physics fluid dynamics prediction. | arXiv |
| DPOT | DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training | Autoregressive denoising operator Transformer with Fourier attention for large-scale PDE pre-training. | ICML |
| PROSE-PDE | Towards a Foundation Model for Partial Differential Equations: Multi-Operator Learning and Extrapolation | Multimodal PDE foundation model supporting multi-operator learning, extrapolation, and equation identification. | Phys. Rev. E |
| PROSE-FD | PROSE-FD: A Multimodal PDE Foundation Model for Learning Fluid Dynamics | Multimodal zero-shot PDE foundation model for shallow water and Navier-Stokes equations across varied geometries. | arXiv |
| OmniArch | OmniArch: Building Foundation Model for Scientific Computing | Multi-scale multi-physics scientific computing foundation model with Fourier encoder-decoder supporting 1D/2D/3D PDE simulation. | ICML |
| Unisolver | Unisolver: PDE-Conditional Transformers Are Universal PDE Solvers | Universal PDE solver using PDE-conditioned Transformer pre-trained with equation, coefficient, and boundary condition information. | NeurIPS |
| PINO | Physics-Informed Neural Operator for Learning Partial Differential Equations | Physics-informed neural operator combining data-driven and physical constraint losses to learn PDE solution operators with minimal labeled data. | ACM / IMS Journal of Data Science |
| PI-MFM | PI-MFM: Physics-informed multimodal foundation model for solving partial differential equations | Physics-informed multimodal foundation model embedding physical priors to reduce data dependency for PDE solving. | arXiv |
| Walrus | Walrus: A Cross-Domain Foundation Model for Continuum Dynamics | 1.3B-parameter cross-domain continuum dynamics foundation model pre-trained on 19 physical systems covering fluids and solids. | arXiv |
| DISCO | DISCO: Learning to DISCover an Evolution Operator for Multi-Physics-Agnostic Prediction | Multi-physics-agnostic evolution operator discovery method for efficient PDE solving and generalization. | arXiv |
| LFNO | Latent Fourier Neural Operator | Latent-space Fourier neural operator performing transforms in low-dimensional space for improved efficiency. | arXiv |
| RNO | Recurrent Neural Operator | Recurrent neural operator combining recurrent structure with operator learning for temporal PDE dynamics. | arXiv |
| MINO | Masked Implicit Neural Operator | Masked implicit neural operator using masking strategies to enhance generalization in operator learning. | arXiv |
| TNO | Transolver / Transformer Neural Operator | Transformer-based neural operator for PDEs on complex geometries and irregular grids. | ICML 2024 |
| HyPINO | Hybrid Physics-Informed Neural Operator | Hybrid physics-informed neural operator combining physical constraints with data-driven learning for enhanced PDE accuracy. | arXiv |
| PI-Latent-NO | Physics-Informed Latent Neural Operator | Physics-informed latent-space neural operator integrating equation constraints in latent space for PDE solving. | arXiv |
| Transolver | Transolver: A Fast Transformer Solver for PDEs on General Geometries | Physics-Attention Transformer PDE solver supporting arbitrary geometries, achieving multi-benchmark SOTA. | ICML |
| SFNO | Spherical Fourier Neural Operators: Learning Stable Dynamics on the Sphere | Spherical Fourier neural operator serving as the backbone of FourCastNet V2 for global weather prediction. | ICML |
| WIND | WIND: Weather Inverse Diffusion for Zero-Shot Atmospheric Modeling | Zero-shot atmospheric modeling foundation model based on inverse diffusion. | arXiv |
| STAR-MD | Scalable Spatio-Temporal SE(3) Diffusion for Long-Horizon Protein Dynamics | Scalable SE(3)-equivariant diffusion model for simulating long-horizon protein dynamics. | arXiv |
| MORPH | Shape-agnostic PDE Foundation Models | Shape-agnostic autoregressive PDE foundation model handling arbitrary domain geometries. | ICLR |
| PDEformer-2 | Versatile Foundation Model for 2D PDEs | Versatile 2D PDE foundation model encoding equation structure as computational graphs. | arXiv |
| NESTOR | Nested MOE Neural Operator for Large-Scale PDE Pre-Training | Nested mixture-of-experts neural operator for large-scale PDE pre-training. | arXiv |
Large-scale models for multi-physics simulation, mesh-based dynamics, and surrogate modeling.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| GPhyT | Towards a Physics Foundation Model | General physics Transformer trained on 1.8 TB of diverse simulation data with zero-shot generalization to unseen physics scenarios. | arXiv |
| PhysiX | PhysiX: A Foundation Model for Physics Simulations | 4.5B-parameter physics simulation foundation model using discrete tokenizer for multi-scale physical processes with autoregressive generation. | NeurIPS |
| PDE-Transformer | PDE-Transformer: Efficient and Versatile Transformers for Physics Simulations | Scalable Transformer architecture for efficient surrogate modeling across multiple PDE types on regular grids. | arXiv |
| M2PDE | M2PDE: Compositional Generative Multiphysics and Multi-component PDE Simulation | Compositional diffusion-based framework for generative multi-physics and multi-component PDE simulation. | arXiv |
| UPS | Unified PDE Solvers | Unified PDE solver foundation model handling cross-domain, cross-dimension, and cross-resolution spatiotemporal PDEs. | arXiv |
| CompNO | CompNO: A Novel Foundation Model approach for solving Partial Differential Equations | Compositional neural operator splitting monolithic models into composable modules for efficient parametric PDE solving. | Applied Sciences |
| GNS | Learning to Simulate Complex Physics with Graph Networks | DeepMind graph network simulator learning particle interaction rules to generalize across fluids, rigid bodies, and deformable objects. | ICML |
| MeshGraphNets | Learning Mesh-Based Simulation with Graph Networks | Graph network simulation on unstructured meshes for aerodynamics and structural mechanics. | ICLR |
| GeoPT | Scaling Physics Simulation via Lifted Geometric Pre-Training | Geometric pre-training foundation model for scaling physics simulation. | arXiv |
Generative models that learn physical dynamics for interactive environment simulation and embodied AI.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Cosmos | Cosmos World Foundation Model Platform for Physical AI | NVIDIA open-source world foundation model platform generating physics-aware video and world states for robotics and autonomous driving. | arXiv / NVIDIA |
| Genie | Genie: Generative Interactive Environments | DeepMind 11B-parameter unsupervised world model generating interactive virtual worlds from a single image. | ICML |
| Genie 2 | Genie 2: A large-scale foundation world model | Upgraded Genie generating diverse, controllable interactive 3D environments from a single image for embodied AI. | DeepMind |
| Genie 3 | Genie 3: A new frontier for world models | General world model generating consistent interactive 3D worlds from text or images in real time with physical consistency. | DeepMind |
| DIAMOND | Diffusion for World Modeling: Visual Details Matter in Atari | Diffusion-based world modeling achieving high visual fidelity for RL agent training in Atari environments. | NeurIPS 2024 |
| WorldDreamer | WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens | General world model capturing physical dynamics across multiple environments via masked token prediction for video generation. | arXiv |
| GameNGen | Diffusion Models Are Real-Time Game Engines | Google Research neural game engine using diffusion models to simulate complex game environments (DOOM) at 20+ fps. | arXiv preprint (Google) |
| OASIS | Oasis: A Universe in a Transformer | Real-time open-world AI model generating interactive Minecraft-like gameplay at 20 fps via Transformer and diffusion. | Project Page |
| Pandora | Pandora: Towards General World Model with Natural Language Actions and Video States | Hybrid autoregressive-diffusion world model controlling video state generation through natural language actions. | arXiv |
| UniSim | Learning Interactive Real-World Simulators | Universal world simulator learning to simulate diverse human-world interactions from text, actions, and image inputs. | ICLR |
| PAN | PAN: A World Model for General, Interactable, and Long-Horizon World Simulation | Action-conditioned world model for general, interactable, long-horizon simulation with environment dynamics consistency. | arXiv |
| PhysDreamer | PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation | Physics-based 3D object interaction generation via video generation for physically consistent object manipulation. | Lecture Notes in Computer Science |
| Astra | General Interactive World Model with Autoregressive Denoising | General interactive world model combining autoregressive and denoising generation. | ICLR |
Foundation models for quantum state representation, many-body simulation, and quantum dynamics.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FNQS | Foundation Model for Quantum Many-Body States via Transformers | Transformer-based foundation model pre-trained on quantum state data, generalizing across different Hamiltonians and lattice structures. | Nature Communications |
| Attention-Based FM for Quantum States | Attention-Based Foundation Model for Quantum States | Attention-based foundation model using self-attention to capture quantum correlations for cross-system quantum state representation. | arXiv |
| NOQS | Neural Operator for Quantum States | Neural operator learning continuous mappings over quantum state space for efficient quantum simulation and prediction. | arXiv |
| Large Electron Model | Large Electron Model: A Foundation Model for Electron Systems | Foundation model for electron systems pre-trained on large-scale electronic structure data, supporting quantum chemistry and materials physics tasks. | arXiv |
| DysonNet | Constant-Time Local Updates for Neural Quantum States | Neural quantum state architecture achieving O(1) local update efficiency for scalable quantum simulation. | arXiv |
Multi-modal models for tokamak plasma behavior prediction and fusion control.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| TokaMind | TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma | Multi-modal Transformer foundation model fusing multiple diagnostic modalities for tokamak plasma behavior prediction and fusion control. | arXiv |
Foundation models for optical design, thin-film structures, and photonic inverse design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| OptoGPT | OptoGPT: A Foundation Model for Inverse Design in Optical Multilayer Thin Film Structures | GPT-based foundation model for automatic inverse design of optical multilayer thin-film structures to meet target spectral properties. | Opto-Electronic Advances |
| MOCLIP | MOCLIP: Multi-modal Optical Contrastive Learning for Inverse Photonic Design | Multi-modal optical contrastive learning model using a CLIP framework for cross-modal photonic structure inverse design. | arXiv |
Neural architectures that preserve physical symmetries, conservation laws, and geometric structure.
| Model | Paper Title | Description | Link |
|---|
Models for electronic structure calculation, wavefunction prediction, and molecular quantum properties.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Skala | Accurate and scalable exchange-correlation with deep learning | Microsoft quantum chemistry foundation model for electronic structure computation and cross-system molecular property prediction. | arXiv |
| Orbformer | Orbformer: Orbital Transformer for Electronic Structure Prediction | Orbital Transformer predicting electronic structure from molecular orbital representations with physical symmetry priors. | arXiv |
| OrbEvo | Orbital Transformers for Predicting Wavefunctions in TD-DFT | Equivariant graph Transformer predicting real-time TD-DFT wavefunction evolution. | arXiv |
Deep learning platforms for reactive flow and combustion CFD.
| Model | Paper Title | Description | Link |
|---|
Domain-specific foundation models for nuclear reactor control and simulation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| NucReactor-FM | Agentic Physical AI toward a Domain-Specific FM for Nuclear Reactor Control | Domain-specific foundation model for nuclear reactor control combining physics simulation with reinforcement learning. | arXiv |
Foundation models for global weather forecasting and climate prediction at various spatial and temporal scales.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Aurora | Aurora: A Foundation Model of the Earth System | A large-scale earth system foundation model by Microsoft trained on over one million hours of multi-source geophysical data for atmosphere, ocean, and air quality prediction. | Nature |
| Pangu-Weather | Accurate Medium-range Global Weather Forecasting with 3D Neural Networks | A 3D high-resolution AI weather forecasting model by Huawei trained on 43 years of ERA5 reanalysis data, generating 10-day global forecasts in seconds. | Nature |
| GraphCast | Learning Skillful Medium-range Global Weather Forecasting | A graph neural network-based global medium-range weather forecasting model by Google DeepMind at 0.25° resolution, outperforming ECMWF HRES for 10-day forecasts. | Science |
| GenCast | GenCast: Diffusion-based Ensemble Forecasting for Medium-range Weather | A diffusion-based ensemble weather forecasting system by Google DeepMind generating probabilistic 15-day forecasts that surpass ECMWF ENS. | Nature |
| ClimaX | ClimaX: A Foundation Model for Weather and Climate | The first weather and climate foundation model based on Transformer architecture supporting flexible fine-tuning for multiple downstream meteorological tasks. | ICML |
| FengWu | FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead | A multi-modal multi-task weather forecasting system by Shanghai AI Lab that extends deterministic forecast skill to 10.75 days. | arXiv |
| FuXi | FuXi: A Cascade Machine Learning Forecasting System for 15-day Global Weather Forecast | A cascade machine learning weather forecasting system trained on 39 years of ERA5 data, achieving 15-day forecast performance comparable to ECMWF ensemble mean. | npj Climate and Atmospheric Science |
| NeuralGCM | Neural General Circulation Models for Weather and Climate | A neural GCM by Google Research combining differentiable atmospheric dynamics solvers with machine learning for weather-to-climate timescale prediction. | Nature |
| FourCastNet | FourCastNet: A Global Data-driven High-resolution Weather Forecasting System | A high-resolution global weather forecasting model by NVIDIA using adaptive Fourier neural operators (AFNO) at 0.25° resolution. | arXiv |
| ECMWF AIFS | AIFS — ECMWF's Data-driven Forecasting System | ECMWF's operational AI forecasting system combining graph neural networks and Transformers for data-driven weather prediction. | arXiv |
| Stormer | Scaling Transformer Neural Networks for Skillful and Reliable Medium-range Weather Forecasting | A streamlined and efficient Transformer weather forecasting model achieving state-of-the-art performance with less training data. | Advances in Neural Information Processing Systems 37 |
| AtmoRep | AtmoRep: A Stochastic Model of Atmosphere Dynamics Using Large Scale Representation Learning | A task-agnostic atmospheric foundation model based on large-scale representation learning for stochastic atmosphere dynamics. | arXiv |
| WeatherGFT | WeatherGFT: Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling | A hybrid physics-AI weather forecasting model extending predictions to finer temporal resolutions at 30-minute intervals. | NeurIPS |
| WeatherGFM | WeatherGFM: Learning A Weather Generalist Foundation Model via In-context Learning | A weather generalist foundation model unifying forecasting, super-resolution, image translation, and post-processing via in-context learning. | ICLR |
| Prithvi WxC | Prithvi WxC: Foundation Model for Weather and Climate | A 2.3-billion-parameter weather and climate foundation model by IBM and NASA trained on 160 MERRA-2 variables. | arXiv |
| FuXi-2.0 | FuXi-2.0: Advancing machine learning weather forecasting model for practical applications | An upgraded version of FuXi providing hourly global forecasts with a more comprehensive set of meteorological variables. | arXiv |
| ArchesWeather | ArchesWeather: An efficient AI weather forecasting model at 1.5° resolution | A lightweight and efficient AI weather forecasting model at 1.5° resolution using a combination of 2D and column-wise attention. | arXiv |
| W-MAE | W-MAE: Pre-trained Weather Model with Masked Autoencoder | A task-agnostic atmospheric foundation model based on masked autoencoder pre-training for weather data. | arXiv |
| Omni-Weather | Omni-Weather: Unified Multimodal Foundation Model for Weather Generation and Understanding | A unified multimodal foundation model integrating radar, satellite, and numerical data for weather generation and understanding. | arXiv |
Foundation models for satellite imagery analysis, multi-spectral and multi-temporal earth observation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Prithvi-EO-2.0 | Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications | A multi-temporal earth observation foundation model by NASA/IBM trained on 4.2 million global time-series samples supporting Landsat and Sentinel-2. | arXiv |
| SpectralGPT | SpectralGPT: Spectral Remote Sensing Foundation Model | The first spectral remote sensing foundation model using a 3D generative pre-trained Transformer designed for multi-spectral and hyperspectral satellite imagery. | IEEE TPAMI (2024) |
| SatMAE | SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery | A masked autoencoder pre-training framework for temporal and multi-spectral satellite imagery. | NeurIPS 2022 |
| SeaMo | SeaMo: A Multi-Seasonal and Multimodal Remote Sensing Foundation Model | A multi-seasonal multimodal remote sensing foundation model fusing optical, SAR, and meteorological data. | arXiv |
| TerraMind | TerraMind: Large-Scale Generative Multimodality for Earth Observation | A large-scale generative multimodal earth observation foundation model by IBM/ESA/DLR trained on 500 billion tokens. | ICCV 2025 |
| RingMo | RingMo: A Remote Sensing Foundation Model with Masked Image Modeling | A remote sensing foundation model by the Chinese Academy of Sciences using masked image modeling for large-scale pre-training. | IEEE TGRS |
| SatCLIP | SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery | A global general-purpose location encoder by Microsoft using Sentinel-2 contrastive learning to generate location embeddings. | AAAI |
| SkySense | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery | A large-scale multimodal remote sensing foundation model pre-trained on 21.5 million temporal optical and SAR data samples. | CVPR |
| Scale-MAE | Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning | A scale-aware masked autoencoder that explicitly models spatial resolution relationships for multiscale geospatial representation learning. | ICCV 2023 |
| DOFA | Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation | A neural plasticity-inspired multimodal EO foundation model using dynamic wavelength-adaptive hypernetworks to handle diverse sensor data. | arXiv |
| GFM (Prithvi-EO-1.0) | Foundation Models for Generalist Geospatial Artificial Intelligence | NASA/IBM's first-generation earth science foundation model based on self-supervised Vision Transformers trained on HLS data. | arXiv |
| S2MAE | S2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing Data | A spatial-spectral masked autoencoder providing joint spatial-spectral pre-training for spectral remote sensing imagery. | CVPR 2024 |
| RingMoE | RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation | A 14.7-billion-parameter mixture-of-modality-experts remote sensing foundation model pre-trained on 400M+ samples. | arXiv |
| WaveMAE | WaveMAE: Wavelet-decomposition Masked Autoencoder for Multispectral Satellite Imagery | A self-supervised foundation model combining wavelet decomposition with geospatial priors for multispectral satellite imagery. | arXiv |
| RoMA | RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing | A scalable Mamba-architecture remote sensing foundation model addressing ViT limitations in large-scale remote sensing pre-training. | NeurIPS |
| CROMA | CROMA: Contrastive Radar-Optical Masked Autoencoders for Remote Sensing | A contrastive radar-optical masked autoencoder for multimodal remote sensing representation learning. | NeurIPS |
| AnySat | AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities | A unified multi-resolution multimodal earth observation model using JEPA architecture for diverse EO tasks. | CVPR |
| TerraFM | TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation | A scalable self-supervised foundation model for unified multisensor earth observation pre-trained on 18.7 million samples. | ICLR 2026 |
Foundation models for ocean forecasting, eddy-resolving prediction, and marine environment monitoring.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| OceanGPT | OceanGPT: A Large Language Model for Ocean Science Tasks | A domain-specific large language model for ocean science by Zhejiang University using the DoInstruct framework for ocean domain instruction data. | ACL 2024 |
| XiHe | XiHe: A Data-Driven Model for Global Ocean Eddy-Resolving Forecasting | A data-driven global ocean eddy-resolving forecast model at 1/12° resolution trained on 25 years of reanalysis data. | arXiv |
| WV-Net | WV-Net: A Foundation Model for SAR WV-mode Satellite Imagery Trained Using Contrastive Self-supervised Learning | The first foundation model for SAR ocean satellite imagery using self-supervised contrastive learning on synthetic aperture radar data. | arXiv |
| GLONET | GLONET: Mercator's End-to-End Neural Global Ocean Forecasting System | An end-to-end neural network global ocean forecasting system by Mercator Ocean trained on GLORYS12 reanalysis data. | Journal of Geophysical Research: Machine Learning and Computation |
| WenHai | Forecasting the Eddying Ocean with a Deep Neural Network | A deep neural network ocean forecasting system excelling at mesoscale eddy dynamics prediction. | Nature Communications |
| FuXi-Ocean | FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution | A data-driven global ocean forecasting system with 6-hour temporal and 1/12° spatial resolution reaching 1500-meter depth. | NeurIPS |
| ORCA-DL | Data-driven Global Ocean Modeling for Seasonal to Decadal Prediction | A data-driven global ocean model supporting 3D ocean predictions from seasonal to decadal timescales. | Science Advances |
| FuXi-ONS | Data-driven Ensemble Prediction of the Global Ocean | A machine-learning ensemble global ocean forecasting system for 5-to-365-day predictions. | arXiv |
Foundation models for earthquake detection, seismic phase picking, and waveform analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| PhaseNet | PhaseNet: A Deep-Neural-Network-Based Seismic Arrival Time Picking Method | One of the most widely used deep learning models for seismic arrival-time picking in seismology. | Geophysical Journal International |
| EQTransformer | EQTransformer: An Attentive Deep-Learning Model for Simultaneous Earthquake Detection and Phase Picking | An attention-based deep learning model for simultaneous earthquake detection and seismic phase picking. | Nature Communications |
| SeisT | SeisT: A Foundational Deep-Learning Model for Earthquake Monitoring Tasks | A Transformer-based seismic monitoring foundation model supporting multiple earthquake tasks including detection, phase picking, and magnitude estimation. | IEEE Transactions on Geoscience and Remote Sensing |
| SeisLM | SeisLM: a Foundation Model for Seismic Waveforms | A large-scale self-supervised seismic waveform foundation model pre-trained via contrastive learning on massive open-source seismic data. | arXiv |
| SeisMoLLM | SeisMoLLM: Advancing Seismic Monitoring via Cross-modal Transfer with Pre-trained Large Language Model | A seismic monitoring foundation model leveraging cross-modal transfer from GPT-2 architecture for seismic analysis. | arXiv |
| SeismicXM | SeismicXM: A Cross-Task Foundation Model for Single-Station Seismic Waveform Processing | A cross-task seismic waveform processing foundation model by China Earthquake Administration supporting multiple single-station tasks. | SRL |
| U-Trans | U-Trans: A Foundation Model for Seismic Waveform Representation | A U-Net encoder-decoder architecture seismic waveform representation foundation model trained on 2M+ three-component waveforms. | Scientific Reports |
| PhaseNet+ | PhaseNet+: Towards End-to-End Earthquake Monitoring Using a Multitask Deep Learning Model | A multi-task extension of PhaseNet enabling end-to-end earthquake monitoring. | arXiv |
| SeisCLIP | SeisCLIP: Contrastive Multimodal Seismology Foundation Model | A contrastive multimodal seismology foundation model learning joint representations from seismic waveforms and metadata. | arXiv |
Foundation models for hydrological prediction, flood modeling, and river forecasting.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| HydroGAT | HydroGAT: Distributed Heterogeneous Graph Attention Transformer for Spatiotemporal Flood Prediction | A graph attention network-based hydrological prediction foundation model capturing spatial dependencies across watersheds. | arXiv |
| ZeroFlood | ZeroFlood: A Geospatial Foundation Model for Data-Efficient Flood Susceptibility Mapping | A geospatial foundation model for data-efficient flood susceptibility mapping. | arXiv |
| GraphRiverCast | Topology-informed AI Foundation Model for Global River Forecasting | A topology-informed AI foundation model for global river hydrodynamic forecasting. | arXiv |
Models for wildfire danger forecasting and fire spread prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FireCastNet | FireCastNet: earth-as-a-graph for seasonal fire prediction | A deep learning global wildfire danger forecasting model fusing meteorological, vegetation, and terrain data for multi-timescale prediction. | Scientific Reports |
| FireScope | FireScope: Wildfire Risk Prediction with a Chain-of-Thought Oracle | A wildfire risk prediction foundation model and benchmark using multi-modal data for fire risk assessment. | arXiv |
Foundation models for atmospheric pollution forecasting.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AirCast | AirCast: Improving Air Pollution Forecasting Through Multi-Variable Data Alignment | A data-driven air quality forecasting foundation model supporting multi-variable atmospheric pollutant concentration prediction. | arXiv |
| FuXi-Air | FuXi-Air: Urban Air Quality Forecasting Based on Emission-Meteorology-Pollutant multimodal Machine Learning | A multimodal machine learning air quality forecasting extension of the FuXi series integrating emission, meteorological, and observational data. | arXiv |
Foundation models for sea ice monitoring and polar region forecasting.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| SIFM | SIFM: A Foundation Model for Multi-granularity Arctic Sea Ice Forecasting | A sea ice foundation model for multi-granularity Arctic sea ice concentration forecasting from satellite observations. | arXiv |
| IceNet | IceNet: Seasonal Arctic Sea Ice Forecasting with Probabilistic Deep Learning | A probabilistic deep learning model for seasonal Arctic sea ice forecasting that significantly outperforms dynamical physics models. | Nature Communications |
| IceMamba | IceMamba: Seasonal Forecasting of Pan-Arctic Sea Ice with State Space Model | A Mamba state space model-based sea ice forecasting foundation model efficiently processing polar spatiotemporal sequence data. | arXiv |
Large language models specialized for earth science knowledge understanding and reasoning.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| K2 | K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization | The first 7-billion-parameter geoscience LLM based on LLaMA, further pre-trained and instruction-tuned on earth science literature. | Proceedings of the 17th ACM International Conference on Web Search and Data Mining |
| JiuZhou | JiuZhou: Open Foundation Language Models for Geoscience | A multilingual geoscience LLM by Tsinghua University supporting Chinese and English earth science knowledge QA and reasoning. | GitHub |
| GeoGPT | GeoGPT: A Large Language Model for Geospatial Artificial Intelligence | A geospatial AI LLM by Zhejiang Lab combining tool-calling capabilities for geospatial analysis and reasoning. | GitHub |
| GeoGalactica | GeoGalactica: A Scientific Large Language Model for Geoscience | A 30-billion-parameter geoscience LLM based on Galactica architecture pre-trained on earth science corpora. | arXiv |
Foundation models for subsurface characterization, seismic exploration, and well log analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Transparent Earth | The Transparent Earth: A Multimodal Foundation Model for the Earth's Subsurface | A multimodal transformer-based foundation model by LANL for subsurface structure imaging and inversion. | arXiv |
| GEM 3D | Geological Everything Model 3D: A Promptable Foundation Model for Subsurface Understanding | A promptable generative 3D earth model unifyin |
Truncated — view the full README on GitHub.
Curated list of 1,000+ scientific foundation models spanning life sciences, chemistry, physics, medicine, and more.
37
26 commits
updated Jun 17, 2026
One map, every frontier — your starting point for 1,000+ AI-for-Science models.
A comprehensive, bilingual (EN/ZH) guide to foundation models driving the next wave of scientific breakthroughs — 1,000+ models across nine domains, from protein design to weather prediction. Cross-disciplinary models are listed wherever they apply.
Language: English | Chinese
Large-scale pretrained models learning protein representations from amino acid sequences.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| ESM-1b | Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences | Large-scale unsupervised protein language model trained on 250 million sequences to capture evolutionary diversity and biological properties. | PNAS |
| ESM-2 | Evolutionary-scale prediction of atomic-level protein structure with a language model | 15-billion-parameter protein language model enabling atomic-level structure prediction from single sequences (ESMFold). | Science |
| ESM3 | Simulating 500 million years of evolution with a language model | Multimodal protein generative model that jointly processes sequence, structure, and function to simulate 500 million years of evolution. | Science |
| ESMFold | Evolutionary-scale prediction of atomic-level protein structure with a language model | Single-sequence protein structure prediction method based on ESM-2, approximately 60× faster than AlphaFold2. | Science |
| ESM Cambrian (ESMC) | ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning | Next-generation ESM protein language model that reveals intrinsic biological principles of proteins through unsupervised learning. | EvolutionaryScale |
| ProtTrans | ProtTrans: Toward understanding the language of life through self-supervised learning | Suite of large-scale protein pretrained models (ProtBERT, ProtXLNet, ProtT5, etc.) trained on 393 billion amino acids. | IEEE TPAMI |
| ProteinBERT | ProteinBERT: a universal deep-learning model of protein sequence and function | Universal protein model jointly pretrained on sequences and GO annotations for diverse protein property prediction. | Bioinformatics |
| ProtGPT2 | ProtGPT2 is a deep unsupervised language model for protein design | GPT-2-based protein sequence generation model that produces novel sequences resembling natural proteins. | Nature Communications |
| ProGen | Large language models generate functional protein sequences across diverse families | Large-scale protein language model from Salesforce trained on 280 million sequences to generate functional artificial proteins. | Nature Biotechnology |
| ProGen2 | ProGen2: Exploring the boundaries of protein language models | 6.4-billion-parameter protein language model trained on genomic, metagenomic, and protein family data. | Cell Systems |
| ProGen3 | Scaling unlocks broader generation and deeper functional understanding of proteins | Sparse mixture-of-experts protein generative model trained on 1.5 trillion tokens for enhanced protein design. | bioRxiv |
| SaProt | SaProt: Protein language modeling with structure-aware vocabulary | Structure-aware protein language model that encodes 3D structures as discrete tokens via Foldseek for joint training with sequences. | ICLR 2024 |
| xTrimoPGLM | xTrimoPGLM: Unified 100B-scale pre-trained transformer for deciphering the language of protein | 100-billion-parameter unified protein language model supporting both protein understanding and generation tasks. | Nature Methods |
| Ankh | Ankh: Optimized protein language model unlocks general-purpose modelling | Optimized protein language model emphasizing training efficiency to achieve competitive performance with fewer resources. | arXiv |
| ProCyon | ProCyon: A multimodal foundation model for protein phenotypes | Multimodal protein foundation model integrating sequence, structure, and natural language data to predict protein phenotypes. | bioRxiv |
| ProSST | ProSST: Protein language modeling with quantized structure and disentangled attention | Protein language model combining amino acid sequences with quantized 3D structural information. | NeurIPS 2024 |
| InstructPLM | InstructPLM: Aligning protein language models to follow protein design instructions | Instruction-tuned ESM2 model that surpasses ESM3 on protein design tasks through alignment training. | arXiv |
| UniRep | Unified rational protein engineering with sequence-based deep representation learning | RNN-based protein language model for unified representation learning and rational protein engineering. | Nature Methods |
| MSA Transformer | MSA Transformer | Transformer model that operates over multiple sequence alignments for improved protein modeling. | ICML |
| EVE | Disease variant prediction with deep generative models of evolutionary data | Evolutionary variational autoencoder for predicting disease-causing genetic variants from protein family data. | Nature |
| Tranception | Protein fitness prediction with autoregressive transformers and inference-time retrieval | Autoregressive language model with retrieval-time alignment context for protein fitness prediction. | ICML |
| RITA | RITA: a Study on Scaling Up Generative Protein Sequence Models | Scaling study of autoregressive protein generative models up to 1.2 billion parameters. | arXiv |
| ProtSSN | Semantical and Geometrical Protein Encoding for Zero-Shot Engineering | Structure-plus-sequence denoising pretraining framework for zero-shot protein engineering. | eLife |
| ProLLaMA | ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing | Multi-task protein LLM based on the LLaMA architecture for diverse protein language processing tasks. | TAI |
| ProTrek | ProTrek: Navigating the Protein Universe through Tri-Modal Contrastive Learning | Tri-modal protein model learning joint representations of sequence, structure, and function via contrastive learning. | Nature Biotechnology |
| ProtWord | ProtWord: A Discrete Protein Language Model for Functional Discovery and De Novo Design | 150M-parameter discrete protein language model that translates sequences into an 8,192-token vocabulary for functional discovery. | bioRxiv |
| ProtLLM | ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training | Interleaved protein-language large language model with a dynamic protein mounting mechanism. | ACL |
| ProtHyena | Hyena architecture enables fast and efficient protein language modeling | Fast Hyena-based protein language model leveraging implicit convolution for efficient sequence processing. | iMetaOmics |
| DPLM | Diffusion Language Models Are Versatile Protein Learners | Diffusion-based language model for versatile protein sequence generation and understanding. | ICML |
| PTM-Mamba | PTM-Mamba: a PTM-aware protein language model with bidirectional gated Mamba blocks | State-space model with bidirectional gated Mamba blocks for post-translational-modification-aware protein representation. | Nature Methods |
| DeepSequence | Deep generative models of genetic variation capture the effects of mutations | Variational autoencoder over aligned protein families for predicting the effects of mutations. | Nature Methods |
| PoET | PoET: A generative model of protein families as sequences-of-sequences | Sequences-of-sequences family modeling framework with retrieval for protein fitness prediction. | Advances in Neural Information Processing Systems 36 |
| Prot2Text | Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers | Multimodal model generating natural language descriptions of protein functions from structure and sequence. | AAAI 2024 |
| PAIR | Boosting the predictive power of protein representations with a corpus of text annotations | Method that boosts protein representations by incorporating text annotation corpora. | Nature Machine Intelligence |
| MULAN | MULAN: Multimodal protein language model for sequence and structure encoding | Multimodal protein encoder that jointly models sequence and structure information. | Bioinformatics Advances |
| ProteinSage | ProteinSage: From implicit learning to explicit structural constraints for efficient protein language modeling | Protein language model combining implicit learning with explicit structural constraints for improved efficiency. | bioRxiv |
| OneProt | OneProt: Towards Multi-Modal Protein Foundation Models | Multi-modal protein foundation model integrating structure, sequence, text, and binding site data. | arXiv |
| ECNet | ECNet is an evolutionary context-integrated deep learning framework for protein engineering | Deep learning framework integrating evolutionary context for protein engineering and fitness prediction. | Nature Communications |
| PoET-2 | Understanding protein function with a multimodal retrieval-augmented foundation model | Next-generation retrieval-augmented multimodal protein family model improving fitness prediction over PoET. | arXiv |
| FlexRibbon | FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling | 3-billion-parameter model jointly pretrained on amino acid sequences and 3D structures capturing flexible conformations. | bioRxiv |
| ProteinAligner | ProteinAligner: A Tri-Modal Contrastive Learning Framework for Protein Representation Learning | Tri-modal contrastive learning framework integrating protein sequences, structures, and scientific literature. | OpenReview |
| ProteinTalks | ProteinTalks: Multi-Modal Protein Language Model with Natural Language Interaction | Multi-modal protein language model supporting natural language interaction for protein knowledge retrieval. | bioRxiv |
| GearNet | Protein Representation Learning by Geometric Structure Pretraining | Relational graph neural network learning protein structure representations via geometric pretraining and multi-view contrastive learning. | ICLR |
| MIF | Masked inverse folding with sequence transfer for protein representation learning | Self-supervised masked inverse folding pretraining for learning protein representations from structures. | Protein Engineering, Design and Selection |
| PPLM | A paired sequence language model for protein-protein interaction | Paired sequence language model predicting protein-protein interactions from paired amino acid sequences. | Nature Communications |
| AIDO.Protein | Mixture of experts enable efficient and effective protein understanding and design | 16B-parameter mixture-of-experts protein module trained on 1.2 trillion amino acids within the AIDO ecosystem. | bioRxiv |
| BioReason-Pro | BioReason-Pro: Advancing Protein Function Prediction with Multimodal Biological Reasoning | First multimodal reasoning LLM for protein function prediction integrating ESM3 embeddings with GO-GPT ontology modeling; achieves 73.6% F_max on GO term prediction, preferred over UniProt annotations by human experts 79% of the time. | bioRxiv |
Methods for predicting 3D protein structures from sequences.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AlphaFold2 | Highly accurate protein structure prediction with AlphaFold | Revolutionary protein structure prediction model achieving atomic-level accuracy at CASP14. | Nature |
| AlphaFold3 | Accurate structure prediction of biomolecular interactions with AlphaFold 3 | Diffusion-based model predicting biomolecular interaction structures for proteins, nucleic acids, small molecules, and ions. | Nature |
| RoseTTAFold | Accurate prediction of protein structures and interactions using a three-track neural network | Three-track neural network for protein structure prediction as an open-source alternative to AlphaFold2. | Science |
| RoseTTAFold2 | Efficient and accurate prediction of protein structure using RoseTTAFold2 | Upgraded RoseTTAFold combining key features from AlphaFold2 and the original RoseTTAFold. | bioRxiv |
| RoseTTAFold All-Atom | Generalized biomolecular modeling and design with RoseTTAFold All-Atom | All-atom biomolecular modeling framework supporting complex prediction of proteins, nucleic acids, small molecules, and metal ions. | Science |
| OmegaFold | High-resolution de novo structure prediction from primary sequence | MSA-free single-sequence protein structure prediction leveraging pretrained protein language models. | bioRxiv |
Generative models for de novo protein backbone and sequence design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| ProteinMPNN | Robust deep learning-based protein sequence design using ProteinMPNN | Message-passing neural network for protein sequence design with experimentally validated high success rates. | Science |
| RFdiffusion | De novo design of protein structure and function with RFdiffusion | Diffusion model built on RoseTTAFold for de novo protein backbone design. | Nature |
| RFdiffusion3 | RFdiffusion3: All-atom biomolecular design | Upgraded RFdiffusion supporting all-atom-level protein and biomolecular design. | bioRxiv |
| Chroma | Illuminating protein space with a programmable generative model | Programmable protein diffusion generative model from Generate:Biomedicines. | Nature |
| FrameDiff | SE(3) diffusion model with application to protein backbone generation | SE(3)-equivariant diffusion model for protein backbone generation without pretrained structure prediction networks. | ICLR |
| FrameFlow | SE(3) stochastic flow matching for protein backbone generation | SE(3) flow matching model for protein backbone generation. | ICLR |
| FoldingDiff | Protein structure generation via folding diffusion | Diffusion model based on protein folding angle representations that simulates natural folding processes. | Nature Communications |
| Genie | Genie: SE(3)-equivariant generative model for protein backbone design | SE(3)-equivariant DDPM model for protein backbone design. | ICML (Workshop) |
| Genie 2 | Out of many, one: Designing and scaffolding proteins at the scale of the structural universe with Genie 2 | Upgraded Genie capturing a broader and more diverse protein structure space. | arXiv |
| Proteus | Proteus: Exploring protein structure generation for enhanced designability and efficiency | Efficient protein backbone generation model without requiring pretrained structure prediction networks. | bioRxiv |
| FoldFlow | FoldFlow: SE(3) stochastic flow matching for protein backbone generation | Family of stochastic flow matching models for protein backbone generation. | ICLR |
| ProteinGenerator (PG) | Multistate and functional protein design using RoseTTAFold | RoseTTAFold-based diffusion model that simultaneously generates protein sequences and structures. | Nature Biotechnology |
| SeedProteo | SeedProteo: All-atom protein design | Diffusion-based all-atom protein design model integrating structure and sequence information for binder design. | arXiv |
| ESM-IF (ESM-IF1) | Language models generalize beyond natural proteins | Inverse folding model conditioned on backbone structures for generating protein sequences. | bioRxiv |
| EvoDiff | Protein generation with evolutionary diffusion | Sequence-based diffusion model for protein generation leveraging evolutionary data. | bioRxiv/Nature Biotechnology |
| ZymCTRL | ZymCTRL: a conditional language model for the generation of artificial enzymes | Conditional enzyme language model trained on 37 million BRENDA enzyme sequences. | bioRxiv |
| PiFold | PiFold: Toward effective and efficient protein inverse folding | Efficient inverse folding model with a novel PiGNN architecture. | ICLR |
| LM-Design | Structure-informed language models are protein designers | Structure-informed language model for protein design combining PLM and structural context. | ICML |
| ProteinDT | A text-guided protein design framework | Text-guided protein generation framework using multimodal learning. | Nature Machine Intelligence |
| Protpardelle | An all-atom protein generative model | All-atom generative model for producing complete protein structures including side chains. | PNAS |
| Multiflow | Generative Flows on Discrete State-Spaces for Protein Co-Design | Discrete flow matching framework for joint sequence-structure protein co-design. | ICML |
| La-Proteina | La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching | Atomistic protein generation method via partially latent flow matching. | arXiv |
| EvoFlows | Evolutionary Edit-Based Flow-Matching for Protein Engineering | Evolutionary edit-based flow matching approach for directed protein engineering. | arXiv |
| Fold2Seq | Fold2Seq: A joint sequence(1D)-fold(3D) embedding-based generative model for protein design | Joint sequence-fold embedding generative model learning from both 3D structures and 1D sequences. | ICML |
| TERMinator | TERMinator: A neural framework for structure-based protein design using tertiary repeating motifs | Neural network framework for protein design based on tertiary repeating motifs. | Nature Communications |
| AlphaDesign | AlphaDesign: A graph protein design method and benchmark on AlphaFoldDB | Graph-based protein design method benchmarked on the AlphaFold Database. | arXiv |
| Latent-X | Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design | Atom-level frontier model generating all-atom structures and sequences for de novo protein binder design. | arXiv |
| Latent-X2 | Drug-like antibodies with low immunogenicity in human panels designed with Latent-X2 | Generative model designing drug-like antibodies with strong binding affinity and low immunogenicity. | arXiv |
| ProDiT | Generating functional proteins with a multimodal diffusion transformer | Multimodal diffusion Transformer protein design model trained on 214 million proteins. | bioRxiv |
| SimpleDesign | SimpleDesign: Joint Model for Protein Sequence and Structure Codesign | End-to-end joint model for protein sequence-structure co-design without tokenizers. | ICLR |
| PXDesign | PXDesign: Fast De Novo Design of Protein Binders | ByteDance Protenix fast and modular pipeline for de novo protein binder design with 20–73% hit rates. | bioRxiv |
Foundation models for peptide design, antimicrobial peptides, and cyclic peptides.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| PepMLM | Target Sequence-Conditioned Generation of Therapeutic Peptide Binders via Span Masked Language Modeling | ESM-2 fine-tuned span masked LM generating linear peptide binders conditioned on target protein sequences. | Nature Biotechnology / ICLR |
| PepBERT | PepBERT: Lightweight language models for peptide representation | Lightweight dedicated peptide language model for bioactive peptide discovery and representation learning. | bioRxiv |
| PepDoRA | PepDoRA: A Unified Peptide Language Model via Weight-Decomposed Low-Rank Adaptation | Unified peptide language model predicting multiple peptide properties via weight-decomposed low-rank adaptation. | arXiv |
| AMP-Designer | A foundation model approach to guide antimicrobial peptide design in the era of AI-driven scientific discovery | LLM-based foundation model for de novo design of antimicrobial peptides with significant antibacterial activity. | arXiv |
| AMP-Diffusion | AMP-Diffusion: Integrating Latent Diffusion with Protein Language Models for Antimicrobial Peptide Generation | Latent diffusion model integrated with ESM-2 for generating novel antimicrobial peptides. | bioRxiv |
| deepAMP | A Foundation Model Identifies Broad-Spectrum Antimicrobial Peptides against Drug-Resistant Bacterial Infection | Deep generative framework based on peptide language models identifying broad-spectrum antimicrobial peptides. | Nature Communications |
| RFpeptides | Accurate de novo design of high-affinity protein-binding macrocycles | RoseTTAFold-based denoising diffusion model for de novo macrocyclic peptide design. | Nature Chemical Biology |
| AfCycDesign | Cyclic peptide structure prediction and design using AlphaFold2 | AlphaFold2-based method for cyclic peptide structure prediction, redesign, and de novo generation. | Nature Communications |
| CP-Composer | Zero-Shot Cyclic Peptide Design via Composable Geometric Constraints | Framework for zero-shot cyclic peptide design using composable geometric constraints. | ICML |
| CpSDE | Designing Cyclic Peptides via Harmonic SDE with Atom-Bond Modeling | Cyclic peptide design method using harmonic stochastic differential equations with atom-bond modeling. | ICML |
| PDeepPP | A general language model for peptide identification | General deep learning framework for peptide function prediction combining pretrained PLMs with Transformer-CNN architecture. | arXiv |
Models predicting protein–protein interactions and complex structures.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| PLM-interact | PLM-interact: extending protein language models to predict protein-protein interactions | Extension of protein language models for PPI prediction by jointly encoding protein pairs from sequences alone. | Nature Communications |
| IntFold | IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction | Controllable biomolecular structure prediction foundation model matching AlphaFold3 accuracy with support for PPI complexes and allosteric states. | arXiv |
Models capturing protein thermodynamics, conformational dynamics, and molecular dynamics.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| ProTDyn | ProTDyn: a foundation Protein language model for Thermodynamics and Dynamics | Protein thermodynamics and dynamics foundation model unifying conformational ensemble generation and multi-timescale dynamics modeling. | NeurIPS |
| SeqDance | Learning Biophysical Dynamics with Protein Language Models | Protein language model incorporating biophysical dynamics, trained on MD simulations and normal mode analyses of 64,000+ proteins. | bioRxiv |
| ESMDance | Learning Biophysical Dynamics with Protein Language Models | ESM-2 fine-tuned variant for protein conformational dynamics prediction. | bioRxiv |
| MD-LLM-1 | MD-LLM-1: A Large Language Model for Molecular Dynamics | First molecular dynamics LLM, fine-tuning Mistral 7B for protein conformational dynamics prediction. | arXiv |
| VibeGen | Agentic End-to-End De Novo Protein Design for Tailored Dynamics Using a Language Diffusion Model | Language diffusion model framework for end-to-end protein design targeting specific vibrational dynamics. | arXiv |
| DynamicsPLM | Learning Protein Representations with Conformational Dynamics | Protein language model learning from computationally generated conformational dynamics ensembles. | bioRxiv |
| DPLM-2 | DPLM-2: A Multimodal Diffusion Protein Language Model | Multimodal discrete diffusion protein language model jointly modeling amino acid sequences and 3D structures. | NeurIPS 2025 |
| METL | Biophysics-based protein language models for protein engineering | Mutation effect transfer learning framework integrating biophysical modeling and machine learning for protein engineering. | Nature Methods |
Pretrained models for RNA sequence representation, structure inference, and function prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| RNA-FM | Interpretable RNA foundation model from unannotated data for highly accurate RNA structure and function predictions | RNA foundation model pretrained on 23 million non-coding RNA sequences for structure and function prediction. | Nature Methods |
| RiNALMo | RiNALMo: general-purpose RNA language models can generalize well on structure prediction tasks | 650-million-parameter general-purpose RNA language model trained on 36 million non-coding RNA sequences with strong structure prediction generalization. | Nature Communications |
| ERNIE-RNA | ERNIE-RNA: an RNA language model with structure-enhanced representations | Modified BERT-based RNA language model incorporating base-pairing structure information for enhanced representations. | Nature Communications |
| RNA-MSM | Multiple sequence alignment-based RNA language model and its application to structural inference | MSA-based RNA language model leveraging co-evolutionary information from homologous RNAs for structure inference. | Nucleic Acids Research |
| UNI-RNA | UNI-RNA: Universal pre-trained models revolutionize RNA research | Universal pretrained RNA model trained on the largest RNA sequence dataset supporting multiple downstream tasks. | bioRxiv |
| RNAErnie | Multi-purpose RNA language modelling with motif-aware pretraining and type-guided fine-tuning | Multi-purpose RNA language model combining motif-aware pretraining with RNA-type-guided fine-tuning. | Nature Machine Intelligence |
| AIDO.RNA | A large-scale foundation model for RNA function and structure prediction | 1.6-billion-parameter RNA foundation model trained on 42 million non-coding RNA sequences at single-nucleotide resolution. | bioRxiv |
| BiRNA-BERT | BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization | Adaptive tokenization RNA language model overcoming limitations in sequence length and diversity. | Communications Biology |
| mRNABERT | mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset | Language model specifically designed for mRNA sequence engineering with a dual-tokenization scheme. | Nature Communications |
| CodonBERT | CodonBERT: Large language models for mRNA design and optimization | Codon-level tokenized mRNA language model for mRNA design and optimization. | BEACON Benchmark |
| NucleicBERT | NucleicBERT: A large language model for RNA structure prediction | BERT-based self-supervised masked language model for RNA sequence analysis and structure prediction. | bioRxiv |
| MP-RNA | MP-RNA: Unleashing multi-species RNA foundation model via calibrated secondary structure prediction | Multi-species RNA foundation model emphasizing calibrated secondary structure prediction. | EMNLP 2024 |
| OmniGenome | Bridging sequence-structure alignment in RNA foundation models | RNA foundation model precisely aligning RNA sequences with secondary structures for bidirectional mapping. | arXiv |
| structRFM | A fully open structure-guided RNA foundation model for robust structural and functional inference | Fully open-source structure-guided RNA foundation model integrating sequence and secondary structure for robust inference. | bioRxiv |
| SpliceBERT | SpliceBERT: a pre-trained RNA language model for analyzing vertebrate splicing | Pretrained RNA language model specialized for splicing analysis in vertebrates. | Genome Biology |
| PlantRNA-FM | An interpretable RNA foundation model for exploring functional RNA motifs in plants | Interpretable RNA foundation model for plant biology trained on 1,124+ plant species. | Nature Machine Intelligence |
| AllSplice | Perturbation-aware predictive modeling of RNA splicing using bidirectional transformers | Perturbation-aware bidirectional transformer for RNA splicing prediction. | bioRxiv |
| HydraRNA | HydraRNA: A hybrid architecture based full-length RNA language model | Hybrid-architecture full-length RNA language model combining multiple architecture strengths for processing long RNA sequences. | Genome Biology |
| EVA-RNA | EVA-RNA: A Scaling Cross-Species Transcriptomic Foundation Model for Immunology & Inflammation | Cross-species transcriptomic foundation model trained on 500K+ human and mouse samples for immunology and inflammation research. | OpenReview |
| GRNFormer | GRNFormer: A Biologically-Guided Framework for Integrating Gene Regulatory Networks into RNA Foundation Models | Biologically-guided framework integrating gene regulatory networks into RNA foundation model training. | Findings of the Association for Computational Linguistics: ACL 2025 |
| BMFM-RNA | BMFM-RNA: Whole-cell expression decoding improves transcriptomic foundation models | IBM biomedical foundation model RNA module using whole-cell expression decoding to enhance transcriptomic models. | arXiv (IBM) |
| CodonFM | CodonFM: Foundation Models for Codons | Codon foundation model trained on 130 million protein-coding sequences across 20,000+ species by NVIDIA and Arc Institute. | GitHub |
| Orthrus | Orthrus: Evolutionary and Functional RNA Foundation Models | Mamba-based RNA foundation model pretrained with biologically-augmented contrastive learning. | bioRxiv |
Deep learning methods for RNA 2D and 3D structure prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| RhoFold+ | Accurate RNA 3D structure prediction using a language model-based deep learning approach | Pretrained RNA language model-based approach for accurate RNA 3D structure prediction trained on 23.7 million sequences. | Nature Methods |
| RNA-FrameFlow | Flow Matching for de novo 3D RNA Backbone Design | SE(3) flow matching method for de novo 3D RNA backbone structure generation. | arXiv |
| DRfold2 | Ab initio RNA structure prediction with composite language model | Deep learning framework for ab initio RNA 3D structure prediction using a composite language model. | bioRxiv |
| 3DRNALM | Accurate RNA 3D structure prediction using a language model-based framework | Language model-based framework for accurate RNA 3D structure prediction. | Nature Communications |
| NuFold | NuFold: end-to-end approach for RNA tertiary structure prediction | End-to-end deep learning model for RNA tertiary structure prediction. | Nature Communications |
Generative models for RNA therapeutic sequence design and structure co-design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| RNAGenesis | RNAGenesis: A Generalist Foundation Model for Functional RNA Therapeutics | Generalist RNA therapeutics foundation model unifying sequence representation, structure prediction, and functional de novo design. | bioRxiv |
| GEMORNA | Deep generative models design mRNA sequences with enhanced translational capacity and stability | Deep generative model designing mRNA sequences with enhanced translational capacity and stability. | Science |
| EVA | A Long-Context Generative Foundation Model Deciphers RNA Design Principles | 1.4-billion-parameter MoE generative foundation model trained on 114 million full-length RNA sequences for long-context RNA design. | bioRxiv |
| RNACG | RNACG: A Universal RNA Sequence Conditional Generation model based on Flow-Matching | Universal RNA sequence conditional generation model based on flow matching. | arXiv |
| RiboFlow | RiboFlow: Conditional De Novo RNA Co-Design via Synergistic Flow Matching | Synergistic flow matching framework for RNA sequence-structure co-design targeting ligand binding. | NeurIPS |
| RiboGen | RiboGen: RNA Sequence and Structure Co-Generation with Equivariant MultiFlow | Equivariant multi-flow model simultaneously generating RNA sequences and full-atom 3D structures. | ICLR |
| SANDSTORM | Generative and predictive neural networks for the design of functional RNA molecules | Neural network for RNA function prediction integrating sequence and secondary structure data. | Nature Communications |
| GARDN | Generative and predictive neural networks for the design of functional RNA molecules | Generative neural network for functional RNA design, paired with SANDSTORM for prediction. | Nature Communications |
| mRNA-GPT | Large generative mRNA language foundation model for efficient coding sequence generation and design | 302M-parameter GPT-2-based mRNA generative language model for coding sequence generation across three biological domains. | bioRxiv |
| codonGPT | codonGPT: reinforcement learning on a generative language model enables scalable mRNA design | Reinforcement learning plus generative language model for scalable mRNA codon optimization design. | Nucleic Acids Research |
| RiboDecode | Deep generative optimization of mRNA codon sequences for enhanced mRNA translation and therapeutic efficacy | Deep generative optimization framework for mRNA codon sequences enhancing translation efficiency and therapeutic efficacy. | Nature Communications |
Foundation models for DNA sequence understanding, variant effect prediction, and gene regulation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Evo | Sequence modeling and design from molecular to genome scale with Evo | Arc Institute DNA foundation model processing molecular-to-genome-scale sequences (>650K tokens). | Science |
| Evo 2 | Genome modelling and design across all domains of life with Evo 2 | DNA foundation model trained on 9 trillion base pairs covering all domains of life with 1-million-token context window. | Nature |
| DNABERT | DNABERT: Pre-trained bidirectional encoder representations from transformers model for DNA-language in genome | First BERT-based pretrained model for genomic DNA sequences using k-mer tokenization. | Bioinformatics |
| DNABERT-2 | DNABERT-2: Efficient foundation model and benchmark for multi-species genome | Upgraded DNABERT using BPE tokenization instead of k-mer for multi-species genome analysis. | ICLR 2024 |
| Nucleotide Transformer | Nucleotide Transformer: Building and evaluating robust foundation models for human genomics | Large-scale genomic foundation model (50M–2.5B parameters) from InstaDeep trained on 3,200+ human genomes. | Nature Methods |
| HyenaDNA | HyenaDNA: Long-range genomic sequence modeling at single nucleotide resolution | Hyena implicit convolution-based genomic model for single-nucleotide-resolution long-range modeling up to 1 million bp. | NeurIPS 2023 |
| Enformer | Effective gene expression prediction from sequence by integrating long-range interactions | Transformer model from DeepMind/Calico predicting gene expression and chromatin states from DNA sequences. | Nature Methods |
| Caduceus | Caduceus: Bi-directional equivariant long-range DNA sequence modeling | Bidirectional Mamba-based DNA language model supporting reverse complement equivariance. | ICML |
| GenSLMs | GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics | Genome-scale language model pretrained on 110 million prokaryotic gene sequences analyzing SARS-CoV-2 evolution (Gordon Bell Prize). | IJHPCA |
| GROVER | DNA language model GROVER learns sequence context in the human genome | DNA language model trained on the human genome using BPE to define DNA vocabulary and capture CpG methylation features. | Nature Machine Intelligence |
| Sei | A sequence-based global map of regulatory activity for deciphering human genetics | Deep learning framework predicting 21,900+ chromatin features and mapping sequences to 40 regulatory activity classes. | Nature Genetics |
| GPN | DNA language models are powerful predictors of genome-wide variant effects | Unsupervised DNA language model predicting genome-wide variant effects from genomic sequences. | PNAS |
| Borzoi | Borzoi decodes the complex DNA signals governing gene regulation | Deep learning model predicting RNA-seq coverage from DNA sequences to decode gene regulatory signals. | Nature Genetics |
| PlantCaduceus | Cross-species modeling of plant genomes at single-nucleotide resolution using a pretrained DNA language model | Plant-specific DNA language model trained on 16 angiosperm genomes supporting cross-species analysis. | PNAS |
| HybriDNA | HybriDNA: A hybrid Transformer-Mamba2 DNA language model | Hybrid Transformer-Mamba2 DNA language model supporting ultra-long sequences (131kb) at single-nucleotide resolution. | arXiv |
| Nucleotide Transformer v3 (NTv3) | A foundational model for joint sequence-function multi-species prediction | Multi-species long-range genomic prediction and functional annotation foundation model from InstaDeep. | bioRxiv |
| Basenji | Sequential regulatory activity prediction across chromosomes with convolutional neural networks | CNN for predicting gene expression and regulatory activity from DNA sequences. | Genome Research |
| Basenji2 | Cross-species regulatory sequence activity prediction | Cross-species DNA regulatory activity prediction model. | PLoS Computational Biology |
| ChromBPNet | ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility | Base-resolution deep learning model for chromatin accessibility prediction with bias factorization. | bioRxiv |
| scBasset | scBasset: Sequence-based modeling of single-cell ATAC-seq using convolutional neural networks | CNN for modeling single-cell ATAC-seq chromatin accessibility from DNA sequences. | Nature Methods |
| EpiGePT | EpiGePT: a Pretrained Transformer model for epigenomics | Pretrained transformer model for epigenomics data analysis and prediction. | bioRxiv |
| AgroNT | A foundational large language model for edible plant genomes | Crop-specific genomic foundation model for plant genomics and breeding applications. | Communications Biology |
| AlphaGenome | Advancing regulatory variant effect prediction with AlphaGenome | Megabase-scale DNA model from Google DeepMind predicting gene expression and chromatin signals for variant effect analysis. | Nature |
| GENA-LM | GENA-LM: A Family of Open-Source Foundational DNA Language Models for Long Sequences | Open-source family of transformer DNA language models supporting up to 36k bp sequences. | bioRxiv/Bioinformatics |
| DNAGPT | DNAGPT: A Generalized Pre-trained Tool for Multiple DNA Sequence Analysis Tasks | Generative pretrained model for multiple DNA analysis tasks. | bioRxiv/PLoS ONE |
| MoDNA | MoDNA: Motif-Oriented Pre-training For DNA Language Model | Motif-oriented pretrained DNA language model capturing regulatory motif patterns. | ACM BCB |
| GPN-MSA | GPN-MSA: An alignment-based DNA language model for genome-wide variant effect prediction | DNA language model using multi-species alignment for genome-wide variant effect prediction. | Nature Biotechnology |
| MergeDNA | MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging | Hierarchical dynamic tokenization approach for context-aware genome modeling. | AAAI |
| BMFM-DNA | BMFM-DNA: A SNP-aware DNA foundation model to capture variant effects | IBM SNP-aware DNA foundation model for capturing variant effects in genomic sequences. | arXiv |
| EpiAgent | EpiAgent: foundation model for single-cell epigenomics | Foundation model for single-cell ATAC-seq epigenomic data analysis. | Nature Methods |
| Gene42 | Gene42: Long-Range Genomic Foundation Model With Dense Attention | Decoder-only long-range genomic foundation model processing up to 192,000 bp at single-nucleotide resolution. | arXiv |
| dnaHNet | dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning | Scalable hierarchical foundation model for genomic sequence learning. | arXiv |
| JEPA-DNA | JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures | Genomic foundation model based on joint-embedding predictive architecture combining generative and discriminative objectives. | arXiv |
| GeneZip | GeneZip: Region-Aware Compression for Long Context DNA Modeling | Region-aware compression method for long-context DNA sequence modeling. | arXiv |
| ModernGENA | ModernGENA: A modernized BERT-style DNA foundation model | Modernized BERT architecture adapted for DNA foundation modeling (ModernBERT for genomics). | OpenReview |
| OpticalDNA | OpticalDNA: Reimagining DNA sequence analysis as an OCR task | Novel framework reimagining DNA sequence analysis as an optical character recognition task. | arXiv |
| OmniReg-GPT | OmniReg-GPT: A Generative Pre-trained Model for Universal Gene Regulation Prediction | Generative pretrained model for universal gene regulation prediction across species and tissues. | Nature Communications |
| BOTANIC-0 | BOTANIC-0: A Plant Genomic Foundation Model | Family of plant genomic foundation models (100M–1B parameters) pretrained on 1,600+ curated plant genomes. | bioRxiv |
| Species-aware DNA LM | Species-aware DNA Language Modeling | DNA language model incorporating species-specific information during pretraining. | bioRxiv |
| Genos | Genos: A Large Human-Centric Genomic Foundation Model | Large-scale (up to 10B parameters) human-centric genomic foundation model with MoE-Transformer architecture from BGI. | GigaScience |
| GENERator | GENERator: A Long-Context Generative Genomic Foundation Model | Long-context generative genomic foundation model pretrained on 386 billion nucleotides with 98k context length. | arXiv |
| BioReason | BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model | First deep integration of DNA foundation models (Nucleotide Transformer/Evo2) with LLMs for multi-step biological reasoning; raises KEGG disease pathway prediction from 86% to 98% with interpretable reasoning traces. | NeurIPS 2025 |
Foundation models for single-cell transcriptomics, perturbation prediction, and virtual cell modeling.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| scGPT | scGPT: toward building a foundation model for single-cell multi-omics using generative AI | Generative pretrained transformer for single-cell multi-omics, enabling cell annotation, perturbation prediction, and gene network inference from 33M+ cells. | Nature Methods |
| scBERT | scBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data | Large-scale BERT-based pretrained model for automated cell type annotation from scRNA-seq data. | Nature Machine Intelligence |
| Geneformer | Transfer learning enables predictions in network biology | Transformer model pretrained on ~30 million single-cell transcriptomes for transfer learning in gene network biology. | Nature |
| scFoundation | Large-scale foundation model on single-cell transcriptomics | 100M-parameter single-cell transcriptomics foundation model (xTrimoGene) trained on 50M+ human cells. | Nature Methods |
| UCE | Universal Cell Embeddings: A foundation model for cell biology | Universal cell embedding model that creates a unified representation space across species and tissues. | bioRxiv |
| SCimilarity | A cell atlas foundation model for scalable search of similar human cells | Deep metric-learning foundation model for single-cell profiles, enabling rapid similarity search and annotation across a 23.4M-cell human atlas from 412 scRNA-seq studies. | Nature |
| GeneCompass | GeneCompass: Deciphering universal gene regulatory mechanisms with a knowledge-informed cross-species foundation model | Knowledge-enhanced cross-species foundation model trained on 100M+ human and mouse cells for deciphering gene regulatory mechanisms. | Cell Research |
| CellPLM | CellPLM: Pre-training of cell language model beyond single cells | Cell language model that integrates gene-gene and cell-cell interactions, going beyond single-cell level pretraining. | ICLR 2024 |
| tGPT | Generative pretraining from large-scale transcriptomes for single-cell deciphering | Generative pretrained model on 22.3 million single-cell transcriptomes for cell deciphering and clinical translation. | iScience |
| CellFM | CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells | 800M-parameter foundation model pretrained on 100 million human cell transcriptomes. | Nature Communications |
| Nicheformer | Nicheformer: a foundation model for single-cell and spatial omics | Single-cell and spatial omics foundation model trained on SpatialCorpus-110M (110M+ cells), capturing spatial microenvironments. | Nature Methods |
| scMulan | scMulan: A multitask generative pre-trained language model for single-cell analysis | Multitask generative pretrained language model that encodes cells as structured "c-sentences" for single-cell analysis. | RECOMB 2024 |
| Cell2Sentence | Cell2Sentence: Teaching large language models the language of biology | Converts gene expression profiles into natural language sentences, enabling GPT-2 adaptation for single-cell transcriptomics. | ICML 2024 |
| C2S-Scale | Scaling large language models for next-generation single-cell analysis | Scales the Cell2Sentence framework with larger models and broader data for improved single-cell RNA-seq analysis. | bioRxiv |
| GET | GET: A foundation model of transcription across human cell types | Universal expression transformer predicting gene expression from chromatin accessibility across 213 human cell types. | Nature |
| GenePT | GenePT: A simple but effective foundation model for genes and cells built from ChatGPT | Simple and effective gene/cell foundation model using ChatGPT embeddings for gene and cell representation. | bioRxiv |
| LangCell | LangCell: Language-cell pre-training for cell identity understanding | Joint pretraining of natural language and single-cell transcriptomics to enhance cell identity understanding. | arXiv |
| CellVQ | Illuminating cell states by a comprehensive and interpretable single cell foundation model | Comprehensive and interpretable single-cell foundation model trained on 68 million cells for illuminating cell states. | Nature Communications |
| scKGBERT | scKGBERT: A knowledge-enhanced foundation model for single-cell transcriptomics | Knowledge graph-enhanced foundation model for single-cell transcriptomics. | Genome Biology |
| SATURN | Toward universal cell embeddings: integrating scRNA-seq datasets across species | Cross-species universal cell embedding framework that integrates scRNA-seq datasets across organisms. | Nature Methods |
| scTab | scTab: Scaling cross-tissue single-cell annotation models | Scalable deep learning model for cross-tissue cell type annotation trained on 22 million cells. | Nature Communications |
| scPRINT | scPRINT: pre-training on 50 million cells allows robust gene network predictions | Transformer foundation model trained on 50M cells for robust gene network inference. | Nature Communications |
| scPRINT-2 | scPRINT-2: Towards the next-generation of cell foundation models | Next-generation cell foundation model trained on 350M cells across 16 organisms. | bioRxiv |
| scELMo | scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis | Language model embeddings applied to single-cell data analysis. | bioRxiv |
| scVI | scVI: Variational Inference for Single-Cell Gene Expression | Deep generative model providing probabilistic framework for single-cell transcriptomics analysis. | Nature Methods |
| scANVI | scANVI: semi-supervised integration of single-cell multi-omic data | Semi-supervised deep generative model for integrating single-cell multi-omic data. | Molecular Systems Biology |
| totalVI | Joint probabilistic modeling of single-cell multi-omic data with totalVI | Joint probabilistic model for simultaneous RNA and protein single-cell data analysis. | Nature Methods |
| scPoli | Population-level integration of single-cell datasets enables multi-scale analysis | Population-level single-cell dataset integration enabling multi-scale biological analysis. | Nature Methods |
| scHyena | scHyena: Foundation Model for Full-Length Single-Cell RNA-Seq Analysis in Brain | Hyena architecture-based foundation model for full-length scRNA-seq analysis in brain tissue. | arXiv |
| TOSICA | Transformer for one stop interpretable cell type annotation | Transfer learning framework for single-cell omics analysis across datasets and modalities. | Nature Communications |
| xTrimoGene | xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data | Efficient and scalable representation learner for scRNA-seq data using asymmetric encoder-decoder architecture. | NeurIPS 2023 |
| CancerFoundation | A single-cell RNA sequencing foundation model to decipher drug resistance in cancer | Cancer-specific scRNA-seq foundation model for deciphering drug resistance mechanisms. | bioRxiv |
| Cell-GraphCompass | Cell-GraphCompass: Modeling Single Cells with Graph Structure Foundation Model | Graph structure foundation model for single-cell analysis using graph-based cell representations. | National Science Review |
| scLong | scLong: A billion-parameter foundation model for capturing long-range gene context | Billion-parameter foundation model attending to all 28,000 genes simultaneously for long-range context. | Nature Communications |
| Tahoe-x1 | Tahoe-x1: Scaling Perturbation-Trained Single-Cell Foundation Models to 3 Billion Parameters | 3B-parameter perturbation-trained single-cell foundation model for predicting cellular responses. | bioRxiv |
| PULSAR | PULSAR: a Foundation Model for Multi-scale and Multicellular Biology | Multi-scale foundation model integrating 36M+ cells for multicellular biology analysis. | bioRxiv |
| TranscriptFormer | A Cross-Species Generative Cell Atlas Across 1.5 Billion Years of Evolution | Cross-species generative cell atlas foundation model spanning 1.5 billion years of evolution. | bioRxiv |
| TCRfoundation | TCRfoundation: A multimodal foundation model for single-cell immune profiling | Multimodal foundation model integrating gene expression with TCR sequences for immune profiling. | GitHub |
| CELLama | CELLama: Foundation Model for Single Cell and Spatial Transcriptomics | Cell embedding model leveraging language model capabilities for single-cell and spatial transcriptomics. | bioRxiv |
| scPROTEIN | scPROTEIN: versatile deep graph contrastive learning framework for single-cell proteomics | Graph contrastive learning framework for single-cell proteomics embedding and analysis. | Nature Methods |
| TEDDY | TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology | Family of foundation models designed for comprehensive single-cell biology understanding. | ICML Workshop |
| Tabula | Tabula: A Tabular Self-Supervised Foundation Model for Single-Cell Transcriptomics | Tabular self-supervised foundation model tailored for single-cell transcriptomics data. | NeurIPS 2025 |
| ChromFound | ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibility Data | Universal foundation model for single-cell chromatin accessibility (scATAC-seq) data analysis. | NeurIPS |
| EpiFoundation | EpiFoundation: A Foundation Model for Single-Cell ATAC-seq via Peak-to-Gene Alignment | Foundation model for scATAC-seq using peak-to-gene alignment for epigenomic analysis. | bioRxiv |
| SCARF | SCARF: Single Cell ATAC-seq and RNA-seq Foundation model | Multi-modal foundation model jointly modeling scATAC-seq and scRNA-seq data. | bioRxiv |
| CAPTAIN | CAPTAIN: A multimodal foundation model pretrained on co-assayed single-cell RNA and protein | Multimodal foundation model for co-assayed scRNA and protein data integration. | bioRxiv |
| OKR-Cell | OKR-Cell: Open world knowledge aided single-cell foundation model with cross-modal pre-training | Open-world knowledge-enhanced single-cell foundation model with cross-modal pretraining. | arXiv |
| GeneJepa | GeneJepa: A predictive world model of the transcriptome | Predictive world model for transcriptomics based on JEPA architecture. | arXiv |
| VCWorld | VCWorld: Biological world model for virtual cell simulation | Biological world model for simulating virtual cell dynamics and perturbation responses. | arXiv |
| CellHermes | CellHermes: Harmonizing multimodal data for omics understanding | Multimodal data harmonization model for unified omics understanding. | arXiv |
| CellTok | CellTok: Early-fusion multimodal LLM for single-cell transcriptomics via tokenization | Early-fusion multimodal LLM for single-cell transcriptomics using gene expression tokenization. | arXiv |
| sciLaMA | sciLaMA: Single-cell representation learning leveraging prior knowledge from LLMs | Single-cell representation learning leveraging prior knowledge from large language models. | arXiv |
| scConcept | scConcept: Contrastive pretraining for technology-agnostic single-cell representations | Contrastive pretraining framework for technology-agnostic single-cell representations. | arXiv |
| scLinguist | scLinguist: Hyena-based foundation model for cross-modality translation in single-cell multi-omics | Hyena-based foundation model for cross-modality translation in single-cell multi-omics. | arXiv |
| scNET | scNET: Context-specific gene and cell embeddings by integrating scRNA with PPI | Context-specific gene and cell embeddings integrating scRNA-seq with protein-protein interaction networks. | arXiv |
| scLAMBDA | scLAMBDA: Modeling single-cell multi-gene perturbation responses | Model for predicting single-cell responses to multi-gene combinatorial perturbations. | arXiv |
| GeneMamba | GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data | Mamba architecture-based efficient single-cell foundation model with scalable computation. | arXiv |
| Atacformer | Atacformer: Transformer-based foundation model for ATAC-seq data analysis | Transformer-based foundation model for ATAC-seq chromatin accessibility data analysis. | arXiv |
| CLM-X | CLM-X: Cross-Modal Language Model for Single-Cell Multi-Omics | Cross-modal language model for unified single-cell multi-omics representation learning. | arXiv |
| CellOracle | CellOracle: Dissecting cell identity via network inference and in silico gene perturbation | Computational framework using gene regulatory networks to simulate gene perturbation effects on cell identity. | Nature |
| Stack | Stack: In-Context Learning of Single-Cell Biology | Arc Institute single-cell foundation model trained on 149M human cells enabling zero-shot prediction via in-context learning. | bioRxiv |
| Lingshu-Cell | Lingshu-Cell: cellular world model for transcriptome modeling | Masked discrete diffusion cellular world model for transcriptome modeling from Alibaba DAMO. | arXiv |
| OmniCell | OmniCell: Unified Foundation Modeling of Single-Cell and Spatial Transcriptomics | Unified foundation model for both single-cell and spatial transcriptomics analysis. | bioRxiv |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AlphaCell | Towards building a World Model to simulate perturbation-induced cellular dynamics | Virtual cell world model that simulates perturbation-induced cellular dynamics. | bioRxiv |
| CellFluxV2 | CellFluxV2: An Image Generative Foundation Model for Virtual Cell Modeling | Flow matching-based generative foundation model for virtual cell image modeling. | bioRxiv |
| X-Cell | X-Cell: Scaling Causal Perturbation Prediction Across Diverse Cellular Contexts | Large-scale diffusion language model predicting genome-wide transcriptional responses across diverse cellular contexts. | bioRxiv |
Foundation models that integrate molecular, cellular, and tissue-level information across biological scales.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Xpressor | Towards foundation models that learn across biological scales | Cross-scale learning framework integrating molecular, cellular, and tissue-level gene expression via cross-attention. | bioRxiv |
| AIDO | Toward AI-driven digital organism: Multiscale foundation models for predicting, simulating and programming biology at all levels | AI-driven digital organism system integrating DNA→RNA→protein→cell multi-scale foundation models. | arXiv |
| CDT (Central Dogma Transformer) | Central Dogma Transformer: Towards Mechanism-Oriented AI for Cellular Understanding | Architecture integrating pretrained DNA (Enformer), RNA (scGPT), and protein (ProteomeLM) models via directional cross-attention mirroring the central dogma information flow, producing unified Virtual Cell Embeddings. | arXiv |
| CDT-II | Central Dogma Transformer II: An AI Microscope for Understanding Cellular Regulatory Mechanisms | AI microscope with DNA/RNA self-attention and cross-attention for transcriptional control; achieves per-gene mean r=0.84 on K562 CRISPRi data, recovers GFI1B regulatory network (6.6× enrichment), and predicts therapeutic target consequences via gradient attribution. | arXiv |
| CDT-III | Central Dogma Transformer III: Interpretable AI Across DNA, RNA, and Protein | Two-stage Virtual Cell Embedder (VCE-N for nuclear transcription, VCE-C for cytosolic translation) extending to full central dogma with protein prediction; achieves RNA r=0.843 and protein r=0.969, rediscovers 5/7 known Alemtuzumab side effects without clinical data. | arXiv |
Foundation models for antibody engineering, structure prediction, and immune receptor analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| IgBERT | Large scale paired antibody language models | BERT-based antibody language model trained on 2B+ unpaired and 2M paired antibody sequences from OAS. | PLOS Computational Biology |
| IgT5 | Large scale paired antibody language models | T5-based paired antibody language model companion to IgBERT for antibody design and engineering. | PLOS Computational Biology |
| AntiBERTy | Deciphering antibody affinity maturation with language models and weakly supervised learning | BERT-based model trained on 558M antibody sequences for affinity maturation analysis. | arXiv |
| AbLang | AbLang: an antibody language model for completing antibody sequences | Antibody-specific language model trained on OAS for residue prediction and antibody representation. | Bioinformatics Advances |
| AbLang2 | Addressing the antibody germline bias and its effect on language models for improved antibody design | Improved antibody language model for paired heavy-light chains with reduced germline bias. | Bioinformatics |
| IGLOO | Tokenizing Loops of Antibodies | Multimodal antibody loop tokenizer enhancing protein language models for antibody research. | NeurIPS 2025 Workshop |
| Ab-RoBERTa | Antibody Foundational Model : Ab-RoBERTa | RoBERTa-based antibody language model for paratope prediction and antibody design. | arXiv |
| BALM | Accurate Prediction of Antibody Function and Structure Using Bio-Inspired Antibody Language Model | Bio-inspired antibody language model for predicting antibody structure and function. | Science Advances |
| DASM | Separating selection from mutation in antibody language models | Deep amino acid selection model that separates selection from mutation in antibody sequence modeling. | eLife |
| nanoBERT | nanoBERT: a deep learning model for gene agnostic navigation of the nanobody mutational space | Nanobody-specific transformer model for predicting amino acid substitutions in VHH sequences. | Bioinformatics Advances |
| FAbCon | A generative foundation model for antibody sequence understanding | 2.4B-parameter generative foundation model for antibody sequence understanding. | bioRxiv |
| S2ALM | S2ALM: Sequence-Structure Pre-trained Large Language Model for Antibody | Sequence-structure pretrained large language model for comprehensive antibody understanding. | Research |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| IgFold | Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies | Fast antibody structure prediction using pretrained LM on 558M sequences with graph neural networks. | Nature Communications |
| DeepAb | Antibody structure prediction using interpretable deep learning | Interpretable deep learning model for antibody Fv structure prediction specializing in CDR loop modeling. | Patterns |
| ABlooper | ABlooper: fast accurate antibody CDR loop structure prediction with accuracy estimation | Rapid equivariant neural network for antibody CDR loop structure prediction with accuracy estimation. | Bioinformatics |
| AntiFold | AntiFold: Improved structure-based antibody design using inverse folding | Antibody-specific inverse folding model fine-tuned from ESM-IF1 for CDR sequence generation from structures. | Bioinformatics Advances |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| IgGM | A generative foundation model for antibody design | Generative foundation model for comprehensive antibody design. | bioRxiv |
| DiffAb | Antigen-Specific Antibody Design and Optimization with Diffusion-Based Generative Models | Diffusion-based generative model for antigen-specific antibody CDR-H3 design, jointly modeling sequence, structure, and orientation. | NeurIPS 2022 |
| dyMEAN | Full-Atom Antibody Design via dyMEAN | End-to-end full-atom antibody design using dynamic multi-channel equivariant graph network. | ICML |
| MEAN | Conditional Antibody Design as 3D Equivariant Graph Translation | 3D equivariant graph neural network for conditional antibody CDR sequence-structure co-design. | ICLR 2023 |
| RefineGNN | Iterative Refinement Graph Neural Network for Antibody Sequence-Structure Co-design | Iterative refinement GNN for antibody CDR co-design of sequence and 3D structure via autoregressive generation. | ICLR |
| Ophiuchus-Ab | Ophiuchus-Ab: A Versatile Generative Foundation Model for Advanced Antibody-Based Immunotherapy | Diffusion language model for antibody immunotherapy and paired antibody repertoire generation. | bioRxiv |
| NanoAbLLaMA | NanoAbLLaMA: construction of nanobody libraries with protein large language models | LLaMA2-based language model fine-tuned for nanobody (VHH) library construction and design. | Frontiers in Chemistry |
| CoSiNE | CoSiNE: Conditionally Site-Independent Neural Evolution of Antibody Sequences | Conditionally site-independent neural evolution model explicitly modeling antibody affinity maturation. | arXiv |
| AbBFN2 | AbBFN2: A flexible antibody foundation model based on Bayesian Flow Networks | Flexible antibody foundation model based on Bayesian flow networks for multi-objective unified modeling. | bioRxiv |
| AntibodyDesignBFN | AntibodyDesignBFN: High-Fidelity Fixed-Backbone Antibody Design via Discrete Bayesian Flow Networks | High-fidelity fixed-backbone antibody design using discrete Bayesian flow networks. | arXiv |
| AbAffinity | AbAffinity: A Large Language Model for Predicting Antibody Binding Affinity | Large language model for predicting antibody-antigen binding affinity. | arXiv |
| CALM | CALM: Cross-attention Adaptive Immune Receptor–Antigen Language Model | Cross-attention language model for antibody-antigen specificity prediction. | bioRxiv |
| JAM-2 | JAM-2: Fully computational design of drug-like antibodies | Fully computational model for designing drug-like antibodies developed by Nabla Bio. | Technical Report |
| Chai-2 | Chai-2: Zero-shot antibody discovery | Zero-shot antibody discovery model developed by Chai Discovery. | bioRxiv |
| Model | Paper Title | Description | Link |
|---|---|---|---|
| TCR-BERT | TCR-BERT: learning the grammar of T-cell receptors for flexible antigen-binding analyses | Modified BERT trained on TCR sequences via self-supervised learning for antigen-specificity prediction. | PMLR v240 |
| tcrLM | tcrLM: a lightweight protein language model for predicting T cell receptor and epitope binding specificity | Lightweight BERT-based LM pretrained on 100M+ TCR CDR3 sequences for TCR-epitope binding prediction. | arXiv |
| TCR-GPT | TCR-GPT: Integrating Autoregressive Model and Reinforcement Learning for T-Cell Receptor Repertoires Generation | Decoder-only transformer for TCR sequence generation using autoregressive modeling with reinforcement learning. | arXiv |
| SCEPTR | Contrastive learning of T cell receptor representations | Lightweight BERT-like transformer for TCR analysis using autocontrastive and masked-language pretraining. | Cell Systems |
| ERGO-II | Prediction of Specific TCR-Peptide Binding From Large Dictionaries of TCR-Peptide Pairs | Deep learning model (LSTM + autoencoder) for TCR-peptide binding prediction using NLP techniques. | Frontiers in Immunology |
| mvTCR | Multi-modal generative modeling for joint analysis of single-cell T cell receptor and gene expression data | Multimodal variational autoencoder integrating single-cell TCR sequences with gene expression data. | Nature Communications |
| NetTCR-2.0 | NetTCR-2.0 enables accurate prediction of TCR-peptide binding | Deep learning model for TCR-peptide-MHC binding prediction using paired TCRα and β sequences. | Communications Biology |
| TCR-TRANSLATE | Conditional generation of real antigen-specific T cell receptor sequences | ML framework for generating antigen-specific TCR sequences including for unseen epitopes. | Nature Machine Intelligence |
Foundation models for enzyme function prediction, kinetics modeling, and de novo enzyme design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| EnzyGen | Generative Enzyme Design Guided by Functionally Important Sites and Small-Molecule Substrates | Generative enzyme design model leveraging functional sites and substrate information. | arXiv |
| RFdiffusion2 | Atom-level enzyme active site scaffolding using RFdiffusion2 | Atom-level enzyme active site scaffolding tool based on RFdiffusion architecture. | Nature Methods |
| CLEAN | Enzyme function prediction using contrastive learning | Contrastive learning model for enzyme EC number prediction, outperforming BLAST and traditional methods. | Science |
| CLEAN-Contact | Improved enzyme functional annotation prediction using contrastive learning with structural inference | Extension of CLEAN integrating protein contact maps for improved enzyme function annotation. | Communications Biology |
| EnzBERT | Predicting enzymatic function of protein sequences with attention | BERT-based model for predicting enzyme EC numbers from protein sequences using attention mechanisms. | Bioinformatics |
| EnzymeFlow | EnzymeFlow: Generating Reaction-specific Enzyme Catalytic Pockets through Flow Matching and Co-Evolutionary Dynamics | Generative model using flow matching to design reaction-specific enzyme catalytic pockets. | NeurIPS |
| EnzymeCAGE | EnzymeCAGE: A Geometric Foundation Model for Enzyme Retrieval with Evolutionary Insights | Geometric foundation model trained on ~1M enzyme-reaction pairs for enzyme retrieval and function prediction. | bioRxiv |
| CatPred | CatPred: a comprehensive framework for deep learning in vitro enzyme kinetic parameters | Deep learning framework for predicting enzyme kinetic parameters (kcat, Km, Ki) from sequences. | Nature Communications |
| UniKP | UniKP: a unified framework for the prediction of enzyme kinetic parameters | Unified deep learning framework using pretrained protein LMs to predict kcat, Km, and catalytic efficiency. | Nature Communications |
| TurNuP | Turnover number predictions for kinetically uncharacterized enzymes using machine and deep learning | Deep learning model for predicting enzyme turnover numbers (kcat) for uncharacterized enzymes. | Nature Communications |
| EnzyControl | EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation | Framework for substrate-specific enzyme backbone generation with functional control. | arXiv |
| ProtDETR | Interpretable Enzyme Function Prediction via Residue-Level Detection | Attention-based framework for residue-level enzyme EC number prediction inspired by object detection. | arXiv |
| EZpred | EZpred: improving deep learning-based enzyme function prediction using unlabeled sequence homologs | Deep learning framework leveraging unlabeled homolog sequences for improved enzyme EC prediction. | bioRxiv |
| BEC-Pred | A general model for predicting enzyme functions based on enzymatic reactions | BERT-based model predicting enzyme EC numbers from SMILES representations of substrates and products. | Journal of Cheminformatics |
| HIT-EC | HIT-EC: Trustworthy prediction of enzyme commission numbers using a hierarchical interpretable transformer | Hierarchical interpretable transformer for trustworthy enzyme EC number prediction. | Nature Communications |
| ENZYME-UNIFIED | ENZYME-UNIFIED: Learning Holistic Representations of Enzyme Function with a Hybrid Interaction Model | Holistic enzyme function representation learning via hybrid interaction modeling. | OpenReview |
Foundation models for spatially resolved gene expression, tissue architecture, and histology-omics integration.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Novae | Novae: a graph-based foundation model for spatial transcriptomics data | Graph-based foundation model for spatial transcriptomics trained on 30 million cells. | Nature Methods |
| SpaFoundation | SpaFoundation: a visual foundation model for spatial transcriptomics | Visual foundation model for spatial transcriptomics using 1.84M histological images. | bioRxiv |
| STPath | STPath: a generative foundation model for integrating spatial transcriptomics and WSIs | Generative foundation model integrating spatial transcriptomics with whole-slide histology images. | npj Digital Medicine |
| OmiCLIP | A visual-omics foundation model to bridge histopathology with spatial transcriptomics | Visual-omics foundation model bridging histopathology images and spatial transcriptomics data. | Nature Methods |
| SpatialFusion | SpatialFusion: A lightweight multimodal foundation model for spatial transcriptomics | Lightweight multimodal foundation model integrating gene expression, histopathology, and pathway data. | bioRxiv |
| SAGE-FM | SAGE-FM: A lightweight and interpretable foundation model for spatial transcriptomics | GCN-based lightweight and interpretable foundation model for spatial transcriptomics. | arXiv |
| scGPT-spatial | scGPT-spatial: Continual Pretraining of Single-Cell FM for Spatial Transcriptomics | Extension of scGPT for spatial transcriptomics via continual pretraining. | bioRxiv |
| SpatialScope | SpatialScope: integrating spatial and single-cell transcriptomics data using deep generative models | Deep generative model for integrating spatial transcriptomics with scRNA-seq data. | Nature Communications |
| stFormer | stFormer: a foundation model for spatial transcriptomics | Transformer foundation model integrating ligand-receptor interactions into spatial gene representations. | bioRxiv |
| STAGE | STAGE: A Foundation Model for Spatial Transcriptomics Analysis via Graph Embeddings | Foundation model using graph embeddings and hierarchical prototypes for spatial transcriptomics. | OpenReview |
| STORM | STORM: A multimodal foundation model of spatial transcriptomics and histology | Multimodal spatial transcriptomics and histology foundation model trained on 1.2M spatially-resolved profiles across 18 organs. | arXiv |
| SEAL | SEAL: Spatial Expression-Aligned Learning for pathology foundation models | Spatial expression-aligned learning framework enhancing pathology foundation models with spatial transcriptomics data. | arXiv |
| MINT | MINT: Molecularly Informed Training with Spatial Transcriptomics Supervision for Pathology Foundation Models | Molecularly informed training with spatial transcriptomics supervision for pathology foundation models. | arXiv |
| HINGE | HINGE: Adapting Pre-trained Single-Cell Foundation Models to Spatial Gene Expression | Adapts pretrained single-cell foundation models to spatial gene expression using histological image conditioning. | arXiv |
| HEIST | HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data | Graph foundation model for both spatial transcriptomics and proteomics data analysis. | arXiv |
| TISSUENARRATOR | TISSUENARRATOR: Generative modeling of spatial transcriptomics with LLMs | LLM-based generative modeling framework for spatial transcriptomics data. | arXiv |
| SpaTranslator | SpaTranslator: Deep generative framework for universal spatial multi-omics cross-modality translation | Deep generative framework for universal cross-modality translation of spatial multi-omics data. | arXiv |
| SpatialProp | SpatialProp: Tissue perturbation modeling with spatially resolved single-cell transcriptomics | Tissue perturbation modeling using spatially resolved single-cell transcriptomics. | arXiv |
| SWITCH | SWITCH: Integrative deep learning of spatial multi-omics | Integrative deep learning framework for spatial multi-omics data analysis. | arXiv |
| CancerSTFormer | CancerSTFormer enables multi-scale analysis of spot-resolution spatial transcriptomes | Multi-scale spatial transcriptomics foundation model for cancer at 50µm and 250µm resolution. | bioRxiv |
| STAGATE | Deciphering spatial domains from spatially resolved transcriptomics with adaptive graph attention auto-encoder | Graph attention auto-encoder for spatial domain identification integrating gene expression and spatial location. | Nature Communications |
| CellViT | CellViT: Vision Transformers for precise cell segmentation and classification | Vision Transformer for precise cell/nuclei segmentation in H&E whole-slide images. | Medical Image Analysis |
| STAMP | Interpretable spatially aware dimension reduction of spatial transcriptomics with STAMP | Deep generative model for spatially-aware interpretable dimension reduction of spatial transcriptomics. | Nature Methods |
| SpaGT | Spatially informed graph transformers for spatially resolved transcriptomics | Graph transformer integrating spatial coordinates and gene expression for spatial domain identification. | Communications Biology |
| BrainBeacon | BrainBeacon: A Cross-Species Foundation Model for Single-cell Spatial Transcriptomics of Brain | Cross-species brain spatial transcriptomics foundation model integrating multi-species data for digital twin brain. | bioRxiv |
| SToFM | SToFM: A multi-scale foundation model for spatial transcriptomics | Multi-scale spatial transcriptomics foundation model integrating macroscopic tissue morphology and microscopic cellular environments. | ICML 2025 |
| OmniCell | OmniCell: Unified Foundation Modeling of Single-Cell and Spatial Transcriptomics | Unified foundation model for both single-cell and spatial transcriptomics analysis. | bioRxiv |
| PAST | PAST: A multimodal single-cell foundation model for histopathology and spatial transcriptomics in cancer | Multimodal single-cell foundation model integrating histopathology images and spatial transcriptomics data for cancer analysis. | arXiv |
Foundation models for glycan structure representation, protein-glycan interactions, and carbohydrate analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| GlycanGT | GlycanGT: A Foundation Model for Glycan Graphs with Pretrained Representation and Generative Learning | First glycan graph foundation model using graph transformer for glycan representation and generative learning. | bioRxiv |
| SweetBERT | Exploring BERT-based models for IUPAC glycan nomenclature | BERT-based glycan sequence language model encoding IUPAC glycan nomenclature and branching structures. | ICLR 2025 Workshop |
| GlycanAA | Modeling All-Atom Glycan Structures via Hierarchical Message Passing and Multi-Scale Pre-training | All-atom glycan modeling framework using hierarchical message passing and multi-scale pretraining. | ICML |
| MCNet | Atom-level machine learning of protein-glycan interactions and cross-chiral recognition | Atom-level machine learning model for protein-glycan interaction prediction including mirror-image glycan recognition. | Science Advances |
| DeepGlycanSite | Highly accurate carbohydrate-binding site prediction with DeepGlycanSite | High-accuracy deep learning model for predicting carbohydrate-binding sites on proteins. | Nature Communications |
| SweetNet | Using graph convolutional neural networks to learn a representation for glycans | Graph convolutional neural network for glycan representation learning handling complex branching structures. | Cell Reports |
| GlycoBERT | Transformer-based Deep Learning for Glycan Structure Inference from MS/MS | BERT-based transformer for inferring glycan structures from tandem mass spectrometry data. | bioRxiv |
Foundation models for metabolomic profiling, spectral analysis, and multi-disease prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MetaboLM | MetaboLM: a metabolomic language model for multi-disease early prediction and risk stratification | Transformer-based metabolomics language model trained on ~84,000 healthy plasma metabolomes for multi-disease early prediction. | Nature Communications |
| DSCF | Deep spectral component filtering as a foundation model for spectral analysis demonstrated in metabolic profiling | Self-supervised deep spectral component filtering foundation model for metabolic profiling analysis. | Nature Machine Intelligence |
Foundation models for cryo-electron microscopy image processing, density map analysis, and structure refinement.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CryoFM | CryoFM: A Flow-based Foundation Model for Cryo-EM Densities | Flow matching-based foundation model learning high-quality biomolecular density map distributions. | NeurIPS 2024 / bioRxiv |
| Cryo-IEF | A comprehensive foundation model for cryo-EM image processing | Comprehensive cryo-EM image processing foundation model pretrained via contrastive learning on 65M particle images. | Nature Methods |
| CryoLVM | CryoLVM: Self-supervised Learning from Cryo-EM Density Maps | Self-supervised cryo-EM density map foundation model using JEPA architecture. | arXiv |
| CryoNet.Refine | CryoNet.Refine: A One-step Diffusion Model for Rapid Refinement of Structural Models with Cryo-EM Density Map Restraints | One-step diffusion model for rapid structural model refinement with cryo-EM density map constraints. | arXiv |
| CryoDRGN-AI | CryoDRGN-AI: neural ab initio reconstruction for cryo-EM | Neural ab initio reconstruction method for heterogeneous cryo-EM data. | Nature Methods |
Foundation models for metagenomic sequencing, microbiome analysis, and pathogen monitoring.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| METAGENE-1 | Metagenomic Foundation Model for Pandemic Monitoring | 7B-parameter autoregressive transformer trained on >1.5 trillion bases of metagenomic DNA/RNA for pathogen monitoring. | arXiv |
| MGM | MGM as a large-scale pretrained foundation model for microbiome analyses in diverse contexts | Large-scale microbiome foundation model trained on >263,000 microbiome samples for diverse contexts. | Advanced Science |
| BiomeGPT | BiomeGPT: A foundation model for the human gut microbiome | Transformer-based human gut microbiome foundation model trained on >13,300 metagenomic samples. | bioRxiv |
| GenomeOcean | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies | Efficient 4B-parameter genome foundation model trained on large-scale metagenomic assemblies (>600 Gbp). | bioRxiv |
| MicroGenomer | MicroGenomer: A Foundation Model for Transferable Microbial Genome Representations | Transferable microbial genome representation model trained on >234.5B base pairs for multi-scale analysis. | bioRxiv (BGI Research) |
| Generanno | Generanno: A Genomic Foundation Model for Metagenomic Annotation | Genomic foundation model for metagenomic annotation trained on 715B prokaryotic base pairs. | bioRxiv |
| FGBERT | FGBERT: Function-Driven Pre-trained Gene Language Model for Metagenomics | Function-driven pretrained gene language model using protein-level context-aware tokenizer for metagenomics. | arXiv |
| MetagenBERT | MetagenBERT: a Transformer Architecture using Foundational DNA Read Embedding Models for novel Metagenome Representation | Transformer framework using DNABERT-2/DNABERT-S for metagenome representation from raw DNA reads. | arXiv |
| Darwin-7B | Darwin-7B: A Multi-Omic Foundation Model for the Human Gut Microbiome via Sparsified Quality-Aware Tokenization | 7B-parameter multi-omic foundation model for the human gut microbiome, trained with sparsified quality-aware tokenization. | ICLR 2026 Workshop |
| ViraLM | ViraLM: virus discovery through genome foundation model | Virus genome foundation model for virus discovery from metagenomic sequences. | Bioinformatics |
Foundation models for phylogenetic tree inference and evolutionary genomic modeling.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Phyla | Evolutionary Reasoning Does Not Arise in Standard Usage of Protein Language Models | Phylogenetic inference foundation model (Phyla, originally proposed in v1 of the same preprint) using hybrid state-space transformer with tree loss function. | bioRxiv |
| PhyloGPN | A Phylogenetic Approach to Genomic Language Modeling | Phylogenetic tree-based genomic language model using multi-species whole-genome alignments and evolutionary models. | Lecture Notes in Computer Science |
Foundation models for integrating DNA, RNA, protein, and other multi-omic data modalities.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| OmniBioTE | Large-Scale Multi-omic Biosequence Transformers for Modeling Protein-Nucleic Acid Interactions | Large-scale multi-omic biosequence transformer trained on >250B tokens of protein and nucleic acid sequences. | PLOS ONE |
| Omni-DNA | Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning | Unified genomic foundation model supporting DNA/RNA/protein cross-modal multi-task learning from Microsoft. | NeurIPS 2025 (Microsoft) |
| OmniNA | OmniNA: A foundation model for nucleotide sequences | Nucleotide sequence foundation model pretrained on >91.7M sequences (>1 trillion bases) for cross-species understanding. | bioRxiv |
| spEMO | Leveraging multi-modal foundation models for analysing spatial multi-omic and histopathology data | Multi-modal foundation model framework integrating spatial multi-omics with histopathology image data. | Nature Biomedical Engineering |
| scMamba | scMamba: A Scalable Foundation Model for Single-Cell Multi-Omics Integration Beyond Highly Variable Feature Selection | Scalable Mamba-based foundation model for single-cell multi-omics integration without feature selection. | arXiv |
Foundation models for DNA methylation, chromatin modifications, and epigenetic regulation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CpGPT | CpGPT: a Foundation Model for DNA Methylation | DNA methylation foundation model predicting CpG site methylation states for aging and disease research. | bioRxiv |
| MethylGPT | MethylGPT: A Foundation Model for the DNA Methylome | DNA methylome foundation model pretrained on large-scale methylation data for epigenetic age prediction and cancer classification. | bioRxiv |
| scDNAm-GPT | scDNAm-GPT: A Foundation Model for Single-Cell DNA Methylation Analysis | Single-cell DNA methylation analysis foundation model for resolving epigenetic heterogeneity at single-cell resolution. | bioRxiv |
Foundation models for mass spectrometry-based proteomics, metabolomics, and compound identification.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| DIA-BERT | DIA-BERT: pre-trained end-to-end transformer models for enhanced DIA proteomics data analysis | Transformer-based foundation model for data-independent acquisition proteomics, improving peptide identification and quantification. | Nature Communications |
| DreaMS | DreaMS: Deep Representations Empowering the Annotation of Mass Spectra | Deep representation learning foundation model for mass spectra annotation and metabolite identification. | Nature Biotechnology |
| LSM-MS2 | LSM-MS2: Large-Scale Mass Spectrometry Foundation Model | Large-scale tandem mass spectrometry foundation model pretrained on millions of MS2 spectra for compound identification. | ChemRxiv |
| OmniNovo | OmniNovo: A Universal Foundation Model for De Novo Peptide Sequencing | Universal foundation model for de novo peptide sequencing directly from mass spectrometry data. | arXiv |
| MS-FM | Foundation model for mass spectrometry proteomics | Unified mass spectrometry proteomics foundation model pretrained on de novo sequencing data. | arXiv |
| InstaNovo | InstaNovo: diffusion-powered de novo peptide sequencing | Diffusion-powered model for de novo peptide sequencing from mass spectrometry data. | Nature Machine Intelligence |
Foundation models for neural activity prediction, brain imaging, and computational neuroscience.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| VFAM | Foundation model of neural activity predicts response to new stimulus types | Neural activity foundation model trained on large-scale mouse visual cortex data, predicting responses to novel stimulus types. | Nature |
Foundation models for designing synthetic regulatory elements and engineering biological systems.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| DNA-Diffusion | Designing synthetic regulatory elements using DNA-Diffusion | Generative diffusion model for designing synthetic DNA regulatory elements. | Nature Genetics |
Foundation models for molecular property prediction, representation learning, and chemical language modeling on SMILES and molecular graphs.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MoLFormer | Large-Scale Chemical Language Representations Capture Molecular Structure and Properties | Transformer-based chemical language model pretrained on 1.1B SMILES with linear attention and rotary embeddings for molecular property prediction. | Nature Machine Intelligence |
| GP-MoLFormer | GP-MoLFormer: A Foundation Model For Molecular Generation | Transformer-based generative foundation model with 46.8M parameters trained on 1.1B SMILES for molecular generation tasks. | Digital Discovery |
| ChemBERTa | ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction | RoBERTa-based chemical language model pretrained on 10M PubChem SMILES for molecular property prediction. | NeurIPS ML4Molecules Workshop |
| ChemBERTa-2 | ChemBERTa-2: Towards Chemical Foundation Models | Evolved ChemBERTa pretrained on 77M PubChem SMILES with optimized pretraining strategies including multi-task regression. | arXiv |
| ChemFM | ChemFM as a Scaling Law Guided Foundation Model Pre-trained on Informative Chemicals | 3B-parameter chemical foundation model trained on 178M UniChem molecules using self-supervised causal language modeling guided by scaling laws. | Communications Chemistry |
| ChemDFM | Developing ChemDFM as a Large Language Foundation Model for Chemistry | LLaMA-13B-based chemistry LLM trained on 34B tokens of chemical literature and fine-tuned with 2.7M instruction pairs. | Cell Reports Physical Science |
| Uni-Mol | Uni-Mol: A Universal 3D Molecular Representation Learning Framework | Universal 3D molecular representation learning framework that directly leverages molecular 3D structures for pretraining and achieves SOTA on property prediction. | ICLR |
| Uni-Mol2 | Uni-Mol2: Exploring Molecular Pretraining Model at Scale | Largest 3D molecular foundation model (1.1B parameters) with dual-track Transformer integrating atomic, graph, and 3D geometric features trained on 884M molecules. | NeurIPS |
| MolBERT | MolBERT: Molecular Representation Learning with Language Models and Domain-Relevant Auxiliary Tasks | BERT-based molecular representation model using SMILES with self-supervised auxiliary tasks for meaningful molecular embeddings. | arXiv |
| GROVER | Self-Supervised Graph Transformer on Large-Scale Molecular Data | Self-supervised graph Transformer combining GNN message passing with Transformer attention, pretrained on large-scale molecular data. | NeurIPS |
| GEM | Geometry-Enhanced Molecular Representation Learning for Property Prediction | Geometry-enhanced molecular representation learning framework that exploits 3D spatial structure information for improved property prediction. | Nature Machine Intelligence |
| Graphormer | Do Transformers Really Perform Bad for Graph Representation? | Graph Transformer framework from Microsoft that won 1st place on OGB-LSC molecular tasks, introducing spatial and edge encodings for graph structure modeling. | NeurIPS |
| MoleculeSTM | Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing | Multi-modal model jointly learning molecular structures and text descriptions for text-driven molecular retrieval and editing. | Nature Machine Intelligence |
| 3D-MoLM | Towards 3D Molecule-Text Interpretation in Language Models | Pioneering framework integrating a 3D molecular encoder with language models via a 3D molecule-text projector for LLM-based molecular understanding. | ICLR |
| MolE | MolE: A Foundation Model for Molecular Graphs Using Disentangled Attention | Molecular graph foundation model from Recursion using disentangled-attention Transformers, pretrained in two self-supervised stages on ~842M molecules. | Nature Communications |
| MiniMol | MiniMol: A Parameter-Efficient Foundation Model for Molecular Learning | Parameter-efficient molecular foundation model (only 10M parameters) pretrained on 3,300+ diverse bioactivity datasets from 6M molecules. | ICML |
| SMILES-Mamba | SMILES-Mamba: Chemical Mamba Foundation Models for Drug ADMET Prediction | Mamba-architecture chemical foundation model with two-stage training (self-supervised pretraining + supervised fine-tuning) for drug ADMET prediction. | NeurIPS 2024 Workshop |
| SMI-TED | SMI-TED: Large-Scale Foundation Model for Materials and Chemistry | Large-scale SMILES encoder-decoder foundation model from IBM, self-supervised on 91M PubChem SMILES for chemistry and materials science. | ICLR 2024 Workshop |
| KPGT | A Knowledge-Guided Pre-training Framework for Improving Molecular Representation | Knowledge-guided graph Transformer pretraining framework integrating chemical knowledge to enhance molecular representation learning. | Nature Communications |
| GIN (Pretrained) | Strategies for Pre-Training Graph Neural Networks | Pioneering work proposing GNN pretraining strategies (node-level + graph-level) with GIN pretrained on 2M molecules for property prediction. | ICLR |
| MIST | Foundation Models for Discovery and Exploration in Chemical Space | Family of large-scale molecular foundation models (Molecular Insight SMILES Transformers) trained on vast unlabeled molecules, predicting 400+ structure-property relationships. | arXiv |
| M2UMol | Multi-to-Uni Modal Knowledge Transfer Pre-training for Molecular Representation Learning | Multi-modal to uni-modal knowledge transfer pretraining framework that distills diverse molecular modality knowledge into a 2D encoder. | Nature Communications |
| Omni-Mol | Exploring Universal Convergent Space for Omni-Molecular Tasks | Unified language model enabling any-to-any modality molecular tasks in a universal convergent space. | NeurIPS |
| TamGen | TamGen: drug design with target-aware molecule generation through a chemical language model | GPT-style chemical language model for target-aware molecule generation and drug design. | Nature Communications |
| MoleculeGPT | MoleculeGPT: Instruction Following LLMs for Molecular Property Prediction | LLM fine-tuned with molecular instruction data for natural-language-driven molecular property prediction. | NeurIPS 2024 Workshop |
| SAFE-GPT | SAFE: A Molecular-Centric Foundation Model with SAFE Representation | Foundation model using Sequential Attachment-based Fragment Embedding (SAFE) molecular representation for generative chemistry. | Digital Discovery |
| DrugGPT | DrugGPT: A GPT-based Strategy for Designing Potential Ligands Targeting Specific Proteins | GPT-based drug design model that generates drug-like molecules targeting specific protein binding pockets. | bioRxiv |
| MultiPUFFIN | Multimodal domain-constrained foundation model | Multi-modal domain-constrained foundation model integrating SMILES, molecular graphs, and 3D geometry for molecular understanding. | arXiv |
| FragCLM | Foundation chemical language model for fragment-based drug discovery | Foundation chemical language model trained on the ZINC-22 fragment dataset for comprehensive fragment-based drug discovery. | arXiv |
Foundation models for chemical reaction prediction, retrosynthetic planning, and synthesis route design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Molecular Transformer | Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction | Pioneering seq2seq Transformer that frames chemical reaction prediction as SMILES translation with uncertainty calibration. | ACS Central Science |
| Chemformer | Chemformer: a pre-trained transformer for computational chemistry | BART-based pretrained Transformer for reaction prediction, retrosynthesis, and other computational chemistry tasks over molecular SMILES. | Machine Learning: Science and Technology |
| RXNFP | Mapping the space of chemical reactions using attention-based neural networks | Transformer model from IBM that learns chemical reaction fingerprints for reaction classification and reaction space mapping. | Nature Machine Intelligence |
| T5Chem | Unified Deep Learning Model for Multitask Reaction Predictions with Explanation | T5-based unified Transformer supporting multi-task chemical reaction predictions including forward synthesis, retrosynthesis, and yield prediction. | Journal of Chemical Information and Modeling |
| Llamole | Multimodal Large Language Models for Inverse Molecular Design with Retrosynthetic Planning | Multi-modal LLM integrating graph diffusion Transformer and GNN for inverse molecular design with retrosynthetic route planning. | NeurIPS |
| RetroSynFormer | Retrosynformer: Planning Multi-step Chemical Synthesis Routes via a Decision Transformer | Decision Transformer for multi-step retrosynthetic planning that models retrosynthesis as a sequence prediction problem. | Digital Discovery |
| SynLlama | SynLlama: Generating Synthesizable Molecules and Their Analogs with Large Language Models | LLM fine-tuned from Meta Llama 3 for generating synthesizable molecules along with complete synthesis routes. | ACS Central Science |
| SynFormer | Generative Artificial Intelligence for Navigating Synthesizable Chemical Space | Generative model framework for exploring synthesizable chemical space, ensuring generated molecules are synthetically accessible. | PNAS |
| ReactionT5 | ReactionT5: A Pre-trained Transformer Model for Accurate Chemical Reaction Prediction with Limited Data | T5-based pretrained Transformer for chemical reaction prediction, pretrained on the Open Reaction Database and excelling with limited data. | Journal of Cheminformatics |
| DeepRetro | DeepRetro Discovers Retrosynthetic Pathways Through Iterative Large Language Model Reasoning | Advanced retrosynthesis framework combining LLM reasoning, reaction templates, and expert feedback for iterative pathway discovery. | Scientific Reports |
| RXNGraphormer | A unified pre-trained deep learning framework for cross-task reaction performance prediction | Unified pretrained reaction graph Transformer integrating GNN and Transformer to learn bond formation/breaking mechanisms across tasks. | Nature Machine Intelligence |
| RSGPT | RSGPT: a generative transformer for retrosynthesis planning pre-trained on ten billion datapoints | Generative Transformer for retrosynthesis planning pretrained on 10 billion datapoints for large-scale synthetic route prediction. | Nature Communications |
| Chem-R | Chem-R: Learning to Reason as a Chemist | Chemical reasoning model that emulates chemists' deep thinking processes through a three-phase training framework. | NeurIPS |
Foundation models for molecular docking, binding affinity prediction, and protein-ligand complex structure prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Pearl | Pearl: A Foundation Model for Placing Every Atom in the Right Location | Protein-ligand structure prediction foundation model from Genesis Molecular AI using large-scale synthetic data and SO(3)-equivariant architecture, surpassing AlphaFold 3. | NeurIPS |
| DiffDock | DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking | Diffusion-based molecular docking method that models protein-ligand docking as a generative problem on SE(3) without requiring prior binding site knowledge. | ICLR |
| DiffDock-L | Fine-Tuning DiffDock-L for Allosteric Kinase Docking | Large-scale DiffDock variant fine-tuned for allosteric kinase binding site docking. | J. Chem. Inf. Model. |
| NeuralPLexer | State-specific Protein-Ligand Complex Structure Prediction with a Multiscale Deep Generative Model | Multi-scale deep generative model that predicts protein-ligand complex 3D structures directly from protein sequence and ligand graph, including conformational changes. | Nature Machine Intelligence |
| Uni-Mol Docking V2 | Uni-Mol Docking V2: Towards Realistic and Accurate Binding Pose Prediction | Molecular docking method in the Uni-Mol series using pretrained molecular and pocket encoders to predict protein-ligand binding poses with >77% success rate. | Lecture Notes in Computer Science |
| Umol | Structure Prediction of Protein-Ligand Complexes from Sequence Information with Umol | AI system predicting full-flexibility, all-atom protein-ligand complex structures solely from amino acid sequence and SMILES. | Nature Communications |
| LigUnity | A Foundation Model for Protein-Ligand Affinity Prediction Through Unified Representation | Unified representation learning foundation model for protein-ligand affinity prediction supporting both virtual screening and lead optimization. | bioRxiv preprint |
| PhysDock | PhysDock: A Physics-Guided All-Atom Diffusion Model for Protein-Ligand Complex Prediction | Physics-guided all-atom diffusion model for protein-ligand complex prediction integrating detailed atomic-level flexibility modeling. | bioRxiv |
| Boltz-2 | Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction | Advanced model for accurate and efficient protein-ligand binding affinity prediction building on AlphaFold3 and Boltz-1 architectures. | bioRxiv |
Equivariant and invariant neural network architectures for learning 3D molecular representations, energy prediction, and force fields.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| SchNet | SchNet: A Continuous-filter Convolutional Neural Network for Modeling Quantum Interactions | Continuous-filter convolutional neural network learning rotationally invariant representations of quantum interactions for molecular energy and force prediction. | NeurIPS 2017 |
| ViSNet | ViSNet: An Equivariant Geometry-Enhanced Graph Neural Network with Vector-Scalar Interactive Message Passing | Equivariant geometry-enhanced GNN with vector-scalar interactive message passing that avoids expensive higher-order tensor operations via runtime geometric computation. | Nature Communications |
| EPT | An equivariant pretrained transformer for unified 3D molecular representation learning | E(3)-equivariant all-atom pretrained Transformer for unified 3D molecular representation learning across diverse scientific domains. | Nature Communications |
Generative models for de novo molecular design, 3D conformation generation, and structure-based molecule generation using diffusion, VAEs, autoregressive, and flow-based approaches.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MolGPT | MolGPT: Molecular Generation Using a Transformer-Decoder Model | GPT-based molecular generation model using a Transformer decoder to autoregressively generate SMILES satisfying specific property constraints. | J. Chem. Inf. Model. |
| cMolGPT | cMolGPT: A Conditional Generative Pre-Trained Transformer for Target-Specific de novo Molecular Generation | Conditional molecular GPT extending MolGPT with target-specific controls for de novo molecular generation. | Molecules |
| GenMol | GenMol: A Drug Discovery Generalist with Discrete Diffusion | General-purpose molecular generation model from NVIDIA using masked discrete diffusion over SAFE representations for multi-stage drug discovery. | ICLR |
| NExT-Mol | NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generation | Foundation model integrating 1D SELFIES language modeling with 3D diffusion for 3D molecule generation. | ICLR 2025 |
| DiTMC | Sampling 3D Molecular Conformers with Diffusion Transformers | Diffusion Transformer framework for sampling accurate 3D molecular conformers integrating discrete molecular graphs with continuous coordinates. | NeurIPS |
| SynCoGen | Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling | Framework for synthesizable 3D molecule generation that jointly models molecular building blocks, chemical reactions, and atomic coordinates. | ICLR |
Foundation models for interpreting and predicting molecular spectra including NMR, IR, Raman, and mass spectrometry.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MolSpectLLM | MolSpectLLM: A Large Language Model for Molecular Spectroscopy Interpretation | Large language model for molecular spectroscopy interpretation, linking NMR, IR, and MS spectral data to molecular structures for spectrum-to-structure reasoning. | arXiv |
Chemical language models applied to food-related molecular property prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FART | A chemical language model for molecular taste prediction | Chemical language model predicting molecular taste properties from SMILES representations. | npj Science of Food |
Foundation models for electrochemical applications including battery electrolyte design.
| Model | Paper Title | Description | Link |
|---|
Machine-learned interatomic potentials for molecular dynamics simulations.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MACE | MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields | Higher-order equivariant message passing framework using atomic cluster expansion for accurate and efficient force field computation. | NeurIPS 2022 |
| MACE-MP-0 | A Foundation Model for Atomistic Materials Chemistry | Pre-trained universal force field covering 89 elements, trained on Materials Project data for general materials chemistry simulation. | J. Chem. Phys. |
| CHGNet | CHGNet as a Pretrained Universal Neural Network Potential for Charge-Informed Atomistic Modelling | Pre-trained universal graph neural network potential with charge information, trained on Materials Project DFT data. | Nature Machine Intelligence |
| M3GNet | A Universal Graph Deep Learning Interatomic Potential for the Periodic Table | Universal graph deep learning interatomic potential trained on Materials Project relaxation data, covering all periodic table elements. | Nature Computational Science |
| SevenNet | SevenNet: Scalable Graph Neural Network Interatomic Potential | Scalable GNN interatomic potential based on NequIP architecture with LAMMPS parallel MD support. | J. Chem. Theory Comput. |
| Orb | Orb: A Fast, Scalable Neural Network Potential | Fast and scalable neural network potential by Orbital Materials, 3–6× faster than existing universal potentials while maintaining SOTA accuracy. | arXiv |
| NequIP | E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials | E(3)-equivariant GNN using equivariant convolutions instead of invariant descriptors, achieving high accuracy with minimal training data. | Nature Communications |
| Allegro | Learning local equivariant representations for large-scale atomistic dynamics | Highly scalable E(3)-equivariant architecture using local equivariant representations to support large-scale molecular dynamics. | Nature Communications |
| Allegro-FM | Allegro-FM: Toward an Equivariant Foundation Model for Exascale Molecular Dynamics Simulations | Equivariant foundation model targeting exascale molecular dynamics simulations based on the Allegro architecture. | J. Phys. Chem. Lett. |
| DPA-2 | DPA-2: a large atomic model as a multi-task learner | Large-scale Deep Potential multi-task atomic model pre-trained across diverse chemical and materials systems with fine-tuning support. | npj Computational Materials |
| ANI-1 | ANI-1: an extensible neural network potential with DFT accuracy at force field computational cost | Pioneering extensible neural network potential achieving DFT accuracy at force-field computational cost for H/C/N/O organic molecules. | Chemical Science |
| ANI-2x | Extending the Applicability of the ANI Deep Learning Molecular Potential to Sulfur and Halogens | Extension of ANI to sulfur and halogens (F/Cl), broadening coverage to a wider organic molecular space. | J. Chem. Theory Comput. |
| AIMNet2 | AIMNet2: A Neural Network Potential to Meet Your Neutral, Charged, Organic, and Elemental-Organic Needs | Highly transferable neural network potential supporting neutral and charged organic molecules across 14 elements. | Chemical Science |
| GRACE | Graph Atomic Cluster Expansion | Universal MLIP framework based on graph atomic cluster expansion, covering 97 elements. | npj Comp. Mater. |
| Orb-v3 | Orb-v3: Atomistic Simulation at Scale | Major upgrade of Orb with improved accuracy and efficiency for large-scale atomistic simulation. | arXiv |
| PET-MAD | Lightweight universal interatomic potential for advanced materials | Lightweight universal interatomic potential covering the full periodic table for advanced materials simulations. | Nature Communications |
| Grappa | Machine-learned molecular mechanics via E(3)-equivariant neural networks | E(3)-equivariant neural network approach to machine-learned molecular mechanics force fields. | Chemical Science |
Predicting physical, electronic, and structural properties of crystalline and molecular materials.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CGCNN | Crystal Graph Convolutional Neural Networks for an Accurate and Interpretable Prediction of Material Properties | Pioneering crystal graph convolutional neural network for direct, interpretable prediction of material properties from crystal structures. | Physical Review Letters |
| MEGNet | Graph Networks as a Universal Machine Learning Framework for Molecules and Crystals | Universal materials graph network supporting property prediction for both molecules and crystals with global state features. | Chemistry of Materials |
| ALIGNN | Atomistic Line Graph Neural Network for Improved Materials Property Predictions | NIST atomistic line graph neural network that explicitly models bond angles, outperforming CGCNN and MEGNet. | npj Computational Materials |
| MultiMat | Multimodal Foundation Models for Material Property Prediction and Discovery | Multimodal foundation model integrating crystal structure, density of states, charge density, and text for comprehensive materials property prediction. | Newton |
| SMI-TED | SMI-TED: Large-Scale Foundation Model for Materials and Chemistry | IBM large-scale SMILES encoder-decoder model pre-trained on 91M PubChem SMILES for materials and chemistry applications. | ICLR 2024 Workshop |
| DARWIN 1.5 | DARWIN 1.5: Large Language Models as Materials Science Adapted Learners | Open-source materials science LLM that predicts material properties and facilitates discovery from natural language input. | arXiv |
| MatBERT | Quantifying the Advantage of Domain-Specific Pre-training on Named Entity Recognition Tasks in Materials Science | BERT model pre-trained on materials science literature (LBNL), outperforming general models on materials NLP tasks. | Patterns |
| MatSciBERT | MatSciBERT: A Materials Domain Language Model for Text Mining and Information Extraction | Domain-specific BERT trained on materials science literature for enhanced text mining and information extraction. | npj Computational Materials |
| MOFTransformer | A Multi-modal Pre-training Transformer for Universal Transfer Learning in Metal-Organic Frameworks | Multi-modal pre-trained Transformer for MOF property prediction, trained on 1M hypothetical MOFs with atomic graph and energy grid embeddings. | Nature Machine Intelligence |
| CrystalFormer | Space Group Informed Transformer for Crystalline Materials Generation | Autoregressive Transformer guided by space group symmetry and Wyckoff positions for crystalline materials generation. | Science Bulletin |
| MatInFormer | Materials Informatics Transformer: A Language Model for Interpretable Materials Properties Prediction | Materials informatics Transformer leveraging LLM techniques for interpretable materials property prediction. | arXiv |
| KPGT | A Knowledge-Guided Pre-training Framework for Improving Molecular Representation | Knowledge-guided graph Transformer pre-training framework using chemical knowledge to enhance molecular representation learning. | Nature Communications |
| Matformer | Periodic Graph Transformers for Crystal Material Property Prediction | Periodic graph Transformer with periodicity-aware multi-graph attention for crystal material property prediction. | NeurIPS 2022 |
| PotNet | Complete and Efficient Graph Transformers for Crystal Material Property Prediction | Complete and efficient crystal graph Transformer achieving full graph representation via interatomic potential information. | ICLR |
| LLM-Prop | LLM-Prop: Predicting Physical And Electronic Properties of Crystalline Solids From Their Text Descriptions | Uses large language models to predict physical and electronic properties of crystals from text descriptions. | arXiv |
| EScAIP | EScAIP: Efficiently Scaled Attention Interatomic Potential | Efficiently scaled attention-based interatomic potential achieving high accuracy and scalability for materials property prediction. | ICLR |
| AlloyGPT | End-to-end prediction and design of additively manufacturable alloys | Autoregressive language model for end-to-end alloy design and property prediction. | npj Computational Materials |
| aLLoyM | aLLoyM: a large language model for alloy phase diagram prediction | Large language model for predicting alloy phase diagrams. | npj Computational Materials |
| MaskTerial | MaskTerial: a foundation model for automated 2D material flake detection | Foundation model for automated detection of 2D material flakes. | Digital Discovery |
| CLOUD | CLOUD: A Scalable and Physics-Informed Foundation Model for Crystal Representation Learning | Scalable physics-informed crystal representation foundation model trained on 6M+ crystal structures. | Nature Communications |
| LLaMat | A family of large language models for materials research with insights into model adaptability in continued pretraining | Family of large language models adapted for materials science tasks. | Nature Machine Intelligence |
Large-scale pre-trained models spanning molecules, materials, and catalysts across the periodic table.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| UMA | UMA: A Family of Universal Models for Atoms | Meta FAIR universal atomic model family trained on 500M+ 3D atomic structures spanning molecules, materials, and catalysts. | NeurIPS |
| Zatom-1 | Zatom-1: A Multimodal Flow Foundation Model for 3D Molecules and Materials | Open-source multimodal flow foundation model unifying generation and prediction for 3D molecules and materials. | arXiv |
| MIST | Foundation Models for Discovery and Exploration in Chemical Space | Large-scale molecular foundation model family (Molecular Insight SMILES Transformers) predicting 400+ structure-property relationships. | arXiv |
| MatterSim | MatterSim: A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures | Microsoft deep learning atomic model covering all elements at 0–5000 K and 0–1000 GPa. | arXiv |
| GNoME | Scaling Deep Learning for Materials Discovery | Google DeepMind GNN materials explorer discovering 2.2M new stable inorganic crystal structures. | Nature |
| JMP | From Molecules to Materials: Pre-training Large Generalizable Models for Atomic Property Prediction | Meta FAIR joint multi-domain pre-training on ~120M atomic systems spanning molecules and materials. | ICLR |
| ATOMICA | Learning Universal Representations of Intermolecular Interactions with ATOMICA | Geometric deep learning model learning universal atomic-level representations of intermolecular interactions. | bioRxiv |
| eSEN | Efficient Scalable Equivariant Networks | Scalable equivariant architecture forming the backbone of UMA, achieving SOTA on molecular and materials benchmarks. | arXiv |
Generative models for discovering and designing novel crystal structures.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| CDVAE | Crystal Diffusion Variational Autoencoder for Periodic Material Generation | Crystal diffusion VAE combining diffusion processes with VAE for end-to-end stable periodic crystal structure generation. | ICLR |
| DiffCSP | Crystal Structure Prediction by Joint Equivariant Diffusion | Joint equivariant diffusion model simultaneously diffusing atom coordinates and lattice parameters for crystal structure prediction. | NeurIPS 2023 |
| SyMat | Towards Symmetry-Aware Generation of Periodic Materials | Symmetry-aware periodic material generation model explicitly leveraging space group symmetry constraints. | NeurIPS 2023 |
| MatterGen | MatterGen: A Generative Model for Inorganic Materials Design | Microsoft diffusion model for inverse design of inorganic crystals conditioned on chemistry, symmetry, and property constraints. | Nature |
| Crystal-GFN | Crystal-GFN: Sampling Crystals with Desirable Properties and Constraints | GFlowNet-based crystal sampling framework that efficiently explores crystal space under property and composition constraints. | arXiv |
| FlowMM | FlowMM: Generating Materials with Riemannian Flow Matching | Riemannian flow matching on crystal manifolds for geometry-aware materials structure generation. | ICML |
| FlowLLM | FlowLLM: Flow Matching for Material Generation with Large Language Models as Base Distributions | Combines LLM base distributions with flow matching, leveraging chemical priors for improved crystal generation. | NeurIPS 2024 |
| CrystalFlow | CrystalFlow: A Flow-Based Generative Model for Crystalline Materials | Flow-based generative model achieving high-fidelity crystal structure generation via normalizing flows. | Nature Communications |
| WyckoffDiff | WyckoffDiff: Diffusion in the Wyckoff Space for Crystal Structure Generation | Diffusion model operating in Wyckoff position space with symmetry-aware representations for improved structural validity. | ICML |
| MatterGPT | MatterGPT: A Generative Transformer for Multi-Property Inverse Design of Solid-State Materials | Autoregressive Transformer supporting multi-property conditioned inverse design of solid-state materials. | arXiv |
| CrystaLLM | CrystaLLM: Large Language Model for Crystallography | LLM that generates crystal structures directly from CIF text without explicit geometric encoding. | Nature Communications |
| UniMat | Scalable Diffusion for Materials Generation | Scalable diffusion model for crystal materials generation with a unified representation across varying crystal sizes. | ICLR |
| DAO-G / DAO-P | Siamese Foundation Models for Crystal Structure Prediction | Siamese pre-training framework: DAO-G for crystal generation and DAO-P for property prediction. | arXiv |
| MOFGPT | Transformer-based generative model for de novo MOF design | Transformer generative model for de novo design of metal-organic frameworks. | arXiv |
| Matra-Genoa | Autoregressive generative material Transformer | Autoregressive Transformer for generative materials design. | npj Computational Materials |
Models for catalytic reaction prediction, adsorption energies, and surface chemistry.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AdsorbML | AdsorbML: A Leap in Efficiency for Adsorption Energy Calculations using Generalizable Machine Learning Potentials | Generalizable ML potentials for efficient adsorption energy calculation, accelerating catalyst screening with OC20 pre-trained models. | npj Computational Materials |
| CatBERTa | CatBERTa: A RoBERTa-based Catalyst Property Prediction Model | RoBERTa-based model predicting catalyst adsorption energies and activities from textual descriptions. | arXiv |
| eSCN | Reducing SO(3) Convolutions to SO(2) for Efficient Equivariant GNNs | Efficient equivariant spherical channel network reducing SO(3) to SO(2) convolutions for major computational speedup. | ICML |
| SCN | Spherical Channels for Modeling Atomic Interactions | Spherical channel network using spherical harmonics for atomic interaction modeling, excelling on OC20 catalyst tasks. | NeurIPS 2022 |
| EquiformerV2 | EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations | Improved equivariant Transformer supporting higher-degree representations, achieving SOTA on OC20/OC22 benchmarks. | ICLR |
| eqV2 | Improved EquiformerV2 for OC20/OC22 | Improved EquiformerV2 variant for general atomic property prediction on Open Catalyst datasets. | arXiv |
| CatDRX | Reaction-conditioned generative model for catalyst design and optimization with CatDRX | Reaction-conditioned generative model for designing catalysts tailored to specific reactions. | Communications Chemistry |
Deep learning models for predicting DFT Hamiltonians and electronic properties.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| DeepH | Deep-learning density functional theory Hamiltonian for efficient ab initio electronic-structure calculation | Deep learning model that directly predicts DFT Hamiltonian matrices to accelerate ab initio electronic structure calculations. | Nature Computational Science |
| DeepH-E3 | DeepH-E3: E(3)-Equivariant Deep Learning for Efficient ab initio Electronic Structure | E(3)-equivariant version of DeepH for more accurate and efficient Hamiltonian matrix element prediction. | Nature Communications |
| HamGNN | HamGNN: Graph Neural Networks for Predicting Hamiltonian Matrix | Graph neural network for Hamiltonian matrix prediction via equivariant message passing at DFT-level accuracy. | arXiv |
| NextHAM | NextHAM: Next-Generation Hamiltonian Prediction with Equivariant Graph Neural Networks | Next-generation equivariant GNN for electronic structure prediction of larger-scale materials systems. | arXiv |
| MACE-H | Equivariant electronic Hamiltonian prediction with many-body message passing | MACE-based equivariant GNN for predicting electronic Hamiltonians. | npj Computational Materials |
Language models and graph networks for polymer informatics and design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| polyBERT | polyBERT: a chemical language model to enable fully machine-driven ultrafast polymer informatics | BERT-based chemical language model trained on polymer SMILES for ultrafast polymer property prediction. | Nature Communications |
| polyGNN | polyGNN: Multitask Graph Neural Networks for Polymer Informatics | Multitask graph neural network for simultaneous polymer property prediction and structure-property learning. | Chemistry of Materials |
| polyBART | polyBART: A Generative Transformer for Polymer Design | BART-based generative Transformer for conditional polymer generation and property-guided inverse design. | arXiv |
| POLYT5 | POLYT5: an encoder-decoder foundation chemical language model for generative polymer design | T5 encoder-decoder foundation chemical language model for polymer design. | npj Artificial Intelligence |
Pre-trained models for battery research, electrode materials, and energy storage.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| BatteryBERT | BatteryBERT: A Pretrained Language Model for Battery Research | BERT fine-tuned on battery literature for text mining and information extraction in battery research. | J. Chem. Inf. Model. |
| BatteryFormer | BatteryFormer: Graph Transformer for Battery Material Property Prediction | Graph Transformer predicting battery electrode capacity, voltage, and cycle life properties. | arXiv |
Foundational graph neural network architectures underlying many materials science models.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| SchNet | SchNet: A Continuous-Filter Convolutional Neural Network for Modeling Quantum Interactions | Pioneering continuous-filter convolutional network encoding interatomic distances as continuous representations for quantum interactions. | NeurIPS |
Foundation models for metamaterial structure-property relationships.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| MetaFO | Toward a robust and generalizable metamaterial foundation model | Bayesian Transformer metamaterial foundation model for zero-shot structure-property prediction. | npj Computational Materials |
Models for predicting superconducting properties and critical temperatures.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| BEE-NET | Developing a complete AI-accelerated workflow for superconductor discovery | Equivariant GNN predicting Eliashberg spectral functions and critical temperatures for superconductor discovery. | npj Computational Materials |
| DeeperBand | A deep learning approach to search for superconductors from electronic bands | Symmetry-aware 3D Vision Transformer predicting superconductivity from electronic band structures. | IOPscience |
Foundation models and deep learning architectures for jet tagging, particle tracking, and collider event analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FM4NPP | A Scaling Foundation Model for Nuclear and Particle Physics | Large-scale self-supervised foundation model for sparse detector data, trained on 11M+ collision events achieving SOTA on sPHENIX experiments. | ICLR |
| OmniLearn | OmniLearn: A Method to Simultaneously Facilitate All Jet Physics Tasks | Multi-task jet physics foundation model learning universal representations via multi-class classification pre-training. | arXiv |
| OmniLearned | Foundation Model Framework for All Tasks Involving Jet Physics | Upgraded OmniLearn framework trained on 1B+ jet events with Transformer architecture for jet classification, regression, and generation. | Physical Review D |
| OmniJet-α | OmniJet-α: The first cross-task foundation model for particle physics | First cross-task particle physics foundation model supporting both jet generation and jet tagging. | Machine Learning: Science and Technology |
| Bumblebee | Bumblebee: Foundation Model for Particle Physics Discovery | BERT-inspired particle physics foundation model embedding four-momentum vectors without positional encoding to capture generative and reconstruction-level information. | NeurIPS 2024 Workshop |
| EveNet | EveNet: A Foundation Model for Particle Collision Data Analysis | Event-level collision data foundation model pre-trained on 500M simulated events with hybrid self-supervised learning for multi-task analysis. | arXiv |
| HEP-JEPA | HEP-JEPA: A foundation model for collider physics using joint embedding predictive architecture | Collider physics foundation model using joint embedding predictive architecture (JEPA) for self-supervised jet tagging. | arXiv |
| JetCLR | Symmetries, Safety, and Self-Supervision | Contrastive self-supervised jet representation learning framework using permutation-invariant Transformer encoder with symmetry augmentation. | SciPost Phys. |
| CaloFM | Foundation Model for Calorimetry via MoE | Mixture-of-experts foundation model for calorimeter simulation. | arXiv |
| JetFormer | Scalable Transformer for Jet Tagging | Scalable Transformer architecture for jet tagging. | arXiv |
| PanopTag | PanopTag: Simultaneously Tagging All Jets in a Particle Collision Event | First method to simultaneously tag all jets in a collision event using encoder-decoder Transformer with event-level context. | arXiv |
| TrackingBERT | A Language Model for Particle Tracking | BERT-based foundation model for particle track reconstruction by tokenizing detector data for LHC tracking. | arXiv |
Neural operators and foundation models for solving partial differential equations and fluid simulations.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FNO | Fourier Neural Operator for Parametric Partial Differential Equations | Pioneering neural operator learning function mappings in Fourier space, resolution-independent and efficient for parametric PDEs. | arXiv |
| DeepONet | Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators | Deep operator network with branch/trunk architecture learning continuous nonlinear operators for PDE solving. | Nature Machine Intelligence |
| FourCastNet | FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators | NVIDIA high-resolution global weather model based on adaptive Fourier neural operators at 0.25° resolution. | arXiv |
| Poseidon | Poseidon: Efficient Foundation Models for PDEs | Efficient PDE foundation model using multi-scale operator Transformer with temporal conditional layer normalization. | NeurIPS 2024 |
| MPP | Multiple Physics Pretraining for Physical Surrogate Models | Task-agnostic Transformer pre-trained autoregressively on multiple spatiotemporal physical systems for enhanced generalization. | NeurIPS |
| ICON / ICON-LM | In-context operator learning with data prompts for differential equation problems | In-context operator learning network solving multiple PDE families via data prompts without retraining. | PNAS 2023 |
| VICON | VICON: Vision In-Context Operator Networks for Multi-Physics Fluid Dynamics | Vision in-context operator network applying Vision Transformer to multi-physics fluid dynamics prediction. | arXiv |
| DPOT | DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training | Autoregressive denoising operator Transformer with Fourier attention for large-scale PDE pre-training. | ICML |
| PROSE-PDE | Towards a Foundation Model for Partial Differential Equations: Multi-Operator Learning and Extrapolation | Multimodal PDE foundation model supporting multi-operator learning, extrapolation, and equation identification. | Phys. Rev. E |
| PROSE-FD | PROSE-FD: A Multimodal PDE Foundation Model for Learning Fluid Dynamics | Multimodal zero-shot PDE foundation model for shallow water and Navier-Stokes equations across varied geometries. | arXiv |
| OmniArch | OmniArch: Building Foundation Model for Scientific Computing | Multi-scale multi-physics scientific computing foundation model with Fourier encoder-decoder supporting 1D/2D/3D PDE simulation. | ICML |
| Unisolver | Unisolver: PDE-Conditional Transformers Are Universal PDE Solvers | Universal PDE solver using PDE-conditioned Transformer pre-trained with equation, coefficient, and boundary condition information. | NeurIPS |
| PINO | Physics-Informed Neural Operator for Learning Partial Differential Equations | Physics-informed neural operator combining data-driven and physical constraint losses to learn PDE solution operators with minimal labeled data. | ACM / IMS Journal of Data Science |
| PI-MFM | PI-MFM: Physics-informed multimodal foundation model for solving partial differential equations | Physics-informed multimodal foundation model embedding physical priors to reduce data dependency for PDE solving. | arXiv |
| Walrus | Walrus: A Cross-Domain Foundation Model for Continuum Dynamics | 1.3B-parameter cross-domain continuum dynamics foundation model pre-trained on 19 physical systems covering fluids and solids. | arXiv |
| DISCO | DISCO: Learning to DISCover an Evolution Operator for Multi-Physics-Agnostic Prediction | Multi-physics-agnostic evolution operator discovery method for efficient PDE solving and generalization. | arXiv |
| LFNO | Latent Fourier Neural Operator | Latent-space Fourier neural operator performing transforms in low-dimensional space for improved efficiency. | arXiv |
| RNO | Recurrent Neural Operator | Recurrent neural operator combining recurrent structure with operator learning for temporal PDE dynamics. | arXiv |
| MINO | Masked Implicit Neural Operator | Masked implicit neural operator using masking strategies to enhance generalization in operator learning. | arXiv |
| TNO | Transolver / Transformer Neural Operator | Transformer-based neural operator for PDEs on complex geometries and irregular grids. | ICML 2024 |
| HyPINO | Hybrid Physics-Informed Neural Operator | Hybrid physics-informed neural operator combining physical constraints with data-driven learning for enhanced PDE accuracy. | arXiv |
| PI-Latent-NO | Physics-Informed Latent Neural Operator | Physics-informed latent-space neural operator integrating equation constraints in latent space for PDE solving. | arXiv |
| Transolver | Transolver: A Fast Transformer Solver for PDEs on General Geometries | Physics-Attention Transformer PDE solver supporting arbitrary geometries, achieving multi-benchmark SOTA. | ICML |
| SFNO | Spherical Fourier Neural Operators: Learning Stable Dynamics on the Sphere | Spherical Fourier neural operator serving as the backbone of FourCastNet V2 for global weather prediction. | ICML |
| WIND | WIND: Weather Inverse Diffusion for Zero-Shot Atmospheric Modeling | Zero-shot atmospheric modeling foundation model based on inverse diffusion. | arXiv |
| STAR-MD | Scalable Spatio-Temporal SE(3) Diffusion for Long-Horizon Protein Dynamics | Scalable SE(3)-equivariant diffusion model for simulating long-horizon protein dynamics. | arXiv |
| MORPH | Shape-agnostic PDE Foundation Models | Shape-agnostic autoregressive PDE foundation model handling arbitrary domain geometries. | ICLR |
| PDEformer-2 | Versatile Foundation Model for 2D PDEs | Versatile 2D PDE foundation model encoding equation structure as computational graphs. | arXiv |
| NESTOR | Nested MOE Neural Operator for Large-Scale PDE Pre-Training | Nested mixture-of-experts neural operator for large-scale PDE pre-training. | arXiv |
Large-scale models for multi-physics simulation, mesh-based dynamics, and surrogate modeling.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| GPhyT | Towards a Physics Foundation Model | General physics Transformer trained on 1.8 TB of diverse simulation data with zero-shot generalization to unseen physics scenarios. | arXiv |
| PhysiX | PhysiX: A Foundation Model for Physics Simulations | 4.5B-parameter physics simulation foundation model using discrete tokenizer for multi-scale physical processes with autoregressive generation. | NeurIPS |
| PDE-Transformer | PDE-Transformer: Efficient and Versatile Transformers for Physics Simulations | Scalable Transformer architecture for efficient surrogate modeling across multiple PDE types on regular grids. | arXiv |
| M2PDE | M2PDE: Compositional Generative Multiphysics and Multi-component PDE Simulation | Compositional diffusion-based framework for generative multi-physics and multi-component PDE simulation. | arXiv |
| UPS | Unified PDE Solvers | Unified PDE solver foundation model handling cross-domain, cross-dimension, and cross-resolution spatiotemporal PDEs. | arXiv |
| CompNO | CompNO: A Novel Foundation Model approach for solving Partial Differential Equations | Compositional neural operator splitting monolithic models into composable modules for efficient parametric PDE solving. | Applied Sciences |
| GNS | Learning to Simulate Complex Physics with Graph Networks | DeepMind graph network simulator learning particle interaction rules to generalize across fluids, rigid bodies, and deformable objects. | ICML |
| MeshGraphNets | Learning Mesh-Based Simulation with Graph Networks | Graph network simulation on unstructured meshes for aerodynamics and structural mechanics. | ICLR |
| GeoPT | Scaling Physics Simulation via Lifted Geometric Pre-Training | Geometric pre-training foundation model for scaling physics simulation. | arXiv |
Generative models that learn physical dynamics for interactive environment simulation and embodied AI.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Cosmos | Cosmos World Foundation Model Platform for Physical AI | NVIDIA open-source world foundation model platform generating physics-aware video and world states for robotics and autonomous driving. | arXiv / NVIDIA |
| Genie | Genie: Generative Interactive Environments | DeepMind 11B-parameter unsupervised world model generating interactive virtual worlds from a single image. | ICML |
| Genie 2 | Genie 2: A large-scale foundation world model | Upgraded Genie generating diverse, controllable interactive 3D environments from a single image for embodied AI. | DeepMind |
| Genie 3 | Genie 3: A new frontier for world models | General world model generating consistent interactive 3D worlds from text or images in real time with physical consistency. | DeepMind |
| DIAMOND | Diffusion for World Modeling: Visual Details Matter in Atari | Diffusion-based world modeling achieving high visual fidelity for RL agent training in Atari environments. | NeurIPS 2024 |
| WorldDreamer | WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens | General world model capturing physical dynamics across multiple environments via masked token prediction for video generation. | arXiv |
| GameNGen | Diffusion Models Are Real-Time Game Engines | Google Research neural game engine using diffusion models to simulate complex game environments (DOOM) at 20+ fps. | arXiv preprint (Google) |
| OASIS | Oasis: A Universe in a Transformer | Real-time open-world AI model generating interactive Minecraft-like gameplay at 20 fps via Transformer and diffusion. | Project Page |
| Pandora | Pandora: Towards General World Model with Natural Language Actions and Video States | Hybrid autoregressive-diffusion world model controlling video state generation through natural language actions. | arXiv |
| UniSim | Learning Interactive Real-World Simulators | Universal world simulator learning to simulate diverse human-world interactions from text, actions, and image inputs. | ICLR |
| PAN | PAN: A World Model for General, Interactable, and Long-Horizon World Simulation | Action-conditioned world model for general, interactable, long-horizon simulation with environment dynamics consistency. | arXiv |
| PhysDreamer | PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation | Physics-based 3D object interaction generation via video generation for physically consistent object manipulation. | Lecture Notes in Computer Science |
| Astra | General Interactive World Model with Autoregressive Denoising | General interactive world model combining autoregressive and denoising generation. | ICLR |
Foundation models for quantum state representation, many-body simulation, and quantum dynamics.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FNQS | Foundation Model for Quantum Many-Body States via Transformers | Transformer-based foundation model pre-trained on quantum state data, generalizing across different Hamiltonians and lattice structures. | Nature Communications |
| Attention-Based FM for Quantum States | Attention-Based Foundation Model for Quantum States | Attention-based foundation model using self-attention to capture quantum correlations for cross-system quantum state representation. | arXiv |
| NOQS | Neural Operator for Quantum States | Neural operator learning continuous mappings over quantum state space for efficient quantum simulation and prediction. | arXiv |
| Large Electron Model | Large Electron Model: A Foundation Model for Electron Systems | Foundation model for electron systems pre-trained on large-scale electronic structure data, supporting quantum chemistry and materials physics tasks. | arXiv |
| DysonNet | Constant-Time Local Updates for Neural Quantum States | Neural quantum state architecture achieving O(1) local update efficiency for scalable quantum simulation. | arXiv |
Multi-modal models for tokamak plasma behavior prediction and fusion control.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| TokaMind | TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma | Multi-modal Transformer foundation model fusing multiple diagnostic modalities for tokamak plasma behavior prediction and fusion control. | arXiv |
Foundation models for optical design, thin-film structures, and photonic inverse design.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| OptoGPT | OptoGPT: A Foundation Model for Inverse Design in Optical Multilayer Thin Film Structures | GPT-based foundation model for automatic inverse design of optical multilayer thin-film structures to meet target spectral properties. | Opto-Electronic Advances |
| MOCLIP | MOCLIP: Multi-modal Optical Contrastive Learning for Inverse Photonic Design | Multi-modal optical contrastive learning model using a CLIP framework for cross-modal photonic structure inverse design. | arXiv |
Neural architectures that preserve physical symmetries, conservation laws, and geometric structure.
| Model | Paper Title | Description | Link |
|---|
Models for electronic structure calculation, wavefunction prediction, and molecular quantum properties.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Skala | Accurate and scalable exchange-correlation with deep learning | Microsoft quantum chemistry foundation model for electronic structure computation and cross-system molecular property prediction. | arXiv |
| Orbformer | Orbformer: Orbital Transformer for Electronic Structure Prediction | Orbital Transformer predicting electronic structure from molecular orbital representations with physical symmetry priors. | arXiv |
| OrbEvo | Orbital Transformers for Predicting Wavefunctions in TD-DFT | Equivariant graph Transformer predicting real-time TD-DFT wavefunction evolution. | arXiv |
Deep learning platforms for reactive flow and combustion CFD.
| Model | Paper Title | Description | Link |
|---|
Domain-specific foundation models for nuclear reactor control and simulation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| NucReactor-FM | Agentic Physical AI toward a Domain-Specific FM for Nuclear Reactor Control | Domain-specific foundation model for nuclear reactor control combining physics simulation with reinforcement learning. | arXiv |
Foundation models for global weather forecasting and climate prediction at various spatial and temporal scales.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Aurora | Aurora: A Foundation Model of the Earth System | A large-scale earth system foundation model by Microsoft trained on over one million hours of multi-source geophysical data for atmosphere, ocean, and air quality prediction. | Nature |
| Pangu-Weather | Accurate Medium-range Global Weather Forecasting with 3D Neural Networks | A 3D high-resolution AI weather forecasting model by Huawei trained on 43 years of ERA5 reanalysis data, generating 10-day global forecasts in seconds. | Nature |
| GraphCast | Learning Skillful Medium-range Global Weather Forecasting | A graph neural network-based global medium-range weather forecasting model by Google DeepMind at 0.25° resolution, outperforming ECMWF HRES for 10-day forecasts. | Science |
| GenCast | GenCast: Diffusion-based Ensemble Forecasting for Medium-range Weather | A diffusion-based ensemble weather forecasting system by Google DeepMind generating probabilistic 15-day forecasts that surpass ECMWF ENS. | Nature |
| ClimaX | ClimaX: A Foundation Model for Weather and Climate | The first weather and climate foundation model based on Transformer architecture supporting flexible fine-tuning for multiple downstream meteorological tasks. | ICML |
| FengWu | FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead | A multi-modal multi-task weather forecasting system by Shanghai AI Lab that extends deterministic forecast skill to 10.75 days. | arXiv |
| FuXi | FuXi: A Cascade Machine Learning Forecasting System for 15-day Global Weather Forecast | A cascade machine learning weather forecasting system trained on 39 years of ERA5 data, achieving 15-day forecast performance comparable to ECMWF ensemble mean. | npj Climate and Atmospheric Science |
| NeuralGCM | Neural General Circulation Models for Weather and Climate | A neural GCM by Google Research combining differentiable atmospheric dynamics solvers with machine learning for weather-to-climate timescale prediction. | Nature |
| FourCastNet | FourCastNet: A Global Data-driven High-resolution Weather Forecasting System | A high-resolution global weather forecasting model by NVIDIA using adaptive Fourier neural operators (AFNO) at 0.25° resolution. | arXiv |
| ECMWF AIFS | AIFS — ECMWF's Data-driven Forecasting System | ECMWF's operational AI forecasting system combining graph neural networks and Transformers for data-driven weather prediction. | arXiv |
| Stormer | Scaling Transformer Neural Networks for Skillful and Reliable Medium-range Weather Forecasting | A streamlined and efficient Transformer weather forecasting model achieving state-of-the-art performance with less training data. | Advances in Neural Information Processing Systems 37 |
| AtmoRep | AtmoRep: A Stochastic Model of Atmosphere Dynamics Using Large Scale Representation Learning | A task-agnostic atmospheric foundation model based on large-scale representation learning for stochastic atmosphere dynamics. | arXiv |
| WeatherGFT | WeatherGFT: Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling | A hybrid physics-AI weather forecasting model extending predictions to finer temporal resolutions at 30-minute intervals. | NeurIPS |
| WeatherGFM | WeatherGFM: Learning A Weather Generalist Foundation Model via In-context Learning | A weather generalist foundation model unifying forecasting, super-resolution, image translation, and post-processing via in-context learning. | ICLR |
| Prithvi WxC | Prithvi WxC: Foundation Model for Weather and Climate | A 2.3-billion-parameter weather and climate foundation model by IBM and NASA trained on 160 MERRA-2 variables. | arXiv |
| FuXi-2.0 | FuXi-2.0: Advancing machine learning weather forecasting model for practical applications | An upgraded version of FuXi providing hourly global forecasts with a more comprehensive set of meteorological variables. | arXiv |
| ArchesWeather | ArchesWeather: An efficient AI weather forecasting model at 1.5° resolution | A lightweight and efficient AI weather forecasting model at 1.5° resolution using a combination of 2D and column-wise attention. | arXiv |
| W-MAE | W-MAE: Pre-trained Weather Model with Masked Autoencoder | A task-agnostic atmospheric foundation model based on masked autoencoder pre-training for weather data. | arXiv |
| Omni-Weather | Omni-Weather: Unified Multimodal Foundation Model for Weather Generation and Understanding | A unified multimodal foundation model integrating radar, satellite, and numerical data for weather generation and understanding. | arXiv |
Foundation models for satellite imagery analysis, multi-spectral and multi-temporal earth observation.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Prithvi-EO-2.0 | Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications | A multi-temporal earth observation foundation model by NASA/IBM trained on 4.2 million global time-series samples supporting Landsat and Sentinel-2. | arXiv |
| SpectralGPT | SpectralGPT: Spectral Remote Sensing Foundation Model | The first spectral remote sensing foundation model using a 3D generative pre-trained Transformer designed for multi-spectral and hyperspectral satellite imagery. | IEEE TPAMI (2024) |
| SatMAE | SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery | A masked autoencoder pre-training framework for temporal and multi-spectral satellite imagery. | NeurIPS 2022 |
| SeaMo | SeaMo: A Multi-Seasonal and Multimodal Remote Sensing Foundation Model | A multi-seasonal multimodal remote sensing foundation model fusing optical, SAR, and meteorological data. | arXiv |
| TerraMind | TerraMind: Large-Scale Generative Multimodality for Earth Observation | A large-scale generative multimodal earth observation foundation model by IBM/ESA/DLR trained on 500 billion tokens. | ICCV 2025 |
| RingMo | RingMo: A Remote Sensing Foundation Model with Masked Image Modeling | A remote sensing foundation model by the Chinese Academy of Sciences using masked image modeling for large-scale pre-training. | IEEE TGRS |
| SatCLIP | SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery | A global general-purpose location encoder by Microsoft using Sentinel-2 contrastive learning to generate location embeddings. | AAAI |
| SkySense | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery | A large-scale multimodal remote sensing foundation model pre-trained on 21.5 million temporal optical and SAR data samples. | CVPR |
| Scale-MAE | Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning | A scale-aware masked autoencoder that explicitly models spatial resolution relationships for multiscale geospatial representation learning. | ICCV 2023 |
| DOFA | Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation | A neural plasticity-inspired multimodal EO foundation model using dynamic wavelength-adaptive hypernetworks to handle diverse sensor data. | arXiv |
| GFM (Prithvi-EO-1.0) | Foundation Models for Generalist Geospatial Artificial Intelligence | NASA/IBM's first-generation earth science foundation model based on self-supervised Vision Transformers trained on HLS data. | arXiv |
| S2MAE | S2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing Data | A spatial-spectral masked autoencoder providing joint spatial-spectral pre-training for spectral remote sensing imagery. | CVPR 2024 |
| RingMoE | RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation | A 14.7-billion-parameter mixture-of-modality-experts remote sensing foundation model pre-trained on 400M+ samples. | arXiv |
| WaveMAE | WaveMAE: Wavelet-decomposition Masked Autoencoder for Multispectral Satellite Imagery | A self-supervised foundation model combining wavelet decomposition with geospatial priors for multispectral satellite imagery. | arXiv |
| RoMA | RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing | A scalable Mamba-architecture remote sensing foundation model addressing ViT limitations in large-scale remote sensing pre-training. | NeurIPS |
| CROMA | CROMA: Contrastive Radar-Optical Masked Autoencoders for Remote Sensing | A contrastive radar-optical masked autoencoder for multimodal remote sensing representation learning. | NeurIPS |
| AnySat | AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities | A unified multi-resolution multimodal earth observation model using JEPA architecture for diverse EO tasks. | CVPR |
| TerraFM | TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation | A scalable self-supervised foundation model for unified multisensor earth observation pre-trained on 18.7 million samples. | ICLR 2026 |
Foundation models for ocean forecasting, eddy-resolving prediction, and marine environment monitoring.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| OceanGPT | OceanGPT: A Large Language Model for Ocean Science Tasks | A domain-specific large language model for ocean science by Zhejiang University using the DoInstruct framework for ocean domain instruction data. | ACL 2024 |
| XiHe | XiHe: A Data-Driven Model for Global Ocean Eddy-Resolving Forecasting | A data-driven global ocean eddy-resolving forecast model at 1/12° resolution trained on 25 years of reanalysis data. | arXiv |
| WV-Net | WV-Net: A Foundation Model for SAR WV-mode Satellite Imagery Trained Using Contrastive Self-supervised Learning | The first foundation model for SAR ocean satellite imagery using self-supervised contrastive learning on synthetic aperture radar data. | arXiv |
| GLONET | GLONET: Mercator's End-to-End Neural Global Ocean Forecasting System | An end-to-end neural network global ocean forecasting system by Mercator Ocean trained on GLORYS12 reanalysis data. | Journal of Geophysical Research: Machine Learning and Computation |
| WenHai | Forecasting the Eddying Ocean with a Deep Neural Network | A deep neural network ocean forecasting system excelling at mesoscale eddy dynamics prediction. | Nature Communications |
| FuXi-Ocean | FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution | A data-driven global ocean forecasting system with 6-hour temporal and 1/12° spatial resolution reaching 1500-meter depth. | NeurIPS |
| ORCA-DL | Data-driven Global Ocean Modeling for Seasonal to Decadal Prediction | A data-driven global ocean model supporting 3D ocean predictions from seasonal to decadal timescales. | Science Advances |
| FuXi-ONS | Data-driven Ensemble Prediction of the Global Ocean | A machine-learning ensemble global ocean forecasting system for 5-to-365-day predictions. | arXiv |
Foundation models for earthquake detection, seismic phase picking, and waveform analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| PhaseNet | PhaseNet: A Deep-Neural-Network-Based Seismic Arrival Time Picking Method | One of the most widely used deep learning models for seismic arrival-time picking in seismology. | Geophysical Journal International |
| EQTransformer | EQTransformer: An Attentive Deep-Learning Model for Simultaneous Earthquake Detection and Phase Picking | An attention-based deep learning model for simultaneous earthquake detection and seismic phase picking. | Nature Communications |
| SeisT | SeisT: A Foundational Deep-Learning Model for Earthquake Monitoring Tasks | A Transformer-based seismic monitoring foundation model supporting multiple earthquake tasks including detection, phase picking, and magnitude estimation. | IEEE Transactions on Geoscience and Remote Sensing |
| SeisLM | SeisLM: a Foundation Model for Seismic Waveforms | A large-scale self-supervised seismic waveform foundation model pre-trained via contrastive learning on massive open-source seismic data. | arXiv |
| SeisMoLLM | SeisMoLLM: Advancing Seismic Monitoring via Cross-modal Transfer with Pre-trained Large Language Model | A seismic monitoring foundation model leveraging cross-modal transfer from GPT-2 architecture for seismic analysis. | arXiv |
| SeismicXM | SeismicXM: A Cross-Task Foundation Model for Single-Station Seismic Waveform Processing | A cross-task seismic waveform processing foundation model by China Earthquake Administration supporting multiple single-station tasks. | SRL |
| U-Trans | U-Trans: A Foundation Model for Seismic Waveform Representation | A U-Net encoder-decoder architecture seismic waveform representation foundation model trained on 2M+ three-component waveforms. | Scientific Reports |
| PhaseNet+ | PhaseNet+: Towards End-to-End Earthquake Monitoring Using a Multitask Deep Learning Model | A multi-task extension of PhaseNet enabling end-to-end earthquake monitoring. | arXiv |
| SeisCLIP | SeisCLIP: Contrastive Multimodal Seismology Foundation Model | A contrastive multimodal seismology foundation model learning joint representations from seismic waveforms and metadata. | arXiv |
Foundation models for hydrological prediction, flood modeling, and river forecasting.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| HydroGAT | HydroGAT: Distributed Heterogeneous Graph Attention Transformer for Spatiotemporal Flood Prediction | A graph attention network-based hydrological prediction foundation model capturing spatial dependencies across watersheds. | arXiv |
| ZeroFlood | ZeroFlood: A Geospatial Foundation Model for Data-Efficient Flood Susceptibility Mapping | A geospatial foundation model for data-efficient flood susceptibility mapping. | arXiv |
| GraphRiverCast | Topology-informed AI Foundation Model for Global River Forecasting | A topology-informed AI foundation model for global river hydrodynamic forecasting. | arXiv |
Models for wildfire danger forecasting and fire spread prediction.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| FireCastNet | FireCastNet: earth-as-a-graph for seasonal fire prediction | A deep learning global wildfire danger forecasting model fusing meteorological, vegetation, and terrain data for multi-timescale prediction. | Scientific Reports |
| FireScope | FireScope: Wildfire Risk Prediction with a Chain-of-Thought Oracle | A wildfire risk prediction foundation model and benchmark using multi-modal data for fire risk assessment. | arXiv |
Foundation models for atmospheric pollution forecasting.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| AirCast | AirCast: Improving Air Pollution Forecasting Through Multi-Variable Data Alignment | A data-driven air quality forecasting foundation model supporting multi-variable atmospheric pollutant concentration prediction. | arXiv |
| FuXi-Air | FuXi-Air: Urban Air Quality Forecasting Based on Emission-Meteorology-Pollutant multimodal Machine Learning | A multimodal machine learning air quality forecasting extension of the FuXi series integrating emission, meteorological, and observational data. | arXiv |
Foundation models for sea ice monitoring and polar region forecasting.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| SIFM | SIFM: A Foundation Model for Multi-granularity Arctic Sea Ice Forecasting | A sea ice foundation model for multi-granularity Arctic sea ice concentration forecasting from satellite observations. | arXiv |
| IceNet | IceNet: Seasonal Arctic Sea Ice Forecasting with Probabilistic Deep Learning | A probabilistic deep learning model for seasonal Arctic sea ice forecasting that significantly outperforms dynamical physics models. | Nature Communications |
| IceMamba | IceMamba: Seasonal Forecasting of Pan-Arctic Sea Ice with State Space Model | A Mamba state space model-based sea ice forecasting foundation model efficiently processing polar spatiotemporal sequence data. | arXiv |
Large language models specialized for earth science knowledge understanding and reasoning.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| K2 | K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization | The first 7-billion-parameter geoscience LLM based on LLaMA, further pre-trained and instruction-tuned on earth science literature. | Proceedings of the 17th ACM International Conference on Web Search and Data Mining |
| JiuZhou | JiuZhou: Open Foundation Language Models for Geoscience | A multilingual geoscience LLM by Tsinghua University supporting Chinese and English earth science knowledge QA and reasoning. | GitHub |
| GeoGPT | GeoGPT: A Large Language Model for Geospatial Artificial Intelligence | A geospatial AI LLM by Zhejiang Lab combining tool-calling capabilities for geospatial analysis and reasoning. | GitHub |
| GeoGalactica | GeoGalactica: A Scientific Large Language Model for Geoscience | A 30-billion-parameter geoscience LLM based on Galactica architecture pre-trained on earth science corpora. | arXiv |
Foundation models for subsurface characterization, seismic exploration, and well log analysis.
| Model | Paper Title | Description | Link |
|---|---|---|---|
| Transparent Earth | The Transparent Earth: A Multimodal Foundation Model for the Earth's Subsurface | A multimodal transformer-based foundation model by LANL for subsurface structure imaging and inversion. | arXiv |
| GEM 3D | Geological Everything Model 3D: A Promptable Foundation Model for Subsurface Understanding | A promptable generative 3D earth model unifyin |
Truncated — view the full README on GitHub.