A subjective learning guide for generative AI research including curated list of articles and projects
Generative AI is a hot topic today :fire: and this roadmap is designed to help beginners quickly gain basic knowledge and find useful resources of Generative AI. Even experts are welcome to refer to this roadmap to recall old knowledge and develop new ideas.
This section should help you learn or regain the basic knowledge of neural networks (e.g., backpropagation), get you familiar with the transformer architecture, and describe some common transformer-based models.
Are you very familiar with the following classic neural network structures?
:pencil: If so, you should be able to answer these questions:
Backpropagation (BP) is the base of NN training. You will not be an AI expert if you don't understand BP. There are many textbooks and online tutorials teaching BP, but unfortunately, most of them don't present formulas in vectorized/tensorized forms. The BP formula of an NN layer is indeed as neat as its forward pass formula. This is exactly how BP is implemented and should be implemented. To understand BP, please read the following materials:
:pencil: If you understand BP, you should be able to answer these questions:
Transformer is the base architecture of existing large generative models. It's necessary to understand every component in the transformer. Please read the following materials:
:pencil: If you understand transformers, you should be able to answer these questions:
Einsum is easy and useful [A great tutorial for using einsum/einops]
Open-Endedness is Essential for Artificial Superhuman Intelligence (ICML 2024) [Thoughts on achieving superhuman AI]
Levels of AGI for Operationalizing Progress on the Path to AGI
LLMs are transformers. They can be categorized into encoder-only, encoder-decoder, and decoder-only architectures, as shown in the LLM evolutionary tree below [image source]. Check milestone papers of LLMs.

Encoder-only model can be used to extract sentence features but lacks generative power. Encoder-decoder and decoder-only models are used for text generation. In particular, most existing LLMs prefer decoder-only structures due to stronger repesentational power. Intuitively, encoder-decoder models can be considered a sparse version of decoder-only models and the information decays more from encoder to decoder. Check this paper for more details.
LLMs are typically pretrained from trillions of text tokens by model publishers to internalize the natural language structure. Today's model developers also conduct instructional fine-tuning and Reinforcement Learning from Human Feedback (RLHF) to teach the model to follow human instructions and generate answers aligned with human preference. The users can then download the published model and finetune it on small personal datasets (e.g., movie dialog). Due to huge amount of data, pretraining requires massive computing resources (e.g., more than thousands of GPUs) which is unaffordable by individuals. On the other hand, fine-tuning is less resource-hungry and can be done with a few GPUs.
The following materials can help you understand the pretraining and fine-tuning process:
More tutorials can be found here.
Prompting techniques for LLMs involve crafting input text in a way that guides the model to generate desired responses or outputs. Here are the useful resources to help you write better prompts:
Evaluation tools for large language models help assess their performance, capabilities, and limitations across different tasks and datasets. Here are some common evaluation strategies:
Automatic Evaluation Metrics: These metrics assess model performance automatically without human intervention. Common metrics include:
Human Evaluation: Human judgment is essential for assessing the quality of generated text comprehensively. Common human evaluation methods include:
Benchmark Datasets: Standardized datasets enable fair comparison of models across different tasks and domains. Examples include:
Model Analysis Tools: Tools for analyzing model behavior and performance include:
A complete list can be found here
Standard evaluation frameworks for existing LLMs include:
Dealing with long contexts poses a challenge for large language models due to limitations in memory and processing capacity. Existing techniques include:
A complete list can be found here
Parameter-Efficient Fine-Tuning (PEFT) methods enable efficient adaptation of large pretrained models to various downstream applications by only fine-tuning a small number of (extra) model parameters instead of all the model's parameters:
More work can be found in Huggingface PEFT paper collection and it's highly recommended to practice with HuggingFace PEFT API.
Model merging refers to merging two or more LLMs trained on different tasks into a single LLM. This technique aims to leverage the strengths and knowledge of different models to create a more robust and capable model. For example, a LLM for code generation and another LLM for math prolem solving can be merged together so that the merged model is capable of doing both code generation and math problem solving.
The model merging is intriguing because it can be effectively achieved with very simple and cheap algorithms (e.g., linear combination of model weights). Here are some representative papers and reading materials:
More papers about model merging can be found here
Accelerating decoding of LLMs is crucial for improving inference speed and efficiency, especially in real-time or latency-sensitive applications. Here are some representative work of speeding up decoding process of LLMs:
More work about accelerating LLM decoding can be found via Link 1 and Link 2.
Knowledge editing aims to efficiently modify LLMs behaviors, such as reducing bias and revising learned correlations. It includes many topics such as knowledge localization and unlearning. Representative work includes:
More papers can be found here.
By receiving massive training, LLMs digest world knowledge and are able to follow input instructions precisely. With these amazing capabilities, LLMs can play as agents that are possible to autonomously (and collaboratively) solve complex tasks, or simulate human interactions. Here are some representative papers of LLM agents:
A complete list of papers, platforms, and evaluation tools can be found here.
LLMs face several open challenges that researchers and developers are actively working to address. These challenges include:
A complete list can be found here.
Diffusion models aim to approxmiate the probability distribution of a given data domain, and provide a way to generate samples from its approximated distribution. Their goals are similar to other popular generative models, such as VAE, GANs, and Normalizing Flows.
The working flow of diffusion models is featured with two process:
Check this awesome blog and more introductory tutorials can be found here. Diffusion models can be used to generate images, audios, videos, and more, and there are many subfields related to diffusion models as shown below [image source]:

Here are some representative papers of diffusion models for image generation:
More papers can be found here.
Here are some representative papers of diffusion models for video generation:
More papers can be found here.
Here are some representative papers of diffusion models for audio generation:
More papers can be found here.
Similar to other large generative models, diffusion models are also pretrained on large amount of web data (e.g., LAION-5B dataset) and consume massive computing resources. Users can download the released weights can further fine-tune the model on personal datasets.
Here are some representative papers of efficient fine-tuning of diffusion models:
More papers can be found here.
It's highly recommended to do some practice with Huggingface Diffusers API.
Here we talk about evaluation of diffusion models for image generation. Many existing image quality metrics can be applied.
More image quality metrics and calculation tools can be found here.
Diffusion models require multiple forward steps over to generate data, which is expensive. Here are some representative papers of diffusion models for efficient generation:
More papers can be found here.
Here are some representative papers of knowledge editing for diffusion models:
More papers can be found here.
Here are some survey papers talking about the challenges faced by diffusion models.
Typical LMMs are constructed by connecting and fine-tuning existing pretrained unimodal models. Some are also pretrained from scratch. Check how LMMs evolve in the image below [image source].

There are many different ways of contructing LMMs. Representative architectures include:
More papers can be found via Link 1 and Link 2.
By combining LMMs with robots, researchers aim to develop AI systems that can perceive, reason about, and act upon the world in a more natural and intuitive way, with potential applications spanning robotics, virtual assistants, autonomous vehicles, and beyond. Here are some representative work of realizing embodied AI with LMMs:
More papers can be found via Link 1 and Link 2.
Here are some popular simulators and datasets to evaluate LMMs performance for embodied AI:
More resources can be found here.
Here are some survey papers talking about open challenges for LMM-enabled embodied AI:
Researchers are trying to explore new models other than transformers. The efforts include implicitly structuring model parameters and defining new model architectures.
Here is an awesome tutorial for state space models.
28 commits
A subjective learning guide for generative AI research including curated list of articles and projects
Generative AI is a hot topic today :fire: and this roadmap is designed to help beginners quickly gain basic knowledge and find useful resources of Generative AI. Even experts are welcome to refer to this roadmap to recall old knowledge and develop new ideas.
This section should help you learn or regain the basic knowledge of neural networks (e.g., backpropagation), get you familiar with the transformer architecture, and describe some common transformer-based models.
Are you very familiar with the following classic neural network structures?
:pencil: If so, you should be able to answer these questions:
Backpropagation (BP) is the base of NN training. You will not be an AI expert if you don't understand BP. There are many textbooks and online tutorials teaching BP, but unfortunately, most of them don't present formulas in vectorized/tensorized forms. The BP formula of an NN layer is indeed as neat as its forward pass formula. This is exactly how BP is implemented and should be implemented. To understand BP, please read the following materials:
:pencil: If you understand BP, you should be able to answer these questions:
Transformer is the base architecture of existing large generative models. It's necessary to understand every component in the transformer. Please read the following materials:
:pencil: If you understand transformers, you should be able to answer these questions:
Einsum is easy and useful [A great tutorial for using einsum/einops]
Open-Endedness is Essential for Artificial Superhuman Intelligence (ICML 2024) [Thoughts on achieving superhuman AI]
Levels of AGI for Operationalizing Progress on the Path to AGI
LLMs are transformers. They can be categorized into encoder-only, encoder-decoder, and decoder-only architectures, as shown in the LLM evolutionary tree below [image source]. Check milestone papers of LLMs.

Encoder-only model can be used to extract sentence features but lacks generative power. Encoder-decoder and decoder-only models are used for text generation. In particular, most existing LLMs prefer decoder-only structures due to stronger repesentational power. Intuitively, encoder-decoder models can be considered a sparse version of decoder-only models and the information decays more from encoder to decoder. Check this paper for more details.
LLMs are typically pretrained from trillions of text tokens by model publishers to internalize the natural language structure. Today's model developers also conduct instructional fine-tuning and Reinforcement Learning from Human Feedback (RLHF) to teach the model to follow human instructions and generate answers aligned with human preference. The users can then download the published model and finetune it on small personal datasets (e.g., movie dialog). Due to huge amount of data, pretraining requires massive computing resources (e.g., more than thousands of GPUs) which is unaffordable by individuals. On the other hand, fine-tuning is less resource-hungry and can be done with a few GPUs.
The following materials can help you understand the pretraining and fine-tuning process:
More tutorials can be found here.
Prompting techniques for LLMs involve crafting input text in a way that guides the model to generate desired responses or outputs. Here are the useful resources to help you write better prompts:
Evaluation tools for large language models help assess their performance, capabilities, and limitations across different tasks and datasets. Here are some common evaluation strategies:
Automatic Evaluation Metrics: These metrics assess model performance automatically without human intervention. Common metrics include:
Human Evaluation: Human judgment is essential for assessing the quality of generated text comprehensively. Common human evaluation methods include:
Benchmark Datasets: Standardized datasets enable fair comparison of models across different tasks and domains. Examples include:
Model Analysis Tools: Tools for analyzing model behavior and performance include:
A complete list can be found here
Standard evaluation frameworks for existing LLMs include:
Dealing with long contexts poses a challenge for large language models due to limitations in memory and processing capacity. Existing techniques include:
A complete list can be found here
Parameter-Efficient Fine-Tuning (PEFT) methods enable efficient adaptation of large pretrained models to various downstream applications by only fine-tuning a small number of (extra) model parameters instead of all the model's parameters:
More work can be found in Huggingface PEFT paper collection and it's highly recommended to practice with HuggingFace PEFT API.
Model merging refers to merging two or more LLMs trained on different tasks into a single LLM. This technique aims to leverage the strengths and knowledge of different models to create a more robust and capable model. For example, a LLM for code generation and another LLM for math prolem solving can be merged together so that the merged model is capable of doing both code generation and math problem solving.
The model merging is intriguing because it can be effectively achieved with very simple and cheap algorithms (e.g., linear combination of model weights). Here are some representative papers and reading materials:
More papers about model merging can be found here
Accelerating decoding of LLMs is crucial for improving inference speed and efficiency, especially in real-time or latency-sensitive applications. Here are some representative work of speeding up decoding process of LLMs:
More work about accelerating LLM decoding can be found via Link 1 and Link 2.
Knowledge editing aims to efficiently modify LLMs behaviors, such as reducing bias and revising learned correlations. It includes many topics such as knowledge localization and unlearning. Representative work includes:
More papers can be found here.
By receiving massive training, LLMs digest world knowledge and are able to follow input instructions precisely. With these amazing capabilities, LLMs can play as agents that are possible to autonomously (and collaboratively) solve complex tasks, or simulate human interactions. Here are some representative papers of LLM agents:
A complete list of papers, platforms, and evaluation tools can be found here.
LLMs face several open challenges that researchers and developers are actively working to address. These challenges include:
A complete list can be found here.
Diffusion models aim to approxmiate the probability distribution of a given data domain, and provide a way to generate samples from its approximated distribution. Their goals are similar to other popular generative models, such as VAE, GANs, and Normalizing Flows.
The working flow of diffusion models is featured with two process:
Check this awesome blog and more introductory tutorials can be found here. Diffusion models can be used to generate images, audios, videos, and more, and there are many subfields related to diffusion models as shown below [image source]:

Here are some representative papers of diffusion models for image generation:
More papers can be found here.
Here are some representative papers of diffusion models for video generation:
More papers can be found here.
Here are some representative papers of diffusion models for audio generation:
More papers can be found here.
Similar to other large generative models, diffusion models are also pretrained on large amount of web data (e.g., LAION-5B dataset) and consume massive computing resources. Users can download the released weights can further fine-tune the model on personal datasets.
Here are some representative papers of efficient fine-tuning of diffusion models:
More papers can be found here.
It's highly recommended to do some practice with Huggingface Diffusers API.
Here we talk about evaluation of diffusion models for image generation. Many existing image quality metrics can be applied.
More image quality metrics and calculation tools can be found here.
Diffusion models require multiple forward steps over to generate data, which is expensive. Here are some representative papers of diffusion models for efficient generation:
More papers can be found here.
Here are some representative papers of knowledge editing for diffusion models:
More papers can be found here.
Here are some survey papers talking about the challenges faced by diffusion models.
Typical LMMs are constructed by connecting and fine-tuning existing pretrained unimodal models. Some are also pretrained from scratch. Check how LMMs evolve in the image below [image source].

There are many different ways of contructing LMMs. Representative architectures include:
More papers can be found via Link 1 and Link 2.
By combining LMMs with robots, researchers aim to develop AI systems that can perceive, reason about, and act upon the world in a more natural and intuitive way, with potential applications spanning robotics, virtual assistants, autonomous vehicles, and beyond. Here are some representative work of realizing embodied AI with LMMs:
More papers can be found via Link 1 and Link 2.
Here are some popular simulators and datasets to evaluate LMMs performance for embodied AI:
More resources can be found here.
Here are some survey papers talking about open challenges for LMM-enabled embodied AI:
Researchers are trying to explore new models other than transformers. The efforts include implicitly structuring model parameters and defining new model architectures.
Here is an awesome tutorial for state space models.
28 commits