Language Modeling Research Hub, a comprehensive compendium for enthusiasts and scholars delving into the fascinating realm of language models (LMs), with a particular focus on large language models (LLMs)
Python
20
197 commits
updated Mar 19, 2025
Welcome to the Language Modeling Research Hub, a comprehensive compendium for enthusiasts and scholars delving into the fascinating realm of language models (LMs), with a particular focus on large language models (LLMs). This repository is meticulously organized to facilitate easy navigation and access to a wealth of information.
Dive in to explore, contribute, and expand the horizons of language modeling with us!
"In theory, theory and practice are the same. In practice, they are not." --Albert Einstein
Reading list and related notes for LLM research, see Reading List for details.
- Key Findings
- Architecture
- Causality
- Code Learning
- Dialogue
- Efficiency
- Human Alignment
- Information Extraction
- Instruction Tuning
- Interpretability
- In Context Learning
- Knowledge Update
- Mixture of Experts (MoE)
- Non-Autoregressive Generation
- Reasoning
- Abstract Reasoning
- Chain of Thought
- Symbolic Reasoning
- Retrieval
- Social
My notes for LM research, see notes for details.
- Tokenization
- Position Encoding
Collection of various open-source LLMs, see Open-source LLMs for details.
- Pretrained Model
- Multitask Supervised Finetuned Model
- Instruction Finetuned Model
- English
- Chinese
- Multilingual
- Human Feedback Finetuned Model
- Domain Finetuned Model
- Open Source Projects
- reproduce/framework
- accelerate
- evaluation
- deployment/demo
Related Collections
Datasets for Pretrain/Finetune/Instruction-tune LLMs, see Datasets for details.
- Pretraining Corpora
- Instruction
Related Collections
Collection of automatic evaluation benchmarks, see Evaluation Benchmarks for details.
- English
- Comprehensive
- Knowledge
- Reason
- Hard Mathematical, Theorem
- Code
- Personalization
- Chinese
- Comprehensive
- Safety
- Multilingual
Related Collections
Evaluations of LLM , The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models"
EvaluationPapers4ChatGPT , Resource, Evaluation and Detection Papers for ChatGPT
LLMs-for-NLG-Evaluation , Awesome LLM for NLG Evaluation Papers
Collection of tricks of writing a perfect prompt, see Prompt for details.
LLM API demos (including mirror links), see API for details.
- openai
see Instruction Tuning for details.
- Experiments
- Datasets
- Collection
- Bootstrap
- Model Cards
- Usage
- Results
constrain LLM to generate specific answer (e.g., some open ended QA, limited vocabulary tasks), see Constrained Generate for details.
- Common method (constrain vocabulary + sample algorithm)
- Trie + Beam search (has issues currently)
Python
51.2%
Jupyter Notebook
46.5%
sed
2.4%
Language Modeling Research Hub, a comprehensive compendium for enthusiasts and scholars delving into the fascinating realm of language models (LMs), with a particular focus on large language models (LLMs)
Python
20
197 commits
updated Mar 19, 2025
Welcome to the Language Modeling Research Hub, a comprehensive compendium for enthusiasts and scholars delving into the fascinating realm of language models (LMs), with a particular focus on large language models (LLMs). This repository is meticulously organized to facilitate easy navigation and access to a wealth of information.
Dive in to explore, contribute, and expand the horizons of language modeling with us!
"In theory, theory and practice are the same. In practice, they are not." --Albert Einstein
Reading list and related notes for LLM research, see Reading List for details.
- Key Findings
- Architecture
- Causality
- Code Learning
- Dialogue
- Efficiency
- Human Alignment
- Information Extraction
- Instruction Tuning
- Interpretability
- In Context Learning
- Knowledge Update
- Mixture of Experts (MoE)
- Non-Autoregressive Generation
- Reasoning
- Abstract Reasoning
- Chain of Thought
- Symbolic Reasoning
- Retrieval
- Social
My notes for LM research, see notes for details.
- Tokenization
- Position Encoding
Collection of various open-source LLMs, see Open-source LLMs for details.
- Pretrained Model
- Multitask Supervised Finetuned Model
- Instruction Finetuned Model
- English
- Chinese
- Multilingual
- Human Feedback Finetuned Model
- Domain Finetuned Model
- Open Source Projects
- reproduce/framework
- accelerate
- evaluation
- deployment/demo
Related Collections
Datasets for Pretrain/Finetune/Instruction-tune LLMs, see Datasets for details.
- Pretraining Corpora
- Instruction
Related Collections
Collection of automatic evaluation benchmarks, see Evaluation Benchmarks for details.
- English
- Comprehensive
- Knowledge
- Reason
- Hard Mathematical, Theorem
- Code
- Personalization
- Chinese
- Comprehensive
- Safety
- Multilingual
Related Collections
Evaluations of LLM , The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models"
EvaluationPapers4ChatGPT , Resource, Evaluation and Detection Papers for ChatGPT
LLMs-for-NLG-Evaluation , Awesome LLM for NLG Evaluation Papers
Collection of tricks of writing a perfect prompt, see Prompt for details.
LLM API demos (including mirror links), see API for details.
- openai
see Instruction Tuning for details.
- Experiments
- Datasets
- Collection
- Bootstrap
- Model Cards
- Usage
- Results
constrain LLM to generate specific answer (e.g., some open ended QA, limited vocabulary tasks), see Constrained Generate for details.
- Common method (constrain vocabulary + sample algorithm)
- Trie + Beam search (has issues currently)
Python
51.2%
Jupyter Notebook
46.5%
sed
2.4%