git clone https://github.com/ARENA-education/ARENA_materials.git
ARENA_materials/install.sh
This GitHub repo hosts the exercises and Streamlit pages for the ARENA program.
You can find a summary of each of the chapters below. For more detailed information (including the different ways you can access the exercises), click on the links in the chapter headings.
Scroll to the end to see our instructions for submitting PRs.
The material on this page covers the first five days of the curriculum. It can be seen as a grounding in all the fundamentals necessary to complete the more advanced sections of this course (such as RL, transformers, mechanistic interpretability, training at scale, and generative models).
Some highlights from this chapter include:
The material on this page covers transformers (what they are, how they are trained, how they are used to generate output) as well as mechanistic interpretability (what it is, what are some of the most important results in the field so far, why it might be important for alignment) and other topics related to interpretability (function vectors & model steering).
Some highlights from this chapter include:
Unlike the first chapter (where all the material was compulsory), all sections of this chapter are optional extensions other than the first two exercise sets. In these first two sets you will build and train transformers, and gain a basic understanding of mechanistic interpretability of transformer models which includes induction heads & use of TransformerLens. After this, you can pick any of the other six sets of exercises you want - there are no prerequisites!
If you've finished the compulsory material and are choosing between the other six sets of exercises, we weakly recommend choosing one of the first three (IOI, superposition, and function vectors). IOI should appeal to the experimentalists, superposition to the theorists / mathematicians, and function vectors to the engineers, so there's something for everyone!
Additionally, each optional set of exercises includes a lot of suggested bonus material / further exploration once you've finished, including suggested papers to read and replicate.
Reinforcement learning is an important field of machine learning. It works by teaching agents to take actions in an environment to maximise their accumulated reward.
In this chapter, you will be learning about some of the fundamentals of RL, and working with OpenAI’s Gym environment to run your own experiments.
Some highlights from this chapter include:
Additionally, the later exercise sets include a lot of suggested bonus material / further exploration once you've finished, including suggested papers to read and replicate.
The material in this chapter covers LLM evaluations (what they are for, how to design and build one). Evals produce empirical evidence on the model's capabilities and behavioral tendencies, which allows developers and regulators to make important decisions about training or deploying the model. In this chapter, you will learn the fundamentals of two types of eval: designing a simple multiple-choice (MC) question evaluation benchmark and building an LLM agent for an agent task to evaluate model capabilities with scaffolding.
Some highlights from this chapter include:
The exercises are written in collaboration with Apollo Research, and designed to give you the foundational skills for doing safety evaluation research on language models.
Coming soon!
If you want to submit a PR to the repo (e.g. fixing a bug or typo), this would be much appreciated! The easiest way to do this for us is by editing the master Python file (not the notebook) in infrastructure/master_files, since these are the files that generate all other pages (both Colabs, Streamlit pages, solutions files). For example, if you want to edit material 2.2 (Q-Learning and DQN), you should edit just infrastructure/master_files/master_2_2.py. After PRs are merged, we then run code which updates all the other files based on this one (so you don't have to worry about any of those other files, only the master Python file!).
If you find the PR confusing (because you're not sure exactly what to edit in these master files), then please either send a message in the #errata Slack channel, or just make a PR on non-master files (e.g. the solutions.py file or the markdown files) and we'll be able to merge it & replicate it on the master files.
40 followers · starred Nov 2024
76 followers · starred Feb 2024
29 followers · starred Oct 2024
90 followers · starred Aug 2024
Jupyter Notebook
74.8%
Python
24.0%
git clone https://github.com/ARENA-education/ARENA_materials.git
ARENA_materials/install.sh
This GitHub repo hosts the exercises and Streamlit pages for the ARENA program.
You can find a summary of each of the chapters below. For more detailed information (including the different ways you can access the exercises), click on the links in the chapter headings.
Scroll to the end to see our instructions for submitting PRs.
The material on this page covers the first five days of the curriculum. It can be seen as a grounding in all the fundamentals necessary to complete the more advanced sections of this course (such as RL, transformers, mechanistic interpretability, training at scale, and generative models).
Some highlights from this chapter include:
The material on this page covers transformers (what they are, how they are trained, how they are used to generate output) as well as mechanistic interpretability (what it is, what are some of the most important results in the field so far, why it might be important for alignment) and other topics related to interpretability (function vectors & model steering).
Some highlights from this chapter include:
Unlike the first chapter (where all the material was compulsory), all sections of this chapter are optional extensions other than the first two exercise sets. In these first two sets you will build and train transformers, and gain a basic understanding of mechanistic interpretability of transformer models which includes induction heads & use of TransformerLens. After this, you can pick any of the other six sets of exercises you want - there are no prerequisites!
If you've finished the compulsory material and are choosing between the other six sets of exercises, we weakly recommend choosing one of the first three (IOI, superposition, and function vectors). IOI should appeal to the experimentalists, superposition to the theorists / mathematicians, and function vectors to the engineers, so there's something for everyone!
Additionally, each optional set of exercises includes a lot of suggested bonus material / further exploration once you've finished, including suggested papers to read and replicate.
Reinforcement learning is an important field of machine learning. It works by teaching agents to take actions in an environment to maximise their accumulated reward.
In this chapter, you will be learning about some of the fundamentals of RL, and working with OpenAI’s Gym environment to run your own experiments.
Some highlights from this chapter include:
Additionally, the later exercise sets include a lot of suggested bonus material / further exploration once you've finished, including suggested papers to read and replicate.
The material in this chapter covers LLM evaluations (what they are for, how to design and build one). Evals produce empirical evidence on the model's capabilities and behavioral tendencies, which allows developers and regulators to make important decisions about training or deploying the model. In this chapter, you will learn the fundamentals of two types of eval: designing a simple multiple-choice (MC) question evaluation benchmark and building an LLM agent for an agent task to evaluate model capabilities with scaffolding.
Some highlights from this chapter include:
The exercises are written in collaboration with Apollo Research, and designed to give you the foundational skills for doing safety evaluation research on language models.
Coming soon!
If you want to submit a PR to the repo (e.g. fixing a bug or typo), this would be much appreciated! The easiest way to do this for us is by editing the master Python file (not the notebook) in infrastructure/master_files, since these are the files that generate all other pages (both Colabs, Streamlit pages, solutions files). For example, if you want to edit material 2.2 (Q-Learning and DQN), you should edit just infrastructure/master_files/master_2_2.py. After PRs are merged, we then run code which updates all the other files based on this one (so you don't have to worry about any of those other files, only the master Python file!).
If you find the PR confusing (because you're not sure exactly what to edit in these master files), then please either send a message in the #errata Slack channel, or just make a PR on non-master files (e.g. the solutions.py file or the markdown files) and we'll be able to merge it & replicate it on the master files.
40 followers · starred Nov 2024
76 followers · starred Feb 2024
29 followers · starred Oct 2024
90 followers · starred Aug 2024
Jupyter Notebook
74.8%
Python
24.0%