Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
Tulu is a series of language models that are trained to act as helpful assistants. The dataset consists of a mix of :
These are made by taking either just the training set of the subsets or the entire section if no splits are present. Tulu V2 is presented as a singular training split. Tulu V2 DPO 70B, and is a fine-tuned version of Llama 2 that was trained on on a mix of publicly available, synthetic and human datasets using Direct Preference Optimization (DPO).
Model Family: Other models and the dataset are found in the Tulu V2 collection.
The length distribution of the dataset can be seen below:

Tulu V1 Mix can be found here.
Note: Some samples contain empty turns as noted in this github issue. We will not remove these from this release to ensure reproducibility but you may wish to explicitly filter them out when training your own models!
The included science data is from the following categories:
Note that some of the examples include an off-by-one error in the sentence indexing that had a small or negligible impact on performance. This was found during testing and will be updated in future versions, with the detailed release of the dataset artifact itself coming in a future release.
We are releasing this dataset under the terms of ODC-BY. By using this, you are also bound by the Common Crawl terms of use in respect of the content contained in the dataset.
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
Tulu is a series of language models that are trained to act as helpful assistants. The dataset consists of a mix of :
These are made by taking either just the training set of the subsets or the entire section if no splits are present. Tulu V2 is presented as a singular training split. Tulu V2 DPO 70B, and is a fine-tuned version of Llama 2 that was trained on on a mix of publicly available, synthetic and human datasets using Direct Preference Optimization (DPO).
Model Family: Other models and the dataset are found in the Tulu V2 collection.
The length distribution of the dataset can be seen below:

Tulu V1 Mix can be found here.
Note: Some samples contain empty turns as noted in this github issue. We will not remove these from this release to ensure reproducibility but you may wish to explicitly filter them out when training your own models!
The included science data is from the following categories:
Note that some of the examples include an off-by-one error in the sentence indexing that had a small or negligible impact on performance. This was found during testing and will be updated in future versions, with the detailed release of the dataset artifact itself coming in a future release.
We are releasing this dataset under the terms of ODC-BY. By using this, you are also bound by the Common Crawl terms of use in respect of the content contained in the dataset.