All Texts are translated with DeepL. (Machine Translated.)
EverythingLM V2 is a diverse instruct dataset consisting of 1k of human-assistant conversations. These sets were generated using principles from both evol-instruct and Orca. The dataset encompasses a wide array of topics and interactions.
Reproducing this dataset would cost roughly $40.
We also leverage various system prompts for evol-instruct and for responding to prompts. This dataset has also been filtered to remove OpenAI alignment.
Included in this repo is the script to generate the dataset.
5 commits
All Texts are translated with DeepL. (Machine Translated.)
EverythingLM V2 is a diverse instruct dataset consisting of 1k of human-assistant conversations. These sets were generated using principles from both evol-instruct and Orca. The dataset encompasses a wide array of topics and interactions.
Reproducing this dataset would cost roughly $40.
We also leverage various system prompts for evol-instruct and for responding to prompts. This dataset has also been filtered to remove OpenAI alignment.
Included in this repo is the script to generate the dataset.
5 commits