nlpai-lab/openassistant-guanaco-ko

Dataset

11

stars

3

commits

1

linked in READMEs

Jun 1, 2023

updated

README

Dataset Summary

Korean translation of Guanaco via the DeepL API

Note: There are cases where multilingual data has been converted to monolingual data during batch translation to Korean using the API.

Below is Guanaco's README.


This dataset is a subset of the Open Assistant dataset, which you can find here: https://huggingface.co/datasets/OpenAssistant/oasst1/tree/main

This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples.

This dataset was used to train Guanaco with QLoRA.

For further information, please see the original dataset.

License: Apache 2.0

Contributors

taeminlee

3 commits

nlpai-lab/openassistant-guanaco-ko

Dataset

11

stars

3

commits

1

linked in READMEs

Jun 1, 2023

updated

README

Dataset Summary

Korean translation of Guanaco via the DeepL API

Note: There are cases where multilingual data has been converted to monolingual data during batch translation to Korean using the API.

Below is Guanaco's README.


This dataset is a subset of the Open Assistant dataset, which you can find here: https://huggingface.co/datasets/OpenAssistant/oasst1/tree/main

This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples.

This dataset was used to train Guanaco with QLoRA.

For further information, please see the original dataset.

License: Apache 2.0

Contributors

taeminlee

3 commits