This dataset is released as part of [SWEET-RL: Training Multi-Turn LLM Agents on
60
9 commits
2 linked in READMEs
updated Mar 20, 2025
This dataset is released as part of SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks research project.
Please refer to our project materials here for training and evaluation details.
If you use data, model, or code from this work, please cite with the following BibTex entry:
@misc{zhou2025sweetrltrainingmultiturnllm,
title={SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks},
author={Yifei Zhou and Song Jiang and Yuandong Tian and Jason Weston and Sergey Levine and Sainbayar Sukhbaatar and Xian Li},
year={2025},
eprint={2503.15478},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.15478},
}
The data is licensed under CC-by-NC. This data is an output from Llama 3.1, and subject to the Llama 3.1 license (https://huggingface.co/meta-llama/Llama-3.1-8B/blob/main/LICENSE). Use of the data to train, fine tune, or otherwise improve an AI model, which is distributed or made available, shall also include "Llama" at the beginning of any such AI model name.
9 commits
This dataset is released as part of [SWEET-RL: Training Multi-Turn LLM Agents on
60
9 commits
2 linked in READMEs
updated Mar 20, 2025
This dataset is released as part of SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks research project.
Please refer to our project materials here for training and evaluation details.
If you use data, model, or code from this work, please cite with the following BibTex entry:
@misc{zhou2025sweetrltrainingmultiturnllm,
title={SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks},
author={Yifei Zhou and Song Jiang and Yuandong Tian and Jason Weston and Sergey Levine and Sainbayar Sukhbaatar and Xian Li},
year={2025},
eprint={2503.15478},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.15478},
}
The data is licensed under CC-by-NC. This data is an output from Llama 3.1, and subject to the Llama 3.1 license (https://huggingface.co/meta-llama/Llama-3.1-8B/blob/main/LICENSE). Use of the data to train, fine tune, or otherwise improve an AI model, which is distributed or made available, shall also include "Llama" at the beginning of any such AI model name.
9 commits