An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
| name | # samples |
|---|---|
| en_helpfulness.json | 419049 |
| en_honesty.json | 112580 |
| en_harmlessness.json | 38873 |
| zh_helpfulness.json | 447750 |
| zh_honesty.json | 142885 |
7 commits
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
| name | # samples |
|---|---|
| en_helpfulness.json | 419049 |
| en_honesty.json | 112580 |
| en_harmlessness.json | 38873 |
| zh_helpfulness.json | 447750 |
| zh_honesty.json | 142885 |
7 commits