This dataset is a collection of approximately 52,00090,000 conversations scraped via the ShareGPT API before it was shut down.
These conversations include both user prompts and responses from OpenAI's ChatGPT.
This repository now contains the new 90K conversations version. The previous 52K may
be found in the old/ directory.
This dataset is expected to primarily consist of messages in English and other Western languages.
{
"id": "2wtLdZm",
"conversations": [
{
"from": "human",
"value": "Are you familiar with the SAS programming language?"
},
{
"from": "gpt",
"value": "<div class=\"markdown prose w-full break-words dark:prose-invert light\"><p>Yes, I am familiar with the SAS programming language. SAS is a statistical software suite that is widely used in data analysis, data management, and business intelligence applications. It has its own programming language, which is used to manipulate and analyze data, generate reports, and automate data-driven workflows. The language is known for its ease of use and powerful capabilities, and it is commonly used in industries such as healthcare, finance, and government.</p></div>"
},
{
"from": "human",
"value": "Can you write a SAS proc that imports a dataset from a CSV saved on my local machine?"
}
]
}
N/A
This is a decently large dataset of realistic human-AI conversations which I believe should be released to the research community.
This data was collected using the ShareGPT API.
ShareGPT users and OpenAI ChatGPT.
N/A
N/A
This dataset may contain personal information, if ShareGPT users were sending such information to ChatGPT. ChatGPT warns users not to submit personal information to it, however, so without further evaluation, we believe that this dataset should contain little or no personal information.
This dataset may be used to train models that are competitive with OpenAI's ChatGPT. Please filter this dataset first, as it may contain canned responses, raw HTML, and other undesirable information.
This dataset exhibits all the biases of OpenAI's ChatGPT models (GPT-3.5 and GPT-4) as well as the biases of the users who uploaded the conversations.
N/A
None.
CC0: No Rights Reserved.
The output of machine learning algorithms is uncopyrightable in the United States and other jurisdictions. Additionally, the OpenAI terms of service do not apply to this dataset as users of this dataset are not accessing the OpenAI service.
TODO
These conversations were allegedly scraped by an anonymous user on 4chan.
The 90K version was sourced from this post. Thanks, anon!
This dataset is a collection of approximately 52,00090,000 conversations scraped via the ShareGPT API before it was shut down.
These conversations include both user prompts and responses from OpenAI's ChatGPT.
This repository now contains the new 90K conversations version. The previous 52K may
be found in the old/ directory.
This dataset is expected to primarily consist of messages in English and other Western languages.
{
"id": "2wtLdZm",
"conversations": [
{
"from": "human",
"value": "Are you familiar with the SAS programming language?"
},
{
"from": "gpt",
"value": "<div class=\"markdown prose w-full break-words dark:prose-invert light\"><p>Yes, I am familiar with the SAS programming language. SAS is a statistical software suite that is widely used in data analysis, data management, and business intelligence applications. It has its own programming language, which is used to manipulate and analyze data, generate reports, and automate data-driven workflows. The language is known for its ease of use and powerful capabilities, and it is commonly used in industries such as healthcare, finance, and government.</p></div>"
},
{
"from": "human",
"value": "Can you write a SAS proc that imports a dataset from a CSV saved on my local machine?"
}
]
}
N/A
This is a decently large dataset of realistic human-AI conversations which I believe should be released to the research community.
This data was collected using the ShareGPT API.
ShareGPT users and OpenAI ChatGPT.
N/A
N/A
This dataset may contain personal information, if ShareGPT users were sending such information to ChatGPT. ChatGPT warns users not to submit personal information to it, however, so without further evaluation, we believe that this dataset should contain little or no personal information.
This dataset may be used to train models that are competitive with OpenAI's ChatGPT. Please filter this dataset first, as it may contain canned responses, raw HTML, and other undesirable information.
This dataset exhibits all the biases of OpenAI's ChatGPT models (GPT-3.5 and GPT-4) as well as the biases of the users who uploaded the conversations.
N/A
None.
CC0: No Rights Reserved.
The output of machine learning algorithms is uncopyrightable in the United States and other jurisdictions. Additionally, the OpenAI terms of service do not apply to this dataset as users of this dataset are not accessing the OpenAI service.
TODO
These conversations were allegedly scraped by an anonymous user on 4chan.
The 90K version was sourced from this post. Thanks, anon!