LinkSoul/instruction_merge_set

Dataset

123

stars

30

commits

15

linked in READMEs

Oct 25, 2023

updated

README

Dataset Card for "instruction_merge_set"

本数据集由以下数据集构成:

数据(id in the merged set)Hugging face 地址notes
OIG (unified-任务名称) 15khttps://huggingface.co/datasets/laion/OIGOpen Instruction Generalist Dataset
Dolly databricks-dolly-15khttps://huggingface.co/datasets/databricks/databricks-dolly-15kan open-source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories
UltraChathttps://huggingface.co/datasets/stingning/ultrachatmulti-round dialogue data
Camelhttps://huggingface.co/datasets/camel-ai/ai_society25K conversations between two gpt-3.5-turbo agents.
camel (同上)https://github.com/camel-ai/camel
ChatDoctor icliniq-15k HealthCareMagic-200khttps://github.com/Kent0n-Li/ChatDoctor200k real conversations between patients and doctors from HealthCareMagic.com 15k real conversations between patients and doctors from iciniq-10k
Dollyhttps://github.com/databrickslabs/dolly
GPT4ALLhttps://github.com/nomic-ai/gpt4all
GPT-4-LLM comparision_data_b alpaca_gpt4_data_zh comparision_data_a alpaca_gpt4_data 5khttps://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLMEnglish Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. Chinese Instruction-Following Data generated by GPT-4 using Chinese prompts translated from Alpaca by ChatGPT. Comparison Data ranked by GPT-4 to train reward models. Answers on Unnatural Instructions Data from GPT-4 to quantify the gap between GPT-4 and instruction-tuned models at scale.
GuanacoDataset guanaco_chat_all-utf8 guanaco_non_chat-utf8 paper_answers-utf8 general_ans-utf8 general_questions-utf8 paper_questions-utf8 30khttps://huggingface.co/datasets/JosephusCheung/GuanacoDatasetThe dataset for the Guanaco model is designed to enhance the multilingual capabilities and address various linguistic tasks. It builds upon the 175 tasks from the Alpaca model by providing rewrites of seed tasks in different languages and adding new tasks specifically designed for English grammar analysis, natural language understanding, cross-lingual self-awareness, and explicit content recognition. The Paper/General-QA dataset is a collection of questions and answers constructed for AI-generated papers or general texts in English, Chinese, Japanese, and German.
HC3 ALLhttps://huggingface.co/datasets/Hello-SimpleAI/HC3human-ChatGPT comparison datasets
instinwild instinwild_en instinwild_ch 5khttps://huggingface.co/datasets/QingyiSi/Alpaca-CoT/tree/main/instinwildInstruction-Finetuning Dataset Collection (Alpaca-CoT)
Instruct-to-Codehttps://huggingface.co/datasets/Graverman/Instruct-to-Code
ShareGPT90K sg_90k_part2 sg_90k_part1https://huggingface.co/datasets/RyokoAI/ShareGPT52K90,000 conversations scraped via the ShareGPT API before it was shut down. These conversations include both user prompts and responses from OpenAI's ChatGPT.
UltraChat ultrachat_material_release_230412 ultrachat_release_230407https://github.com/thunlp/UltraChat
wealth-alpaca-lora final_dataset_clean 4.3khttps://www.kaggle.com/code/gbhacker23/wealth-alpaca-loracombination of Stanford's Alpaca (https://github.com/tatsu-lab/stanford_alpaca) and FiQA (https://sites.google.com/view/fiqa/) with another 1.3k pairs custom generated using GPT3.5, 有instruction
Alpaca alpaca_data 5khttps://github.com/tatsu-lab/stanford_alpacainstruct-tuning
Baize alpaca_chat_data medical_chat_data quora_chat_data stack_overflow_chat_datahttps://github.com/project-baize/baize-chatbotinstruction-following data we used for fine-tuning the Alpaca model.
botbots Reasoning flight_bookings medical_appointments travel_agency restaurants_mixed real_estate car_dealership home_maintenance, job_interview 'insurance_consultation': 16, 'hotels': 400, 'tech_support': 32, 'car_rentals': 32, 'pet_care': 48, 'restaurants': 200, 'legal_consultation': 16, 'event_tickets': 240, 'fitness_personal_training': 16, 'scientific_problems': 100https://github.com/radi-cho/botbotsA dataset consisting of dialogues between two instances of ChatGPT (gpt-3.5-turbo). The CLI commands and dialogue prompts themselves have been written by GPT-4. The dataset covers a wide range of contexts (questions and answers, arguing and reasoning, task-oriented dialogues) and downstream tasks (e.g., hotel reservations, medical advice).
ChatAlpaca chatalpaca_data_10khttps://github.com/cascip/ChatAlpacaa chat dataset, multi-turn instruction-following conversations.
DERA trainhttps://github.com/curai/curai-research/tree/main/DERAThe following repository contains the open-ended question-answering version of MedQA.
GPTeacher Toolformer-dedupe-only-dataset roleplay-simple-deduped-roleplay-dataset gpt4-instruct-dedupe-only-datasethttps://github.com/teknium1/GPTeacherA collection of modular datasets generated by GPT-4, General-Instruct - Roleplay-Instruct - Code-Instruct - and Toolformer
OpenAGIhttps://github.com/agiresearch/OpenAGI
prestohttps://github.com/google-research-datasets/prestoA Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs

Contributors

shiyemin2

30 commits

LinkSoul/instruction_merge_set

Dataset

123

stars

30

commits

15

linked in READMEs

Oct 25, 2023

updated

README

Dataset Card for "instruction_merge_set"

本数据集由以下数据集构成:

数据(id in the merged set)Hugging face 地址notes
OIG (unified-任务名称) 15khttps://huggingface.co/datasets/laion/OIGOpen Instruction Generalist Dataset
Dolly databricks-dolly-15khttps://huggingface.co/datasets/databricks/databricks-dolly-15kan open-source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories
UltraChathttps://huggingface.co/datasets/stingning/ultrachatmulti-round dialogue data
Camelhttps://huggingface.co/datasets/camel-ai/ai_society25K conversations between two gpt-3.5-turbo agents.
camel (同上)https://github.com/camel-ai/camel
ChatDoctor icliniq-15k HealthCareMagic-200khttps://github.com/Kent0n-Li/ChatDoctor200k real conversations between patients and doctors from HealthCareMagic.com 15k real conversations between patients and doctors from iciniq-10k
Dollyhttps://github.com/databrickslabs/dolly
GPT4ALLhttps://github.com/nomic-ai/gpt4all
GPT-4-LLM comparision_data_b alpaca_gpt4_data_zh comparision_data_a alpaca_gpt4_data 5khttps://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLMEnglish Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. Chinese Instruction-Following Data generated by GPT-4 using Chinese prompts translated from Alpaca by ChatGPT. Comparison Data ranked by GPT-4 to train reward models. Answers on Unnatural Instructions Data from GPT-4 to quantify the gap between GPT-4 and instruction-tuned models at scale.
GuanacoDataset guanaco_chat_all-utf8 guanaco_non_chat-utf8 paper_answers-utf8 general_ans-utf8 general_questions-utf8 paper_questions-utf8 30khttps://huggingface.co/datasets/JosephusCheung/GuanacoDatasetThe dataset for the Guanaco model is designed to enhance the multilingual capabilities and address various linguistic tasks. It builds upon the 175 tasks from the Alpaca model by providing rewrites of seed tasks in different languages and adding new tasks specifically designed for English grammar analysis, natural language understanding, cross-lingual self-awareness, and explicit content recognition. The Paper/General-QA dataset is a collection of questions and answers constructed for AI-generated papers or general texts in English, Chinese, Japanese, and German.
HC3 ALLhttps://huggingface.co/datasets/Hello-SimpleAI/HC3human-ChatGPT comparison datasets
instinwild instinwild_en instinwild_ch 5khttps://huggingface.co/datasets/QingyiSi/Alpaca-CoT/tree/main/instinwildInstruction-Finetuning Dataset Collection (Alpaca-CoT)
Instruct-to-Codehttps://huggingface.co/datasets/Graverman/Instruct-to-Code
ShareGPT90K sg_90k_part2 sg_90k_part1https://huggingface.co/datasets/RyokoAI/ShareGPT52K90,000 conversations scraped via the ShareGPT API before it was shut down. These conversations include both user prompts and responses from OpenAI's ChatGPT.
UltraChat ultrachat_material_release_230412 ultrachat_release_230407https://github.com/thunlp/UltraChat
wealth-alpaca-lora final_dataset_clean 4.3khttps://www.kaggle.com/code/gbhacker23/wealth-alpaca-loracombination of Stanford's Alpaca (https://github.com/tatsu-lab/stanford_alpaca) and FiQA (https://sites.google.com/view/fiqa/) with another 1.3k pairs custom generated using GPT3.5, 有instruction
Alpaca alpaca_data 5khttps://github.com/tatsu-lab/stanford_alpacainstruct-tuning
Baize alpaca_chat_data medical_chat_data quora_chat_data stack_overflow_chat_datahttps://github.com/project-baize/baize-chatbotinstruction-following data we used for fine-tuning the Alpaca model.
botbots Reasoning flight_bookings medical_appointments travel_agency restaurants_mixed real_estate car_dealership home_maintenance, job_interview 'insurance_consultation': 16, 'hotels': 400, 'tech_support': 32, 'car_rentals': 32, 'pet_care': 48, 'restaurants': 200, 'legal_consultation': 16, 'event_tickets': 240, 'fitness_personal_training': 16, 'scientific_problems': 100https://github.com/radi-cho/botbotsA dataset consisting of dialogues between two instances of ChatGPT (gpt-3.5-turbo). The CLI commands and dialogue prompts themselves have been written by GPT-4. The dataset covers a wide range of contexts (questions and answers, arguing and reasoning, task-oriented dialogues) and downstream tasks (e.g., hotel reservations, medical advice).
ChatAlpaca chatalpaca_data_10khttps://github.com/cascip/ChatAlpacaa chat dataset, multi-turn instruction-following conversations.
DERA trainhttps://github.com/curai/curai-research/tree/main/DERAThe following repository contains the open-ended question-answering version of MedQA.
GPTeacher Toolformer-dedupe-only-dataset roleplay-simple-deduped-roleplay-dataset gpt4-instruct-dedupe-only-datasethttps://github.com/teknium1/GPTeacherA collection of modular datasets generated by GPT-4, General-Instruct - Roleplay-Instruct - Code-Instruct - and Toolformer
OpenAGIhttps://github.com/agiresearch/OpenAGI
prestohttps://github.com/google-research-datasets/prestoA Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs

Contributors

shiyemin2

30 commits