Hercules-v4.0 is an extensive and diverse dataset that combines various domains to create a powerful tool for training artificial intelligence models. The data sources include conversations, coding examples, scientific explanations, and more. The dataset is sourced from multiple high-quality repositories, each contributing to the robustness of Hercules-v4.0 in different knowledge domains.
OpenOrca/SlimOrcaEvol Instruct 70K & 140Kteknium/GPT4-LLM-Cleanedjondurbin/airoboros-3.2AlekseyKorshuk/camel-chatmlCollectiveCognition/chats-data-2023-09-22glaiveai/glaive-code-assistantglaiveai/glaive-function-calling-v2garage-bAInd/Open-Platypusmeta-math/MetaMathQAmicrosoft/orca-math-word-problems-200kGPTeacher roleplay datasetsBI55/MedTextabacusai/SystemChatm-a-p/Code-Feedbacktotally-not-an-llm/EverythingLM-data-V3Locutusque/arc-cotFuseAI/FuseChat-MixtureLDJnr/Pure-Doveteknium/trismegistus-projectVezora/Tested-22k-Python-AlpacaCrystalcareai/alpaca-gpt4-COTgrimulkan/theory-of-mindCollectiveCognition/chats-data-2023-09-27CollectiveCognition/chats-data-2023-10-16NobodyExistsOnTheInternet/sharegptPIPPAsablo/oasst2_curatedThe dataset amalgamates text from various domains, including structured and unstructured data. It contains dialogues, instructional texts, scientific explanations, coding tasks, and more.
Hercules-v4.0 is designed for training and evaluating AI models capable of handling complex tasks across multiple domains. It is suitable for researchers and developers in academia and industry working on advanced conversational agents, instruction-following models, and knowledge-intensive applications.
The data was collected from reputable sources with an emphasis on diversity and quality. It is expected to be relatively clean but may require additional preprocessing for specific tasks.
Hercules-v4.0 contains X-rated content. Users are solely responsible for the use of the dataset and must ensure that their use complies with all applicable laws and regulations. The dataset maintainers are not responsible for the misuse of the dataset.
By using the Hercules-v4.0 dataset, users agree to the following:
Please make sure to read the license for more information.
6 commits
Hercules-v4.0 is an extensive and diverse dataset that combines various domains to create a powerful tool for training artificial intelligence models. The data sources include conversations, coding examples, scientific explanations, and more. The dataset is sourced from multiple high-quality repositories, each contributing to the robustness of Hercules-v4.0 in different knowledge domains.
OpenOrca/SlimOrcaEvol Instruct 70K & 140Kteknium/GPT4-LLM-Cleanedjondurbin/airoboros-3.2AlekseyKorshuk/camel-chatmlCollectiveCognition/chats-data-2023-09-22glaiveai/glaive-code-assistantglaiveai/glaive-function-calling-v2garage-bAInd/Open-Platypusmeta-math/MetaMathQAmicrosoft/orca-math-word-problems-200kGPTeacher roleplay datasetsBI55/MedTextabacusai/SystemChatm-a-p/Code-Feedbacktotally-not-an-llm/EverythingLM-data-V3Locutusque/arc-cotFuseAI/FuseChat-MixtureLDJnr/Pure-Doveteknium/trismegistus-projectVezora/Tested-22k-Python-AlpacaCrystalcareai/alpaca-gpt4-COTgrimulkan/theory-of-mindCollectiveCognition/chats-data-2023-09-27CollectiveCognition/chats-data-2023-10-16NobodyExistsOnTheInternet/sharegptPIPPAsablo/oasst2_curatedThe dataset amalgamates text from various domains, including structured and unstructured data. It contains dialogues, instructional texts, scientific explanations, coding tasks, and more.
Hercules-v4.0 is designed for training and evaluating AI models capable of handling complex tasks across multiple domains. It is suitable for researchers and developers in academia and industry working on advanced conversational agents, instruction-following models, and knowledge-intensive applications.
The data was collected from reputable sources with an emphasis on diversity and quality. It is expected to be relatively clean but may require additional preprocessing for specific tasks.
Hercules-v4.0 contains X-rated content. Users are solely responsible for the use of the dataset and must ensure that their use complies with all applicable laws and regulations. The dataset maintainers are not responsible for the misuse of the dataset.
By using the Hercules-v4.0 dataset, users agree to the following:
Please make sure to read the license for more information.
6 commits