datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hercules-v3.0
Hercules-v3.0
Dataset Name: Hercules-v3.0
Version: 3.0
Release Date: 2024-2-14
Number of Examples: 1,637,895
Domains: Math, Science, Biology, Physics, Instruction Following, Conversation, Computer Science, Roleplay, and more
Languages: Mostly English, but others can be detected.
Task Types: Question Answering, Conversational Modeling, Instruction Following, Code Generation, Roleplay
Data Source Description
Hercules-v3.0 is an extensive and diverse dataset that… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/Hercules-v3.0.hyperion-v3.0Hyperion-3.0 has significantly improved performance over its predecessors.
"I found that having more code datasets than general purpose datasets ironically decreases performance in both coding and general tasks."
Data sources:
OpenOrca/SlimOrca
cognitivecomputations/dolphin (300k examples)
microsoft/orca-math-word-problems-200k (60k examples)
glaiveai/glaive-code-assistant
Vezora/Tested-22k-Python-Alpaca
Unnatural Instructions
BI55/MedText
LDJnr/Pure-Dove
Various domain-specific datasets by… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hyperion-v3.0.han-instruct-dataset-v3.0
Dataset Card for Han Instruct Dataset v3.0
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
Many questions are collect from Reference desk at Thai wikipedia.
Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v3.0.
