datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
han-instruct-dataset-v4.0
Dataset Card for Han Instruct Dataset v4.0 🪿🪿🪿🪿
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
Data sources:
Reference desk at Thai wikipedia.
Law from… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v4.0.han-instruct-dataset-v1.0
Dataset Card for "han-instruct-dataset-v1.0"
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
Dataset Summary
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. It collect the instruction following in Thai from many source.
Many question are collect from Reference desk at Thai wikipedia.
Data sources:
Reference desk at Thai wikipedia.
Law from justicechannel.org… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v1.0.han-instruction-dataset
Han Instruction Dataset
Han instruction dataset: Thai instruction dataset
🪿 Han (ห่าน or goose) Instruction Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
The final dataset of han instruction dataset was released!
GitHub: https://github.com/wannaphong/han-instruction-dataset
Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruction-dataset.han-instruct-dataset-v2.0
Dataset Card for Han Instruct Dataset v2.0
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collect all Thai instruct dataset that made by human and our old model. The dataset can use to train Instruction Following model like ChatGPT or other.
Many question are collect from Reference desk at Thai wikipedia.
Data sources:
Reference desk… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v2.0.han-instruct-dataset-v3.0
Dataset Card for Han Instruct Dataset v3.0
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
Many questions are collect from Reference desk at Thai wikipedia.
Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v3.0.Med-REFL-DPO
News
[2025/06/10] We are releasing the Med-REFL dataset, which is split into two subsets: Reasoning Enhancement Data and Reflection Enhancement Data.
Introduction
This is the Direct Preference Optimization (DPO) dataset created by the Med-REFL framework, designed to improve the reasoning and reflection capabilities of Large Language Models in the medical field.
The dataset is constructed using a low-cost, scalable pipeline that leverages a Tree-of-Thought (ToT) approach… See the full description on the dataset page: https://huggingface.co/datasets/HANI-LAB/Med-REFL-DPO.
