datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.thai-ocr-evaluation
Thai OCR Evaluation Dataset
Dataset Description
The Thai OCR Evaluation Dataset is designed for evaluating Optical Character Recognition (OCR) models across various domains. It includes images and textual data derived from various open-source websites.
This dataset aims to provide a comprehensive evaluation resource for researchers and developers working on OCR systems, particularly in Thai language processing.
Data Fields
Each sample in the dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-ocr-evaluation.openthaieval
OpenThaiEval: Comprehensive Thai Language Evaluation Benchmark
Overview
OpenThaiEval is a Thai language evaluation benchmark containing 1,232 questions
across 17 exam types, ranging from national standardized tests to international
benchmarks and professional certification exams.
Not all 1,232 rows measure Thai, and not all of them are ours to license. Both
points are set out below rather than left for a reader to discover, because both
change what… See the full description on the dataset page: https://huggingface.co/datasets/iapp/openthaieval.thai-investment-consultant-licensing-exams
Thai Public Investment Consultant (IC) Exams Dataset
Overview
This dataset comprises a collection of exam questions and answers from the Thai Public Investment Consultant (IC) Examinations. It's a valuable resource for developing and evaluating question-answering systems in the finance sector.
Dataset Source
The Stock Exchange of Thailand (SET)
Maintainer
Dr. Kobkrit Viriyayudhakorn
Email: kobkrit@iapp.co.th
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-investment-consultant-licensing-exams.lexitron2_prompt_finetune
Lexitron 2.0 Prompt Finetuning Dataset
This dataset is derived from Lexitron 2.0, a Thai-English dictionary developed by NECTEC. It has been processed and formatted for prompt finetuning tasks. The original dataset is from: https://opend-portal.nectec.or.th/dataset/lexitron-2-0
Maintainer
Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Dataset Description
The dataset consists of two main files:
lexitron2_telex_finetune.qwen2.txt - Thai to English lexicon entries… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/lexitron2_prompt_finetune.thai-qa-rag-answer-dataset
Thai QA RAG Answer Synthesis Dataset
Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined
Rows: 9999 rows.
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.thai-wiki-summary-dataset
Thai Wiki Summary Dataset
Rows: 3,000 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"หน่วยพื้นฐานในการแบ่งเขตแดนในโปแลนด์คือ เทศบาล (กมินา) เมืองก็เป็นเทศบาลด้วยเช่นกัน ทว่ามีตราตั้งให้เป็นเมือง ทั้งเมืองและเทศบาลปกครองโดยนายกเทศมนตรี ทว่าในเทศบาล นายกเทศมนตรีเรียกว่าโวกต์ ( วอยต์ในภาษาโปแลนด์) ส่วนในเมืองเรียกว่าเบอร์มิสตร์ ในเมืองใหญ่ ๆ บางเมืองมีความรับผิดชอบและอำนาจพิเศษ… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-wiki-summary-dataset.thai-qa-multiturn-answer-dataset
Thai QA Multiturns Answer Synthesis Dataset
Rows: 11,992 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}]", "output": "สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ"}
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}, {\"assistant\": \"สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ\"}, {\"human\":… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-multiturn-answer-dataset.kanitakorn-deepseek-v45-v44-openthaieval-bridge-mix
Kanitakorn v45 v44 + OpenThaiEval bridge mix
Train-ready SFT mix for a non-Thai-base DeepSeek/Qwen-style <=14B candidate.
Delta from v44:
reuses all v44 normalized MCQ replay and identity rows unchanged
adds locked, verified synthetic OpenThaiEval-style rows
normalizes those bridge rows so explanation precedes the final answer
keeps single-model training only; no BoN, self-consistency, routing, ensemble, or benchmark-label leakage
Rows: 1240
Bridge rows: 70
Train SHA256:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v45-v44-openthaieval-bridge-mix.OpenThai-NER-Corpus
Thai Named Entity Recognition (NER) Corpus
A comprehensive corpus for Thai Named Entity Recognition tasks with 6,748 annotated sentences across 160 different domains.
Overview
This dataset contains Thai text samples annotated with named entity labels for training and evaluating NER models. The corpus covers a wide variety of domains including government, finance, legal, healthcare, education, and more.
Dataset Statistics
Total Samples: 10,345 annotated… See the full description on the dataset page: https://huggingface.co/datasets/JonusNattapong/OpenThai-NER-Corpus.
