datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lexitron2_prompt_finetune
Lexitron 2.0 Prompt Finetuning Dataset
This dataset is derived from Lexitron 2.0, a Thai-English dictionary developed by NECTEC. It has been processed and formatted for prompt finetuning tasks. The original dataset is from: https://opend-portal.nectec.or.th/dataset/lexitron-2-0
Maintainer
Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Dataset Description
The dataset consists of two main files:
lexitron2_telex_finetune.qwen2.txt - Thai to English lexicon entries… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/lexitron2_prompt_finetune.thai-qa-rag-answer-dataset
Thai QA RAG Answer Synthesis Dataset
Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined
Rows: 9999 rows.
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.thai-wiki-summary-dataset
Thai Wiki Summary Dataset
Rows: 3,000 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"หน่วยพื้นฐานในการแบ่งเขตแดนในโปแลนด์คือ เทศบาล (กมินา) เมืองก็เป็นเทศบาลด้วยเช่นกัน ทว่ามีตราตั้งให้เป็นเมือง ทั้งเมืองและเทศบาลปกครองโดยนายกเทศมนตรี ทว่าในเทศบาล นายกเทศมนตรีเรียกว่าโวกต์ ( วอยต์ในภาษาโปแลนด์) ส่วนในเมืองเรียกว่าเบอร์มิสตร์ ในเมืองใหญ่ ๆ บางเมืองมีความรับผิดชอบและอำนาจพิเศษ… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-wiki-summary-dataset.thai-qa-multiturn-answer-dataset
Thai QA Multiturns Answer Synthesis Dataset
Rows: 11,992 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}]", "output": "สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ"}
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}, {\"assistant\": \"สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ\"}, {\"human\":… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-multiturn-answer-dataset.kanitakorn-deepseek-v45-v44-openthaieval-bridge-mix
Kanitakorn v45 v44 + OpenThaiEval bridge mix
Train-ready SFT mix for a non-Thai-base DeepSeek/Qwen-style <=14B candidate.
Delta from v44:
reuses all v44 normalized MCQ replay and identity rows unchanged
adds locked, verified synthetic OpenThaiEval-style rows
normalizes those bridge rows so explanation precedes the final answer
keeps single-model training only; no BoN, self-consistency, routing, ensemble, or benchmark-label leakage
Rows: 1240
Bridge rows: 70
Train SHA256:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v45-v44-openthaieval-bridge-mix.OpenThai-NER-Corpus
Thai Named Entity Recognition (NER) Corpus
A comprehensive corpus for Thai Named Entity Recognition tasks with 6,748 annotated sentences across 160 different domains.
Overview
This dataset contains Thai text samples annotated with named entity labels for training and evaluating NER models. The corpus covers a wide variety of domains including government, finance, legal, healthcare, education, and more.
Dataset Statistics
Total Samples: 10,345 annotated… See the full description on the dataset page: https://huggingface.co/datasets/JonusNattapong/OpenThai-NER-Corpus.
