datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.thai_exam-reformattedReformatted version of scb10x/thai_exam
Additional Changes:
Fix math incorrect answer
ถ้า \log_{\frac{1}{4}} 256 + \frac{2\log 625}{\log 5} = 3^a เมื่อ a เป็นจำนวนจริง แล้วคำตอบของ a เท่ากับเท่าใด?
a. \log_{3} 2
b. \log_{3} 4
c. \log_{3} \frac{33}{4}
d. \log_{3} 10
e. \log_{3} 12
# Original answer: d (\log_{3} 10)
# Correction : b (\log_{3} 4)
thai-exam-seacrowdthai-culturax-clean-dataset
Thai CulturaX Clean dataset
The data is sourced from the Thai subset of CulturaX dataset, which itself is sourced from mC4 and four OSCAR corpora.
It has about 8,748,575,684 words (without whitespace) and 16,768,585 lines (97 GB).
It was filtered content promoting gambling, adult content, and narcotics.
GitHub for clean: https://github.com/wannaphong/thai-filter-website
Considerations for Using the Data
This dataset is the cleaned version of the CulturaX datasets… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-culturax-clean-dataset.amazon-2023-thai-1m
Amazon 2023 Thai 1M / ชุดข้อมูลสินค้า Amazon ภาษาไทย 1 ล้านรายการ
ไทย | English below
ชุดข้อมูลสินค้าอีคอมเมิร์ซภาษาไทย 1,000,000 รายการ แปลจากชุดข้อมูล Amazon Reviews '23 Extension ของ Google ด้วยโมเดล typhoon-translate-4b เพื่อใช้ทดสอบ/สาธิตระบบค้นหาเชิงความหมาย (semantic search) ภาษาไทย
แหล่งที่มา / Source & Attribution ⚠️ สำคัญ
ชุดข้อมูลนี้เป็นงานแปล (derivative work) จาก:
ชุดข้อมูลต้นทาง
google/extended_amazon_2023_dataset (Amazon Reviews '23… See the full description on the dataset page: https://huggingface.co/datasets/pnpkke/amazon-2023-thai-1m.ThaiExamFinetune Multiple-Choice Question Dataset in Thai
Details:
This dataset is designed for fine-tuning Thai language models, focusing on the Chain-of-Thought (COT) process, which aids in analyzing questions and deriving correct answers step by step. The dataset consists of multiple-choice questions divided into five categories:
O-NET: Ordinary National Educational Test
IC: Investment Consultant
TGAT: Thai General Aptitude Test
TPAT: Thai Professional Aptitude Test… See the full description on the dataset page: https://huggingface.co/datasets/EIRTHAIMED/ThaiExam.ThaiQA-v1
ThaiQA v1
ThaiQA v1 is a Thai Synthetic QA dataset. It was created from synthetic method using open source LLM in Thai language.
We used Nvidia Nemotron 4 (340B) to create this dataset.
Topics:
Technology and Gadgets 100
Travel and Tourism 91
Food and Cooking 99
Sports and Fitness 50
Arts and Entertainment 24
Home and Garden 72
Fashion and Beauty 99
Science and Nature 100
History and Culture 91
Education and Learning 99
Pets and Animals 83
Relationships and Family 78
Personal… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/ThaiQA-v1.kanitakorn-v23-thaiexam-clean-20260614
Qwen v23 ThaiExam Clean Mix
Audited fallback mix for Kanitakorn. It avoids v20/v21 replay, avoids v13+v17 double replay, uses v1 repair once, includes all normalized worker v2, adds clean worker v3, and keeps small IF/math/identity retainers.
Validation
Records: 8,214
inspect_generated_jsonl: 3,214 source rows valid, 0 invalid, 0 duplicate prompts
inspect_sft_mix: 0 role errors, 0 empty errors
Contamination scan: 0 issues with 8,608 benchmark texts loaded… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-v23-thaiexam-clean-20260614.KhanomTanLLM-pretrained-dataset-thai-subset
KhanomTanLLM pretrained dataset (Thai subset)
This daataset collect all raw text for pretraining LLM. (Thai subset)
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0
pythainlp/thai-tnhc2-books
pythainlp/thai-constitution-corpus
pythainlp/thai-it-books
pythainlp/prd_news_3011202
pythainlp/thailand-policy-statements
pythainlp/thai-cc-license
pythainlp/blognone_news
pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.seed-free-synthetic-instruct-thai-v1
Seed-Free Synthetic Instruct Thai v1 (F+C+D+)
This dataset is part of the research paper "Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai" submitted to ACL SRW 2024. It represents the best-performing synthetic dataset (F+C+D+) generated using our novel seed-free framework for low-resource languages, specifically Thai.
Dataset Details
Size: 5,000 instructions
Language: Thai
Task: Instruction-tuning for Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/parinzee/seed-free-synthetic-instruct-thai-v1.thai_instruction_sftMedical-o1-Reasoning-SFT-Thai
Medical-GPT-Reasoning-Thai
Dataset Summary
This dataset contains medical Q&A data in JSON format, designed for fine-tuning AI models in medical reasoning and response generation.representing a medical question, complex chain-of-thought reasoning, and a concise response. All content is in Thai language.
The dataset is derived from a larger medical Q&A collection and has been processed to ensure JSON validity, with multi-line objects combined into single valid entries.… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-Reasoning-SFT-Thai.thai-qa-rag-answer-dataset
Thai QA RAG Answer Synthesis Dataset
Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined
Rows: 9999 rows.
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.Thai-R1-Distill-SFT
Thai R1 Distill SFT
Thai Reasoning Dataset for Supervised Finetuning
Translated by iApp Technology
thai-wiki-summary-dataset
Thai Wiki Summary Dataset
Rows: 3,000 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"หน่วยพื้นฐานในการแบ่งเขตแดนในโปแลนด์คือ เทศบาล (กมินา) เมืองก็เป็นเทศบาลด้วยเช่นกัน ทว่ามีตราตั้งให้เป็นเมือง ทั้งเมืองและเทศบาลปกครองโดยนายกเทศมนตรี ทว่าในเทศบาล นายกเทศมนตรีเรียกว่าโวกต์ ( วอยต์ในภาษาโปแลนด์) ส่วนในเมืองเรียกว่าเบอร์มิสตร์ ในเมืองใหญ่ ๆ บางเมืองมีความรับผิดชอบและอำนาจพิเศษ… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-wiki-summary-dataset.Wisesight-Sentiment-Thai
Dataset Card for Zombitx64 Sentiment Corpus Thai
This dataset card describes the "Zombitx64 Sentiment Corpus Thai," a large-scale, manually curated corpus for Thai sentiment analysis and token classification.
Dataset Details
Dataset Description
The Zombitx64 Sentiment Corpus Thai is a collection of Thai-language social media comments and posts, annotated for sentiment at the sentence or token level. The dataset covers a wide range of topics and emotional… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Wisesight-Sentiment-Thai.thai-place-name-romanisation-benchmark
Thai romanisation, scored on the words place names are made of
A romaniser can do well on dictionary words and still misread a road sign.
This dataset scores one engine on both, so the difference is a number rather
than an impression.
Every distinct pure-Thai word in
pythainlp/thai-romanization-dataset
is romanised by translit.py,
the rule-based RTGS engine behind motdang.net, twice:
with its curated exception lexicon
and with the rules alone.
Results
slice… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-place-name-romanisation-benchmark.thai-qa-multiturn-answer-dataset
Thai QA Multiturns Answer Synthesis Dataset
Rows: 11,992 rows (Cleaned)
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}]", "output": "สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ"}
{"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}, {\"assistant\": \"สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ\"}, {\"human\":… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-multiturn-answer-dataset.mt-bench-thai
MT-Bench Thai
MT-Bench Thai is a dataset for multi-turn benchmarking that covers 9 categories.
Writing
Roleplay
Extraction
Reasoning
Math
Coding
STEM
Social Science
Knowledge III
We introduce the final category, Knowledge III, which evaluates understanding of Thai cultural context.
Dataset Loading
from datasets import load_dataset
ds = load_dataset("ThaiLLM-Leaderboard/mt-bench-thai")
print(ds)
output
DatasetDict({
train: Dataset({
features: ['question_id'… See the full description on the dataset page: https://huggingface.co/datasets/ThaiLLM-Leaderboard/mt-bench-thai.thai-traffic-law-qathai-rtgs-romanisation-lexicon
Thai RTGS romanisation — the exception lexicon
694 Thai forms whose romanisation cannot be derived from the
spelling, each with the reading a rule-based romaniser should use instead.
This is the companion to a rule engine, not a replacement for one. The rules
handle the regular cases; this names the irregulars. Measured: it takes the
romaniser behind motdang.net from 77.51% to 79.62%
agreement on 153,948 Thai words, and from 83.45% to 84.22% on the 21,425 of
them that appear in… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-rtgs-romanisation-lexicon.thai-trade-words-chiang-mai
คำบนป้าย — Thai trade words of Chiang Mai and Chiang Rai
1,782 Thai trade terms taken from the tags on business listings in
Chiang Mai and Chiang Rai — the words the city uses for what a shop does —
each with an RTGS reading and an English gloss. Plus 403
administrative place names (tambon, amphoe, city, province) with their
readings.
ซ่อมมอเตอร์ไซค์ Som Motoesai motorcycle repair 191 places
ตู้น้ำดื่มหยอดเหรียญ Tunam Duem Yotrian coin-op drinking-water… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-trade-words-chiang-mai.gs8k_thai_r1_example
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Think in English, Context Thai for improve thai question
thai-land-tax-full-triplets
Thai Land & Buildings Tax — Full Triplet Dataset
Dataset Name: monoboard/thai-land-tax-full-triplets
Language: Thai (th)
Tasks: Legal Retrieval • RAG • Contrastive Learning • Triplet Loss • Embedding Training
Overview
This dataset provides a legally verified retrieval corpus for training Thai-language retrieval models under the Land and Buildings Tax Act (B.E. 2562) and related regulations.
Each example includes:
A legal question (query)
One or more oracle passages (pos)… See the full description on the dataset page: https://huggingface.co/datasets/monoboard/thai-land-tax-full-triplets.thai-multitask-starter
Thai Multitask 9.6K
ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป
แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output
ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ
ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline
แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง
จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.sql-create-context-thai
Overview
This dataset builds from sql-create-context.
@misc{b-mc2_2023_sql-create-context,
title = {sql-create-context Dataset},
author = {b-mc2},
year = {2023},
url = {https://huggingface.co/datasets/b-mc2/sql-create-context},
note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.},
}
thai-name-spell-v3
