datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.MMMU-Thai
MMMU Thai (MMMU Benchmark Translated to Thai)
MMMU Thai is a dataset for evaluating multimodal models on massive multi-discipline tasks requiring college-level knowledge and deliberate reasoning. This dataset is translated from MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) into Thai.
Dataset Details
MMMU Thai consists of 11,500 meticulously collected multimodal questions from college exams, quizzes, and textbooks… See the full description on the dataset page: https://huggingface.co/datasets/iapp/MMMU-Thai.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.thai_exam-reformattedReformatted version of scb10x/thai_exam
Additional Changes:
Fix math incorrect answer
ถ้า \log_{\frac{1}{4}} 256 + \frac{2\log 625}{\log 5} = 3^a เมื่อ a เป็นจำนวนจริง แล้วคำตอบของ a เท่ากับเท่าใด?
a. \log_{3} 2
b. \log_{3} 4
c. \log_{3} \frac{33}{4}
d. \log_{3} 10
e. \log_{3} 12
# Original answer: d (\log_{3} 10)
# Correction : b (\log_{3} 4)
onet-m6-social
Summary
This is a question-answer dataset for the Grade 12 (M6) Social subject of the Thailand Ordinary National Educational Test (ONET).
The dataset was human-extracted by my team from the official release of publicly available exams National Institute of Educational Testing Service during the years 2016-2022.
The exam consists of 510 multiple-choice questions with corresponding answer keys.
It is important to note that only two questions, Q71 and Q85, from the year 2018, require… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/onet-m6-social.thaiqa_squad`thaiqa_squad` is an open-domain, extractive question answering dataset (4,000 questions in `train` and 74 questions in `dev`) in
[SQuAD](https://rajpurkar.github.io/SQuAD-explorer/) format, originally created by [NECTEC](https://www.nectec.or.th/en/) from
Wikipedia articles and adapted to [SQuAD](https://rajpurkar.github.io/SQuAD-explorer/) format by [PyThaiNLP](https://github.com/PyThaiNLP/).that-backpacker-article-corpus
That Backpacker Article Corpus
This dataset contains a structured corpus of long-form travel articles published on ThatBackpacker.com, authored primarily by Audrey Bergner as part of the Samuel & Audrey Media Network.
The corpus includes 323 article records covering destination guides, multi-day itineraries, hiking, food travel, cultural experiences, city guides, transportation, accommodations, and practical travel planning.
It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/that-backpacker-article-corpus.thai_buddhist_studies_exam
Thai Buddhist Studies Examination (Nak Tham)
This repository contains multiple-choice questions from the Thai Buddhist Studies
(Nak Tham) examination (2020, 2022, 2023). This dataset can be used for a benchmark for evaluating Large Language Models'
understanding of Thai Buddhist concepts and teachings.
Dataset Statistics
Year
Number of Multiple Choice Questions
2020
1,350
2022
1,400
2023
1,350
Phra Udom thought on the exam: We have reviewed the Nak… See the full description on the dataset page: https://huggingface.co/datasets/biodatlab/thai_buddhist_studies_exam.alpaca-cleaned-52k-th
Summary
This is a Thai 🇹🇭-instructed dataset translated from cleaned version of the original Alpaca Dataset released by Stanford using Google Cloud Translation, contain 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine.
This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.
The following issues have been identified in the original release and fixed in this… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/alpaca-cleaned-52k-th.ThaiExamFinetune Multiple-Choice Question Dataset in Thai
Details:
This dataset is designed for fine-tuning Thai language models, focusing on the Chain-of-Thought (COT) process, which aids in analyzing questions and deriving correct answers step by step. The dataset consists of multiple-choice questions divided into five categories:
O-NET: Ordinary National Educational Test
IC: Investment Consultant
TGAT: Thai General Aptitude Test
TPAT: Thai Professional Aptitude Test… See the full description on the dataset page: https://huggingface.co/datasets/EIRTHAIMED/ThaiExam.marketing-benchmark-of-more-than-10-ai-models
Marketing Benchmark of 10+ AI Models
A 5,000-question benchmark for evaluating LLMs across six dimensions of modern
marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical
Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas.
Every question is independently authored by the AdsGPT Marketing Bench team.
Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended
and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.thai-investment-consultant-licensing-exams
Thai Public Investment Consultant (IC) Exams Dataset
Overview
This dataset comprises a collection of exam questions and answers from the Thai Public Investment Consultant (IC) Examinations. It's a valuable resource for developing and evaluating question-answering systems in the finance sector.
Dataset Source
The Stock Exchange of Thailand (SET)
Maintainer
Dr. Kobkrit Viriyayudhakorn
Email: kobkrit@iapp.co.th
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-investment-consultant-licensing-exams.open-math-dataset
Dataset Description
Open Math Dataset is an open mathematics corpus designed for mathematical AI, reasoning, education, and research.
The project is being developed from Sri Lanka with the goal of creating an internationally useful mathematics dataset for developers, researchers, educators, and AI systems.
Mathematics Corpus
The dataset is designed to contain structured mathematical problems and solutions across different mathematical domains and education levels.… See the full description on the dataset page: https://huggingface.co/datasets/tharustack/open-math-dataset.ThaiQA-v1
ThaiQA v1
ThaiQA v1 is a Thai Synthetic QA dataset. It was created from synthetic method using open source LLM in Thai language.
We used Nvidia Nemotron 4 (340B) to create this dataset.
Topics:
Technology and Gadgets 100
Travel and Tourism 91
Food and Cooking 99
Sports and Fitness 50
Arts and Entertainment 24
Home and Garden 72
Fashion and Beauty 99
Science and Nature 100
History and Culture 91
Education and Learning 99
Pets and Animals 83
Relationships and Family 78
Personal… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/ThaiQA-v1.ThaiSafetyBench
ThaiSafetyBench
⚠️ Warning: This dataset contains harmful and toxic language. It is intended for academic purposes only.
[ArXiv Paper] [Github] [Hugging Face Leaderboard 🤗]
The ThaiSafetyBench dataset comprises 1,889 malicious Thai-language prompts across various categories. In addition to translated malicious prompts, it includes prompts tailored to Thai culture, offering deeper insights into culturally specific attacks.
Note: The Monarchy type of harm has been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiSafetyBench.this-that-spatial-bench
spatial-decisions
7,305 multiple-choice decision questions over 6,525 distinct simulated states, in
15 families and two environments. Every answer is computed from the simulator, not
annotated by a person and not taken from a model. That is the point of the set: on a question whose
answer is derived from the rules of the environment, a disagreement is a mistake, and there is
nothing to argue about.
The set was built to replace a much narrower public artefact: a recording of 68… See the full description on the dataset page: https://huggingface.co/datasets/limberc/this-that-spatial-bench.korean-contract-derivation-bench
Korean Contract Derivation Benchmark
A model that derives a discount and a unit-price total without being told to will still return the
VAT-inclusive figure verbatim when asked for the contract amount — and a generic instruction to
"compute the relation" does not fix it. This benchmark isolates that.
The finding
240 synthetic Korean procurement summaries. Half state the contract amount directly among distractor
figures (disambiguate); half state it only as a… See the full description on the dataset page: https://huggingface.co/datasets/ThakiCloud/korean-contract-derivation-bench.Thai-Physics-Data-40KThai-Physics-Data is a Thai-Based physics data with more than 40k lines of data.
Data Sources:
ArtifactAI/arxiv-physics-instruct-tune-30k (CC BY-NC 2.0)
camel-ai/physics
How to load Data (Hugging Face)
from datasets import load_dataset
Thai_Physics_Data = load_dataset("Kongongong/Thai-Physics-Data-40K")
Thai_Physics_Data = Thai_Physics_Data['train']
def format_data():
....
data =[]
format_data()
data = Dataset.from_dict({"text": data})
How to load Data (CSV)
from… See the full description on the dataset page: https://huggingface.co/datasets/Kongongong/Thai-Physics-Data-40K.vncompress-vi-v2
VNCompress-VI v2 — query-conditioned context compression for Vietnamese
⚠️ BẢN ĐANG REBUILD — WORK IN PROGRESS.
corpus/qa/qa_synthetic đã sạch và ổn định. compression.jsonl mới có
486 hàng thật (prompt v3, trích xuất) trên tổng số dự kiến 100.000+ —
quá trình sinh đang tạm dừng vì tỉ lệ drop "không đạt ngân sách token"
rất cao (90% ở batch gần nhất) chưa được xử lý gốc rễ. Đừng dùng
compression.jsonl để báo cáo kết quả benchmark hay train E5/E6 ở quy mô
lớn — số hàng hiện tại… See the full description on the dataset page: https://huggingface.co/datasets/thanthienhai/vncompress-vi-v2.vietnamese-legal-qa
thangvip/vietnamese-legal-qa
Dataset Description
This dataset contains Vietnamese legal documents with automatically generated question-answer pairs. Each document includes comprehension questions of varying difficulty levels (easy, medium, hard) and types (factual, interpretation, analytical, application).
Dataset Structure
Data Fields
doc_name: Name of the legal document
doc_type_name: Type of document (e.g., "Luật" for Law)
article_content:… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/vietnamese-legal-qa.thai-gov-procurement_regulation-17-amend-21
🇹🇭 Dataset Card for Thai Government Procurement Dataset
ℹ️ This dataset is optimized for procurement-related NLP tasks in Thai.
This dataset contains a collection of procurement regulations, instructions, and responses focused on public sector purchasing, contract management, and compliance with Thai government standards. It aims to support natural language processing tasks involving procurement assistance, such as chatbot development, procurement dialogue generation… See the full description on the dataset page: https://huggingface.co/datasets/amornpan/thai-gov-procurement_regulation-17-amend-21.Medical-o1-Reasoning-SFT-Thai
Medical-GPT-Reasoning-Thai
Dataset Summary
This dataset contains medical Q&A data in JSON format, designed for fine-tuning AI models in medical reasoning and response generation.representing a medical question, complex chain-of-thought reasoning, and a concise response. All content is in Thai language.
The dataset is derived from a larger medical Q&A collection and has been processed to ensure JSON validity, with multi-line objects combined into single valid entries.… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-Reasoning-SFT-Thai.thai-qa-rag-answer-dataset
Thai QA RAG Answer Synthesis Dataset
Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined
Rows: 9999 rows.
Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Examples
{"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.Thai-R1-Distill-SFT
Thai R1 Distill SFT
Thai Reasoning Dataset for Supervised Finetuning
Translated by iApp Technology
combined-vietnamese-legal-text
thangvip/combined-vietnamese-legal-text
Dataset Description
This is a combined Vietnamese legal dataset with question-answer pairs formatted in a single text column. It combines two datasets:
thangvip/vietnamese-legal-qa (9,715 examples)
thangvip/law-reading-comprehension-qa-filtered (205,369 examples)
Dataset Structure
Data Fields
text: Combined text containing legal content followed by question-answer pairs in XML-like format
Format… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/combined-vietnamese-legal-text.ThaiBarAssociationEditorial
About Thai Bar Under The Royal Patronage
"The Thai Bar Association originated from the Law School established under the royal initiative of King Chulalongkorn (Rama V).
In 1948, the Thai Bar Association established the Institute of Legal Education with the objective of imparting and promoting legal education and professional expertise in the practice of law. Instruction commenced for the first time in November 1948, with a curriculum modeled after the Council of Legal Education in… See the full description on the dataset page: https://huggingface.co/datasets/mspkrai/ThaiBarAssociationEditorial.
