datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.Thailand-Stock-Symbols-and-Metadata
Thailand Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Thailand.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Thailand-Stock-Symbols-and-Metadata.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.ThaiQA_LST20SuperAI Engineer Season 2 , Machima
Machima_ThaiQA_LST20 เป็นชุดข้อมูลที่สกัดหาคำถาม และคำตอบ จากบทความในชุดข้อมูล LST20 โดยสกัดได้คำถาม-ตอบทั้งหมด 7,642 คำถาม มีข้อมูล 4 คอลัมน์ ประกอบด้วย context, question, answer และ status ตามลำดับ
แสดงตัวอย่างดังนี้
context : ด.ต.ประสิทธิ์ ชาหอมชื่นอายุ 55 ปี ผบ.หมู่งาน ป.ตชด. 24 อุดรธานีถูกยิงด้วยอาวุธปืนอาก้าเข้าที่แขนซ้าย 3 นัดหน้าท้อง 1 นัดส.ต.อ.ประเสริฐ ใหญ่สูงเนินอายุ 35 ปี ผบ.หมู่กก. 1 ปส.2 บช.ปส. ถูกยิงเข้าที่แขนขวากระดูกแตกละเอียดร.ต.อ.ชวพล… See the full description on the dataset page: https://huggingface.co/datasets/SuperAI2-Machima/ThaiQA_LST20.Yord_ThaiQA_LST20พี่ยอด และน้อง ๆ ในทีมบ้านมัณิชมา ร่วมกันสร้างชุดข้อมูล คำถาม - คำตอบ จากชุดข้อมูล LST-20
โดยใช้ POS และ NER เพื่อมาสร้างชุดประโยคคำถาม
ได้ข้อมูลคำถาม - ตอบ ทั้งหมดประมาณ 1,000 แถว
thai_buddhist_studies_exam
Thai Buddhist Studies Examination (Nak Tham)
This repository contains multiple-choice questions from the Thai Buddhist Studies
(Nak Tham) examination (2020, 2022, 2023). This dataset can be used for a benchmark for evaluating Large Language Models'
understanding of Thai Buddhist concepts and teachings.
Dataset Statistics
Year
Number of Multiple Choice Questions
2020
1,350
2022
1,400
2023
1,350
Phra Udom thought on the exam: We have reviewed the Nak… See the full description on the dataset page: https://huggingface.co/datasets/biodatlab/thai_buddhist_studies_exam.pali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.thai-investment-consultant-licensing-exams
Thai Public Investment Consultant (IC) Exams Dataset
Overview
This dataset comprises a collection of exam questions and answers from the Thai Public Investment Consultant (IC) Examinations. It's a valuable resource for developing and evaluating question-answering systems in the finance sector.
Dataset Source
The Stock Exchange of Thailand (SET)
Maintainer
Dr. Kobkrit Viriyayudhakorn
Email: kobkrit@iapp.co.th
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-investment-consultant-licensing-exams.ThaiIDCardSynt
Dataset Details
Dataset Description
Curated by: Matichon Maneegard
Shared by [optional]: Matichon Maneegard
Language(s) (NLP): image-to-text
License: apache-2.0
Dataset Sources [optional]
The dataset was entirely synthetic. It does not contain real information or pertain to any specific person.
Uses
Direct Use
Using for tranning OCR or Multimodal.
Dataset Structure
This dataset contains 98 x 6 = 588 samples, and the… See the full description on the dataset page: https://huggingface.co/datasets/Float16-cloud/ThaiIDCardSynt.pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.rag_thai_laws
Thai Laws Dataset
This dataset contains Thai law texts from the Office of the Council of State, Thailand.
The dataset has been cleaned and processed by the iApp Team to improve data quality and accessibility. The cleaning process included:
Converting system IDs to integer format
Removing leading/trailing whitespace from titles and text
Normalizing newlines to maintain consistent formatting
Removing excessive blank lines
The cleaned dataset is now available on Hugging Face for easy… See the full description on the dataset page: https://huggingface.co/datasets/iapp/rag_thai_laws.Thai_Nutrition_Dataset
Dataset Card "Thai Food Nutrition"
Thai Food Nutrition
Data source from Thai Food Composition Tables 2015 Institute of Nutrition, Mahidol University (INMU), Thailand THAIFOODS and ASEANFOODS Regional Centre https://inmu2.mahidol.ac.th/thaifcd/.
ThaiGovernmentLotteryResultsthai-local-language-translation-dataset
Thai Local Language Translation Dataset
Thai Local Language Translation Dataset is a translation dataset for translate Thai Local Language to Thai Central Language. We create the dataset from Thai Dialect Corpus (Thai dialects ASR corpus). We select train set only from Thai Dialect Corpus.
The dataset support Khummuang, Korat, and Pattani.
Reference
Suwanbandit, A., Naowarat, B., Sangpetch, O., Chuangsuwanich, E. (2023) Thai Dialect Corpus and Transfer-based Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-local-language-translation-dataset.xnli2.0_thaiThai-True-Fake-News
Thai Fake News Dataset
Language: Thai
Task: Fake News Classification
Dataset Description
This dataset contains news articles scraped from the Antifakenewscenter Thailand website using Selenium. It spans news published from 2017 to October 2024. The dataset is designed for fake news classification and consists of two main classes:
True News
Fake News
Each class contains 3002 samples, providing a balanced dataset for model training and evaluation.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/EXt1/Thai-True-Fake-News.Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.thai-wikiqa-tha-qaretrievalref: https://aiforthai.in.th/
KamMuang-Thai-Englishthai-mm-dict
🗂️ Dataset Card: Thai-Myanmar Dictionary (2025)
📝 Dataset Summary
The Thai-Myanmar Dictionary (2025) is a high-quality bilingual lexical dataset created by Htet Myet Lynn.It provides direct word-to-word and phrase mappings between Thai and Myanmar (Burmese), supporting both linguistic use and machine learning applications.
The dataset is released under the MIT License, allowing free usage, modification, redistribution, and integration into both academic and commercial… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-mm-dict.Thai_sentimentrag_thai_laws
Thai Laws Dataset
This dataset contains Thai law texts from the Office of the Council of State, Thailand.
The dataset has been cleaned and processed by the iApp Team to improve data quality and accessibility. The cleaning process included:
Converting system IDs to integer format
Removing leading/trailing whitespace from titles and text
Normalizing newlines to maintain consistent formatting
Removing excessive blank lines
The cleaned dataset is now available on Hugging Face for easy… See the full description on the dataset page: https://huggingface.co/datasets/atyko/rag_thai_laws.thai-instructions-ralliothai_sa
Sentiment Analysis Data for the Thai Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Suriyawongkul et al. (2019).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@software{bact_2019_3457447,
author = {Suriyawongkul, Arthit and
Chuangsuwanich, Ekapol and
Chormai, Pattarawat and
Polpanumas, Charin}… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/thai_sa.OmniL2L-Myanmar-Thai
🇲🇲 🔄 🇹🇭 OmniL2L-Myanmar-Thai
💡 Note: This dataset is a dedicated language-pair component of the main multi-language corpus. To access the complete multi-lingual matrix combining all languages simultaneously, please visit the main repository: kalixlouiis/OmniL2L.
OmniL2L-Myanmar-Thai is a trustworthy, human-verified parallel translation dataset pairing Burmese (Myanmar) with Thai. This dataset is custom-tailored for low-resource machine translation and conversational NLP… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/OmniL2L-Myanmar-Thai.Thai-Thangkarn-sentenceThis dataset was developed based on inspiration from (https://huggingface.co/datasets/Intel/polite-guard)
Thai Thang-karn classification
Dataset type: Synthetic and Annotated
Task: Text Classification
Domain: Classification of Ceremonial [พิธีการ], Official [ทางการ], Semi-Official [กึ่งทางการ], Informal [ไม่เป็นทางการ] and coloquial [กันเอง] categories. (According to Thai Grammar.)
Source code: (https://github.com/nnudee/Thai-Thang-karn_text-classification/tree/main/data-generator)… See the full description on the dataset page: https://huggingface.co/datasets/nnudee/Thai-Thangkarn-sentence.xnli2.0_train_thaithai-w2p
Thai W2P
Thai Word-to-Phoneme (W2P) converter.
GitHub: https://github.com/wannaphong/thai_w2p
thai-mm-dict
🗂️ Dataset Card: Thai-Myanmar Dictionary (2025)
📝 Dataset Summary
The Thai-Myanmar Dictionary (2025) is a high-quality bilingual lexical dataset created by Htet Myet Lynn.It provides direct word-to-word and phrase mappings between Thai and Myanmar (Burmese), supporting both linguistic use and machine learning applications.
The dataset is released under the MIT License, allowing free usage, modification, redistribution, and integration into both academic and commercial… See the full description on the dataset page: https://huggingface.co/datasets/htetmyet/thai-mm-dict.
