datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.ThaiMix
ThaiMix (https://arxiv.org/abs/2512.18834) is a Thai pretraining corpus containing 70 billion tokens across 81 million documents (in the minhash subset). Rather than scraping the web again, ThaiMix combines five publicly available Thai datasets, applies Thai-specific quality filtering, and performs cross-dataset deduplication.
Subsets
Subset
Documents
Tokens
Description
minhash_deduped
81.3M
70.5B
Document-level MinHash deduplication
matched
10.9M… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/ThaiMix.ViMed-PET-CT
ViMed-PET-CT
📅 Update: April 23, 2026
🐛 Bug Fixes: Corrected field mismatches (blank/missing fields) and date/filename inconsistencies. Restored missing metadata for patient 1701 (Dec 2023).
✨ New Feature: Added English translations of reports (/reports_en) using Gemma-4-26B-A4B-it.
ℹ️ About the dataset
🍴 Forked and optimized compression of dacthai2807/ViMed-PET, converting .npy and chunked zip files into .npz files.
📝 Better annotation and guideline.
📂… See the full description on the dataset page: https://huggingface.co/datasets/thainamhoang/ViMed-PET-CT.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.ThaMix
ThaMix (https://arxiv.org/abs/2512.18834) is a Thai pretraining corpus built by combining seven publicly available Thai datasets, applying Thai-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses cross-dataset agreement… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/ThaMix.WangchanX-Legal-ThaiCCL-RAG
🏛️ WangchanX-Legal-ThaiCCL-RAG
[Technical Report]
The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.thai-culturax-clean-dataset
Thai CulturaX Clean dataset
The data is sourced from the Thai subset of CulturaX dataset, which itself is sourced from mC4 and four OSCAR corpora.
It has about 8,748,575,684 words (without whitespace) and 16,768,585 lines (97 GB).
It was filtered content promoting gambling, adult content, and narcotics.
GitHub for clean: https://github.com/wannaphong/thai-filter-website
Considerations for Using the Data
This dataset is the cleaned version of the CulturaX datasets… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-culturax-clean-dataset.Thai-English-Corpus
Thai-English Corpus
Thai-English Corpus is a large-scale bilingual corpus containing Thai and English text collected from publicly available datasets on Hugging Face.
The corpus combines educational content, web documents, Wikipedia articles, legal texts, medical articles, financial documents, software-related text, and other general-domain content into a unified format suitable for language model pretraining and NLP research.
Each document is stored with its original source… See the full description on the dataset page: https://huggingface.co/datasets/Naphon/Thai-English-Corpus.thaisum
Dataset Card for ThaiSum
This dataset was forked from thaisum to HF hub.
Dataset Summary
ThaiSum is a large-scale corpus for Thai text summarization obtained from several online news websites namely Thairath, ThaiPBS, Prachathai, and The Standard. This dataset consists of over 350,000 article and summary pairs written by journalists.
Supported Tasks and Leaderboards
summarization, language modeling
Languages
Thai
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaisum.thai-wiki-dataset-v3
Dataset Card for "thai-wiki-dataset-v3"
This dataset collects all Thai Wikimedia project that cleaned all text for Thai language. Example: Wikipedia, Wikiquote, Wikibooks, Wikisource, and Wiktionary.
Use cause: RAG, and pretraining model.
License: cc-by-sa-3.0
thahabiorg_metadata
📖 Thahabi Books Metadata Dataset
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book and includes bibliographic information such as title, author, category, and source details.
📦 Dataset Structure
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book with full bibliographic and structural information.
📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.that-backpacker-article-corpus
That Backpacker Article Corpus
This dataset contains a structured corpus of long-form travel articles published on ThatBackpacker.com, authored primarily by Audrey Bergner as part of the Samuel & Audrey Media Network.
The corpus includes 323 article records covering destination guides, multi-day itineraries, hiking, food travel, cultural experiences, city guides, transportation, accommodations, and practical travel planning.
It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/that-backpacker-article-corpus.SongLyricsDataset contains songs by artists, the names of the songs, the lyrics of the songs, the release date, the cover photo, and the general popularity of the song.
thailaw
Dataset Card for "thailaw"
English
Thai Law Dataset (Act of Parliament)
Data source from Office of the Council of State, Thailand. https://www.krisdika.go.th/
This part of PyThaiNLP Project.
License Dataset is public domain.
Download https://github.com/PyThaiNLP/thai-law/releases
This hub based on Thailaw v0.2.
Thai
คลังข้อมูลกฎหมายไทย (พระราชบัญญัติ)
ข้อมูลเก็บรวบรวมมาจากเว็บไซต์สำนักงานคณะกรรมการกฤษฎีกา https://www.krisdika.go.th/… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thailaw.thai-financial-datasetThis dataset is the cleaned version of the airesearch/CMDF_VISTEC datasets for pretrining model.
It is financial domain for Thai language.
license: cc-by-4.0
marketing-benchmark-of-more-than-10-ai-models
Marketing Benchmark of 10+ AI Models
A 5,000-question benchmark for evaluating LLMs across six dimensions of modern
marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical
Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas.
Every question is independently authored by the AdsGPT Marketing Bench team.
Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended
and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.Zhihu-KOL-More-Than-100-Upvotes对 https://huggingface.co/datasets/wangrui6/Zhihu-KOL 数据进行了初步整理,保留了100赞及以上的数据。
共271261条。
thai-g2p-v4-dataset
Thai G2P v4 dataset
Thai G2P v4 dataset is a Thai grapheme-to-phoneme dataset that was built from Thai W2P and Wiktionary th-pron transliterator. We split the dataset by using the first 2 characters for the split group.
Author: Wannaphong Phatthiyaphaibun
GitHub: https://github.com/PyThaiNLP/thai-g2p-v4
sources
Thai W2P: https://huggingface.co/datasets/wannaphong/thai-w2p
Wiktionary th-pron transliterator: https://github.com/PyThaiNLP/pythainlp/pull/1437
ThaiQA-v1
ThaiQA v1
ThaiQA v1 is a Thai Synthetic QA dataset. It was created from synthetic method using open source LLM in Thai language.
We used Nvidia Nemotron 4 (340B) to create this dataset.
Topics:
Technology and Gadgets 100
Travel and Tourism 91
Food and Cooking 99
Sports and Fitness 50
Arts and Entertainment 24
Home and Garden 72
Fashion and Beauty 99
Science and Nature 100
History and Culture 91
Education and Learning 99
Pets and Animals 83
Relationships and Family 78
Personal… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/ThaiQA-v1.med-app-instruct
ThaiLLM Medical Instruction with Tool Calling
A synthetic Thai medical instruction-following dataset with tool calling capabilities, designed for training language models to handle healthcare-related queries through a mobile health assistant interface.
Dataset Description
This dataset contains multi-turn conversations between users and an AI health assistant, featuring both direct responses and tool-augmented interactions. The conversations simulate a realistic Thai… See the full description on the dataset page: https://huggingface.co/datasets/ThaiLLM/med-app-instruct.thaigov-v2-corpus-31032024
ThaiGov V2 Corpus
GitHub: https://github.com/PyThaiNLP/thaigov-v2-corpus
English
Data from Thai government website. https://www.thaigov.go.th
This part of PyThaiNLP Project.
Compiled by Mr.Wannaphong Phatthiyaphaibun
License Dataset is public domain.
Data format
1 file, 1 news, which is extracted from 1 url.
topic
(Blank line)
content
content
content
content
content
(Blank line)
ที่มา (URL source) : http://www.thaigov.go.th/news/contents/details/NNN… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaigov-v2-corpus-31032024.thai-local-instruction-v2
Thai local instruction v2
Thai local language instruction dataset v2
List:
korat (ภาษาโคราช)
pattani (ภาษาปักษ์ใต้หรือภาษาใต้)
khummuang (ภาษาเหนือหรือภาษาคำเมือง)
isan (ภาษาอีสาน)
Sources:
th.wiktionary.org (CC BY-SA) for khummuang dictionary.
isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence.
pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences.
Created by Wannaphong Phatthiyaphaibun
thai_wikipedia_clean_20230101
Dataset Card for "thai_wikipedia_clean_20230101"
More Information needed
Thai Wikipedia Database dumps to plain text for NLP work.
This dataset was dump on 1 January 2023 from Thai wikipedia.
GitHub: PyThaiNLP / ThaiWiki-clean
Notebook for upload to HF: https://github.com/PyThaiNLP/ThaiWiki-clean/blob/main/thai_wikipedia_clean_20230101_hf.ipynb
thaigov-corpus
ThaiGov corpus
GitHub: https://github.com/PyThaiNLP/thaigov-corpus
English
Data from Thai government website. https://www.thaigov.go.th
This part of PyThaiNLP Project.
Compiled by Mr.Wannaphong Phatthiyaphaibun
License Dataset is public domain.
Data format
1 file, 1 news, which is extracted from 1 url.
topic
(Blank line)
content
content
content
content
content
(Blank line)
ที่มา (URL source) : http://www.thaigov.go.th/news/contents/details/NNN
Thai… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaigov-corpus.thai-sent-local-v2
Thai sent local v2
List:
korat (ภาษาโคราช)
pattani (ภาษาปักษ์ใต้หรือภาษาใต้)
khummuang (ภาษาเหนือหรือภาษาคำเมือง)
isan (ภาษาอีสาน)
Sources:
th.wiktionary.org (CC BY-SA) for khummuang dictionary.
isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence.
pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences.
Created by Wannaphong Phatthiyaphaibun
thai-tnhc2-books
Thai TNHC2 Books
This dataset collect all books from TNHC2 corpus.
We clean the dataset to use text to pretraining model and nlp task.
All books: 353 books
License: CC-0
TNHC2 Dataset (Original) have many a lots of details (chapter, author's detail and more). The dataset is clean to pretraining model and nlp task.
TNHC2 coepus is a Thai old books corpus that all books are copyright expired in Thai law (50 years after the author's death).
TNHC2 Dataset (Original):… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-tnhc2-books.thai_food_v1.0
Thai Food Recipe dataset v1.0
The Thai Food Recipe dataset is a collection of Thai recipes from old Thai books and social networks.
List Book
ตำรับอาหาร - เตื้อง สนิทวงศ์, ม.ร.ว., 2426-2510 - Work in process (ยังไม่ครบ) in v2.0
ผัดกะเพรา - “ทีมครัวเนื้อหอม” จ.ลำปาง
สูตร "เกี๊ยวกุ้ง"
License: cc0-1.0
thai-it-books
Thai IT books
This dataset collects Thai IT books that are the open access books.
license: cc-by-3.0
thai-oldbooks
Thai Old Books dataset
This dataset collect books from Vajirayana library. All books are copyright expired in Thai law (50 years after the author's death).
All books: 75 books.
License: CC-0
News: I created a new dataset named Thai TNHC2 Books that was cleaned from the TNHC2 corpus but this dataset is clear than Thai TNHC2 Books dataset. If you want to train a model. I suggest you mix the two datasets and delete duplicate books.
List Books
บทละครนอกเรื่องสังข์ทอง
ขุนช้างขุนแผน… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-oldbooks.
