datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.ThaiTrees
ThaiTrees
A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and
social media, automatically parsed under the Universal Dependencies framework.
It is released as three artefacts: a raw text corpus, a frequency lexicon, and
a dependency-parsed corpus in CoNLL-U.
Dataset Summary
ThaiTrees contains 341,967,133 tokens across 366,120 documents in four
domains (news, Wikipedia, spoken transcripts, social media). Document
identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.wangchanx-seed-free-synthetic-instruct-thai-120k
Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k
Dataset Summary
This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.spai-ss6-corpus-thai-exam-qa-answers
SPAI SS6 Thai Exam QA With Answers Index
Index repo for normalized Thai O-NET and exam question-answer records with answer keys.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_exam_qa_with_answers
Rows in canonical config: 8,191
Parquet size in canonical config: 0.01 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-exam-qa-answers.thai-bar-exam-judging
Thai Bar-Exam Judging Corpus
Anonymised free-form Thai legal essays from a bar-exam preparation exercise, with three Bar Council-trained examiners scoring every essay and span-anchored inline commentary on roughly two thirds of the answers. Eight LLM examinees took the same exam under the same conditions; their answers were graded blind by the same examiners. Fifteen of the 150 answers were cross-graded by the two non-primary examiners, producing the 3-rater stability subset that… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-bar-exam-judging.spai-ss6-corpus-mental-health-thai
SPAI SS6 Thai Mental Health Index
Index repo for the imported Thai mental-health dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: mental_health_thai
Rows in canonical config: 21,544
Parquet size in canonical config: 0.02 GB
Source license: unknown
License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-mental-health-thai.spai-ss6-corpus-thai-instruction-sft-suraponn
SPAI SS6 Thai Instruction SFT Suraponn Index
Index repo for the Suraponn Thai instruction SFT dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_instruction_sft_suraponn
Rows in canonical config: 131,907
Parquet size in canonical… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-instruction-sft-suraponn.spai-ss6-corpus-thai-wikipedia-clean
SPAI SS6 Thai Wikipedia Clean Corpus Index
Index repo for the Thai Wikipedia clean corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_wikipedia_clean_20230101
Rows in canonical config: 1,436,054
Parquet size in canonical config: 0.26 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-wikipedia-clean.spai-ss6-corpus-thai-synthetic-qa-v1
SPAI SS6 Thai Synthetic QA V1 Index
Index repo for the ThaiSyntheticQA v1 dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_synthetic_qa_v1
Rows in canonical config: 12,668
Parquet size in canonical config: 0.02 GB
Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synthetic-qa-v1.spai-ss6-corpus-thai-idioms-instruction
SPAI SS6 Thai Idioms Instruction Index
Index repo for the Thai idioms instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_idioms_instruction
Rows in canonical config: 1,152
Parquet size in canonical config: 0.00 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-idioms-instruction.spai-ss6-corpus-thai-medical-care-aiotx
SPAI SS6 Thai Medical Care AIoTx Index
Index repo for the imported AIoTx Thai medical-care dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_medical_care_aiotx
Rows in canonical config: 3,599
Parquet size in canonical config: 0.00 GB
Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-medical-care-aiotx.spai-ss6-corpus-thai-local-instruction-v2
SPAI SS6 Thai Local Instruction V2 Index
Index repo for the Thai local instruction v2 dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_local_instruction_v2
Rows in canonical config: 39,829
Parquet size in canonical config: 0.00 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-local-instruction-v2.spai-ss6-corpus-thaisum
SPAI SS6 ThaiSum Corpus Index
Index repo for the ThaiSum corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thaisum
Rows in canonical config: 391,868
Parquet size in canonical config: 1.30 GB
Source license: mit
License review status:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaisum.spai-ss6-corpus-thai-synonym-instruction
SPAI SS6 Thai Synonym Instruction Index
Index repo for the Thai synonym instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_synonym_instruction
Rows in canonical config: 167
Parquet size in canonical config: 0.00 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synonym-instruction.spai-ss6-corpus-khanomtanllm-thai-subset
SPAI SS6 KhanomTanLLM Thai Subset Index
Index repo for the KhanomTanLLM Thai subset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: khanomtanllm_thai_subset
Rows in canonical config: 464,339
Parquet size in canonical config: 1.80 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-khanomtanllm-thai-subset.spai-ss6-corpus-thai-culturax-clean
SPAI SS6 Thai CulturaX Clean Corpus Index
Index repo for the Thai CulturaX clean corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_culturax_clean
Rows in canonical config: 818,727
Parquet size in canonical config: 1.96 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-culturax-clean.spai-ss6-corpus-thai-tourist-attraction-instruction
SPAI SS6 Thai Tourist Attraction Instruction Index
Index repo for the Thai tourist-attraction instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_tourist_attraction_instruction
Rows in canonical config: 51,662
Parquet size… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-tourist-attraction-instruction.spai-ss6-corpus-thaiqa-lst20-mit
SPAI SS6 ThaiQA LST20 MIT Index
Index repo for the ThaiQA LST20 dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thaiqa_lst20_mit
Rows in canonical config: 7,643
Parquet size in canonical config: 0.01 GB
Source license: mit
License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaiqa-lst20-mit.
