datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.thahabiorg_metadata
📖 Thahabi Books Metadata Dataset
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book and includes bibliographic information such as title, author, category, and source details.
📦 Dataset Structure
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book with full bibliographic and structural information.
📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.SongLyricsDataset contains songs by artists, the names of the songs, the lyrics of the songs, the release date, the cover photo, and the general popularity of the song.
ThaiTrees
ThaiTrees
A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and
social media, automatically parsed under the Universal Dependencies framework.
It is released as three artefacts: a raw text corpus, a frequency lexicon, and
a dependency-parsed corpus in CoNLL-U.
Dataset Summary
ThaiTrees contains 341,967,133 tokens across 366,120 documents in four
domains (news, Wikipedia, spoken transcripts, social media). Document
identifiers are shared… See the full description on the dataset page: https://huggingface.co/datasets/nlp-chula/ThaiTrees.wangchanx-seed-free-synthetic-instruct-thai-120k
Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k
Dataset Summary
This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.agentic-coding-traces
Agentic Coding Mooncake Traces
Synthetic agentic coding benchmark datasets in Mooncake trace (JSONL) format,
generated with AIPerf 0.9.0 for LLM inference benchmarking.
Designed for use with InferenceX via the
agentic-replay scenario-type and aiperf_adapter.py.
Files
File
Sessions
Turns
max_prompt_tokens
Seed
64k/dataset.jsonl
1,000
18,595
65,536
42
128k/dataset.jsonl
1,000
16,957
131,072
42
Format
Each line is a Mooncake trace… See the full description on the dataset page: https://huggingface.co/datasets/thangquang09/agentic-coding-traces.spai-ss6-corpus-thai-exam-qa-answers
SPAI SS6 Thai Exam QA With Answers Index
Index repo for normalized Thai O-NET and exam question-answer records with answer keys.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_exam_qa_with_answers
Rows in canonical config: 8,191
Parquet size in canonical config: 0.01 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-exam-qa-answers.thai-bar-exam-judging
Thai Bar-Exam Judging Corpus
Anonymised free-form Thai legal essays from a bar-exam preparation exercise, with three Bar Council-trained examiners scoring every essay and span-anchored inline commentary on roughly two thirds of the answers. Eight LLM examinees took the same exam under the same conditions; their answers were graded blind by the same examiners. Fifteen of the 150 answers were cross-graded by the two non-primary examiners, producing the 3-rater stability subset that… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-bar-exam-judging.spai-ss6-corpus-mental-health-thai
SPAI SS6 Thai Mental Health Index
Index repo for the imported Thai mental-health dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: mental_health_thai
Rows in canonical config: 21,544
Parquet size in canonical config: 0.02 GB
Source license: unknown
License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-mental-health-thai.spai-ss6-corpus-thai-instruction-sft-suraponn
SPAI SS6 Thai Instruction SFT Suraponn Index
Index repo for the Suraponn Thai instruction SFT dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_instruction_sft_suraponn
Rows in canonical config: 131,907
Parquet size in canonical… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-instruction-sft-suraponn.spai-ss6-corpus-thai-wikipedia-clean
SPAI SS6 Thai Wikipedia Clean Corpus Index
Index repo for the Thai Wikipedia clean corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_wikipedia_clean_20230101
Rows in canonical config: 1,436,054
Parquet size in canonical config: 0.26 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-wikipedia-clean.spai-ss6-corpus-thai-synthetic-qa-v1
SPAI SS6 Thai Synthetic QA V1 Index
Index repo for the ThaiSyntheticQA v1 dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_synthetic_qa_v1
Rows in canonical config: 12,668
Parquet size in canonical config: 0.02 GB
Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synthetic-qa-v1.spai-ss6-corpus-thai-idioms-instruction
SPAI SS6 Thai Idioms Instruction Index
Index repo for the Thai idioms instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_idioms_instruction
Rows in canonical config: 1,152
Parquet size in canonical config: 0.00 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-idioms-instruction.spai-ss6-corpus-thai-medical-care-aiotx
SPAI SS6 Thai Medical Care AIoTx Index
Index repo for the imported AIoTx Thai medical-care dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_medical_care_aiotx
Rows in canonical config: 3,599
Parquet size in canonical config: 0.00 GB
Source license:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-medical-care-aiotx.spai-ss6-corpus-thai-local-instruction-v2
SPAI SS6 Thai Local Instruction V2 Index
Index repo for the Thai local instruction v2 dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_local_instruction_v2
Rows in canonical config: 39,829
Parquet size in canonical config: 0.00 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-local-instruction-v2.spai-ss6-corpus-thaisum
SPAI SS6 ThaiSum Corpus Index
Index repo for the ThaiSum corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thaisum
Rows in canonical config: 391,868
Parquet size in canonical config: 1.30 GB
Source license: mit
License review status:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaisum.spai-ss6-corpus-thai-synonym-instruction
SPAI SS6 Thai Synonym Instruction Index
Index repo for the Thai synonym instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_synonym_instruction
Rows in canonical config: 167
Parquet size in canonical config: 0.00 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synonym-instruction.spai-ss6-corpus-khanomtanllm-thai-subset
SPAI SS6 KhanomTanLLM Thai Subset Index
Index repo for the KhanomTanLLM Thai subset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: khanomtanllm_thai_subset
Rows in canonical config: 464,339
Parquet size in canonical config: 1.80 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-khanomtanllm-thai-subset.spai-ss6-corpus-thai-culturax-clean
SPAI SS6 Thai CulturaX Clean Corpus Index
Index repo for the Thai CulturaX clean corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_culturax_clean
Rows in canonical config: 818,727
Parquet size in canonical config: 1.96 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-culturax-clean.spai-ss6-corpus-thai-tourist-attraction-instruction
SPAI SS6 Thai Tourist Attraction Instruction Index
Index repo for the Thai tourist-attraction instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_tourist_attraction_instruction
Rows in canonical config: 51,662
Parquet size… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-tourist-attraction-instruction.spai-ss6-corpus-thaiqa-lst20-mit
SPAI SS6 ThaiQA LST20 MIT Index
Index repo for the ThaiQA LST20 dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thaiqa_lst20_mit
Rows in canonical config: 7,643
Parquet size in canonical config: 0.01 GB
Source license: mit
License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thaiqa-lst20-mit.
