CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.texttext-retrieval1M<n<10M13 likes25k downloads12h agoHugging Face02wayu-ai /thai-commoncrawl-index Thai Common Crawl Index (2019–2026) An index of every page Common Crawl detected as Thai across 70 monthly crawls, from January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30). 932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains Each row records where the page lives inside Common Crawl's WARC archives — file name, byte offset, and record length — so you can fetch exactly the pages you want with HTTP range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.tabular100M<n<1B0 likes7.5k downloads1mo agoHugging Face03typhoon-ai /thai_exam Dataset Card for Thai_Exam ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows: ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.tabularquestion-answeringn<1K19 likes2.1k downloads2y agoHugging Face04matichon /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.textquestion-answering1K<n<10K0 likes2k downloads5mo agoHugging Face05ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face06openthaigpt /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.textquestion-answering1K<n<10K8 likes1.5k downloads3y agoHugging Face07iapp /thai_handwriting_dataset Thai Handwriting Dataset This dataset combines two major Thai handwriting datasets: BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet) Thai Handwritten Free Dataset by Wang (train-0001.parquet onwards) Maintainer kobkrit@iapp.co.th Dataset Description BEST 2019 Dataset Contains handwritten Thai text images along with their ground truth transcriptions. The images have been processed and standardized for machine learning tasks.… See the full description on the dataset page: https://huggingface.co/datasets/iapp/thai_handwriting_dataset.imagetext-to-image10K<n<100K22 likes1.4k downloads2y agoHugging Face08AdaMLLab /ThaiMix ThaiMix (https://arxiv.org/abs/2512.18834) is a Thai pretraining corpus containing 70 billion tokens across 81 million documents (in the minhash subset). Rather than scraping the web again, ThaiMix combines five publicly available Thai datasets, applies Thai-specific quality filtering, and performs cross-dataset deduplication. Subsets Subset Documents Tokens Description minhash_deduped 81.3M 70.5B Document-level MinHash deduplication matched 10.9M… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/ThaiMix.texttext-generation10M<n<100M1 likes1.1k downloads8mo agoHugging Face09thainamhoang /ViMed-PET-CT ViMed-PET-CT 📅 Update: April 23, 2026 🐛 Bug Fixes: Corrected field mismatches (blank/missing fields) and date/filename inconsistencies. Restored missing metadata for patient 1701 (Dec 2023). ✨ New Feature: Added English translations of reports (/reports_en) using Gemma-4-26B-A4B-it. ℹ️ About the dataset 🍴 Forked and optimized compression of dacthai2807/ViMed-PET, converting .npy and chunked zip files into .npz files. 📝 Better annotation and guideline. 📂… See the full description on the dataset page: https://huggingface.co/datasets/thainamhoang/ViMed-PET-CT.textimage-to-textn<1K0 likes958 downloads4mo agoHugging Face10HaruthaiAi /TreeOil_Painting_ScientificJourney_Thailand_CaseStudy🧪 Tree Oil Painting: A Scientific Journey – Thailand Case Study This dataset documents a rare and detailed forensic investigation of a mysterious 19th-century oil painting, known as The Tree Oil Painting, using scientific methods and AI-assisted analysis. Compiled in Thailand between 2015 and 2025, this work represents a grassroots effort to validate the painting’s origins through physical evidence, pigment mapping, synchrotron spectroscopy, and historical comparison. 🧩 Overview Title: Tree… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/TreeOil_Painting_ScientificJourney_Thailand_CaseStudy.imagen<1K0 likes884 downloads1y agoHugging Face11AdoCleanCode /SPEEED_s3_words_thai_0k-250ktext100K<n<1M0 likes878 downloads7mo agoHugging Face12typhoon-ai /ThaiOCRBench ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai ThaiOCRBench is the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks.Inspired by OCRBench v2, it contains 2,808 human-annotated samples across 13 diverse tasks, including table parsing, chart understanding, full-page OCR, key information extraction, and visual question answering. The benchmark enables standardized zero-shot… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiOCRBench.imageimage-text-to-text1K<n<10K7 likes821 downloads10mo agoHugging Face13mcshao /Thai-understanding Thai-Understanding: Thai-SUP & XLSR-Thai Overview Thai-Understanding is an open-source repository that provides a solution for speech understanding in the Thai language. This repository includes: Thai-SUP: The first open-source Thai speech understanding dataset, which includes over 1,000 hours of data across three tasks: Intent Classification (IC), Named Entity Recognition (NER), and Speech Rephrasing (SR). XLSR-Thai: The first large-scale self-supervised learning (SSL)… See the full description on the dataset page: https://huggingface.co/datasets/mcshao/Thai-understanding.tabular100K<n<1M6 likes815 downloads1y agoHugging Face14wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes807 downloads1mo agoHugging Face15iapp /MMMU-Thai MMMU Thai (MMMU Benchmark Translated to Thai) MMMU Thai is a dataset for evaluating multimodal models on massive multi-discipline tasks requiring college-level knowledge and deliberate reasoning. This dataset is translated from MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) into Thai. Dataset Details MMMU Thai consists of 11,500 meticulously collected multimodal questions from college exams, quizzes, and textbooks… See the full description on the dataset page: https://huggingface.co/datasets/iapp/MMMU-Thai.imagequestion-answering10K<n<100K2 likes669 downloads2y agoHugging Face16SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes617 downloads4mo agoHugging Face17VISAI-AI /gsm8k-thai gsm8k-thai This dataset is a Thai translation of the GSM8k benchmark (https://huggingface.co/datasets/openai/gsm8k), a dataset of grade school math word problems. The translation was performed using Claude 3.5 Sonnet. It is intended for evaluating the performance of language models on mathematical reasoning in the Thai language. The split of training and test data follows the original GSM8k dataset. Annotations source: claude-3.5-sonnet language: en -> th… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/gsm8k-thai.text1K<n<10K0 likes538 downloads2y agoHugging Face18typhoon-ai /thai-dialect-isan-dataset Dataset Card for Thai Dialect Isan Speech Corpus Dataset Description This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language. The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.textautomatic-speech-recognition10K<n<100K5 likes530 downloads10mo agoHugging Face19open-law-data-thailand /ocs-krisdika Open Law Data Thailand: OCS Krisdika Dataset ชุดข้อมูลกฎหมายจาก สำนักงานคณะกรรมการกฤษฎีกา (Office of the Council of State) รวบรวมและจัดทำโดยโครงการ Open Law Data Thailand เพื่อส่งเสริมการเข้าถึงข้อมูลกฎหมายในรูปแบบที่เครื่องอ่านได้ (Machine-Readable) Dataset Structure ข้อมูลถูกจัดเก็บในรูปแบบ JSON Lines (.jsonl) แบ่งไฟล์ตาม ปีและเดือน (YYYY/YYYY-MM.jsonl) เพื่อความสะดวกในการดาวน์โหลดและบริหารจัดการ Data Fields แต่ละบรรทัด (Row) ประกอบด้วยข้อมูลดังนี้: title… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/ocs-krisdika.text-retrieval4 likes479 downloads10mo agoHugging Face20CMKL /Porjai-Thai-voice-dataset-central Porjai-Thai-voice-dataset-central This corpus contains a officially split of 700 hours for Central Thai, and 40 hours for the three dialect each. The corpus is designed such that there are some parallel sentences between the dialects, making it suitable for Speech and Machine translation research. Our demo ASR model can be found at https://www.cmkl.ac.th/research/porjai. The Thai Central data was collected using Wang Data Market. Since parts of this corpus are in the ML-SUPERB… See the full description on the dataset page: https://huggingface.co/datasets/CMKL/Porjai-Thai-voice-dataset-central.audio100K<n<1M18 likes393 downloads2y agoHugging Face21Thai-Binh /trainvideo1K<n<10K0 likes389 downloads2y agoHugging Face22airesearch /thai-ser 🇹🇭 THAI-SER Dataset 🎭 [📝 Paper (preprint)] Published by: AI Research Institute of Thailand (AIResearch) In collaboration with: Vidyasirimedhi Institute of Science and Technology (VISTEC) Digital Economy Promotion Agency (depa) Department of Computer Engineering, Faculty of Engineering, Chulalongkorn University Department of Dramatic Arts, Faculty of Arts, Chulalongkorn University Sponsored by: Advanced Info Services Public Company Limited (AIS), and Siam Commercial… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/thai-ser.audioaudio-classification10K<n<100K6 likes367 downloads2mo agoHugging Face23RJTPP /thai_exam-reformattedReformatted version of scb10x/thai_exam Additional Changes: Fix math incorrect answer ถ้า \log_{\frac{1}{4}} 256 + \frac{2\log 625}{\log 5} = 3^a เมื่อ a เป็นจำนวนจริง แล้วคำตอบของ a เท่ากับเท่าใด? a. \log_{3} 2 b. \log_{3} 4 c. \log_{3} \frac{33}{4} d. \log_{3} 10 e. \log_{3} 12 # Original answer: d (\log_{3} 10) # Correction : b (\log_{3} 4) textquestion-answering1K<n<10K0 likes362 downloads1y agoHugging Face24kunato /thai-exam-seacrowdtextn<1K1 likes344 downloads2y agoHugging Face25n-order /Thai-dialect-corpusaudio100K<n<1M1 likes338 downloads2y agoHugging Face26wayu-ai /thai-contextasr-bench Thai Contextual-Biasing ASR Benchmark TL;DR Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant? Each utterance comes with a bias list: entity strings (brands, person names, places) that may or may not be spoken in the audio, written the way a real Thai user would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.audioautomatic-speech-recognition1K<n<10K1 likes337 downloads2mo agoHugging Face27kunato /thaillm-leaderboard-dataset0 likes330 downloads2y agoHugging Face28nakhun /thaisumThaiSum is a large-scale corpus for Thai text summarization obtained from several online news websites namely Thairath, ThaiPBS, Prachathai, and The Standard. This dataset consists of over 350,000 article and summary pairs written by journalists.summarization100K<n<1M10 likes292 downloads3y agoHugging Face29airesearch /WangchanX-Legal-ThaiCCL-RAG 🏛️ WangchanX-Legal-ThaiCCL-RAG [Technical Report] The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.texttext-generation10K<n<100K13 likes289 downloads2y agoHugging Face30Korn1886 /thai_handwriting_dataset Thai Handwriting Dataset This dataset combines two major Thai handwriting datasets: BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet) Thai Handwritten Free Dataset by Wang (train-0001.parquet onwards) Maintainer kobkrit@iapp.co.th Dataset Description BEST 2019 Dataset Contains handwritten Thai text images along with their ground truth transcriptions. The images have been processed and standardized for machine… See the full description on the dataset page: https://huggingface.co/datasets/Korn1886/thai_handwriting_dataset.text-to-image10K<n<100K0 likes285 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.