CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads2h agoHugging Face02ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face03open-law-data-thailand /ocs-krisdika Open Law Data Thailand: OCS Krisdika Dataset ชุดข้อมูลกฎหมายจาก สำนักงานคณะกรรมการกฤษฎีกา (Office of the Council of State) รวบรวมและจัดทำโดยโครงการ Open Law Data Thailand เพื่อส่งเสริมการเข้าถึงข้อมูลกฎหมายในรูปแบบที่เครื่องอ่านได้ (Machine-Readable) Dataset Structure ข้อมูลถูกจัดเก็บในรูปแบบ JSON Lines (.jsonl) แบ่งไฟล์ตาม ปีและเดือน (YYYY/YYYY-MM.jsonl) เพื่อความสะดวกในการดาวน์โหลดและบริหารจัดการ Data Fields แต่ละบรรทัด (Row) ประกอบด้วยข้อมูลดังนี้: title… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/ocs-krisdika.text-retrieval4 likes1.4k downloads10mo agoHugging Face04AdaMLLab /ThaiMix ThaiMix (https://arxiv.org/abs/2512.18834) is a Thai pretraining corpus containing 70 billion tokens across 81 million documents (in the minhash subset). Rather than scraping the web again, ThaiMix combines five publicly available Thai datasets, applies Thai-specific quality filtering, and performs cross-dataset deduplication. Subsets Subset Documents Tokens Description minhash_deduped 81.3M 70.5B Document-level MinHash deduplication matched 10.9M… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/ThaiMix.texttext-generation10M<n<100M1 likes1.2k downloads8mo agoHugging Face05thainamhoang /ViMed-PET-CT ViMed-PET-CT 📅 Update: April 23, 2026 🐛 Bug Fixes: Corrected field mismatches (blank/missing fields) and date/filename inconsistencies. Restored missing metadata for patient 1701 (Dec 2023). ✨ New Feature: Added English translations of reports (/reports_en) using Gemma-4-26B-A4B-it. ℹ️ About the dataset 🍴 Forked and optimized compression of dacthai2807/ViMed-PET, converting .npy and chunked zip files into .npz files. 📝 Better annotation and guideline. 📂… See the full description on the dataset page: https://huggingface.co/datasets/thainamhoang/ViMed-PET-CT.textimage-to-textn<1K0 likes1k downloads5mo agoHugging Face06SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes515 downloads4mo agoHugging Face07airesearch /WangchanX-Legal-ThaiCCL-RAG 🏛️ WangchanX-Legal-ThaiCCL-RAG [Technical Report] The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.texttext-generation10K<n<100K13 likes280 downloads2y agoHugging Face08pythainlp /thai-culturax-clean-dataset Thai CulturaX Clean dataset The data is sourced from the Thai subset of CulturaX dataset, which itself is sourced from mC4 and four OSCAR corpora. It has about 8,748,575,684 words (without whitespace) and 16,768,585 lines (97 GB). It was filtered content promoting gambling, adult content, and narcotics. GitHub for clean: https://github.com/wannaphong/thai-filter-website Considerations for Using the Data This dataset is the cleaned version of the CulturaX datasets… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-culturax-clean-dataset.texttext-generation10M<n<100M5 likes272 downloads2y agoHugging Face09Naphon /Thai-English-Corpus Thai-English Corpus Thai-English Corpus is a large-scale bilingual corpus containing Thai and English text collected from publicly available datasets on Hugging Face. The corpus combines educational content, web documents, Wikipedia articles, legal texts, medical articles, financial documents, software-related text, and other general-domain content into a unified format suitable for language model pretraining and NLP research. Each document is stored with its original source… See the full description on the dataset page: https://huggingface.co/datasets/Naphon/Thai-English-Corpus.texttext-generation100M<n<1B0 likes269 downloads3mo agoHugging Face10pythainlp /thaisum Dataset Card for ThaiSum This dataset was forked from thaisum to HF hub. Dataset Summary ThaiSum is a large-scale corpus for Thai text summarization obtained from several online news websites namely Thairath, ThaiPBS, Prachathai, and The Standard. This dataset consists of over 350,000 article and summary pairs written by journalists. Supported Tasks and Leaderboards summarization, language modeling Languages Thai Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaisum.textsummarization100K<n<1M5 likes263 downloads3y agoHugging Face11nakhun /thaisumThaiSum is a large-scale corpus for Thai text summarization obtained from several online news websites namely Thairath, ThaiPBS, Prachathai, and The Standard. This dataset consists of over 350,000 article and summary pairs written by journalists.summarization100K<n<1M10 likes257 downloads3y agoHugging Face12pythainlp /thai-wiki-dataset-v3 Dataset Card for "thai-wiki-dataset-v3" This dataset collects all Thai Wikimedia project that cleaned all text for Thai language. Example: Wikipedia, Wikiquote, Wikibooks, Wikisource, and Wiktionary. Use cause: RAG, and pretraining model. License: cc-by-sa-3.0 texttext-generation100K<n<1M11 likes247 downloads3y agoHugging Face13pythainlp /thailaw Dataset Card for "thailaw" English Thai Law Dataset (Act of Parliament) Data source from Office of the Council of State, Thailand. https://www.krisdika.go.th/ This part of PyThaiNLP Project. License Dataset is public domain. Download https://github.com/PyThaiNLP/thai-law/releases This hub based on Thailaw v0.2. Thai คลังข้อมูลกฎหมายไทย (พระราชบัญญัติ) ข้อมูลเก็บรวบรวมมาจากเว็บไซต์สำนักงานคณะกรรมการกฤษฎีกา https://www.krisdika.go.th/… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thailaw.texttext-generation10K<n<100K12 likes143 downloads3y agoHugging Face14phoneee /thai-legal-corpus Thai Legal Corpus v1 Cleaned and deduplicated Thai legal text corpus for Continual Pre-Training (CPT), with structured citation metadata for legal analysis. Dataset Description Metric Value Total records 176,543 Total size 6.03 GB Avg doc length 11,952 chars Sources 3 Year range 1874-2026 Splits train 158,887 / val 8,826 / test 8,830 Sources Source Records krisdika 6,743 supreme_court 127,012 thailaw 42,788… See the full description on the dataset page: https://huggingface.co/datasets/phoneee/thai-legal-corpus.text-generation100K<n<1M0 likes143 downloads7mo agoHugging Face15pythainlp /thai-financial-datasetThis dataset is the cleaned version of the airesearch/CMDF_VISTEC datasets for pretrining model. It is financial domain for Thai language. license: cc-by-4.0 texttext-generation100K<n<1M2 likes141 downloads3y agoHugging Face16opendatalab /WanJuan-Thai 💡 Introduction WanJuan-Thai (万卷丝路-泰语) corpus, with a volume exceeding 155GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Thai.text-generation3 likes123 downloads1y agoHugging Face17pythainlp /thai-g2p-v4-dataset Thai G2P v4 dataset Thai G2P v4 dataset is a Thai grapheme-to-phoneme dataset that was built from Thai W2P and Wiktionary th-pron transliterator. We split the dataset by using the first 2 characters for the split group. Author: Wannaphong Phatthiyaphaibun GitHub: https://github.com/PyThaiNLP/thai-g2p-v4 sources Thai W2P: https://huggingface.co/datasets/wannaphong/thai-w2p Wiktionary th-pron transliterator: https://github.com/PyThaiNLP/pythainlp/pull/1437 texttext-generation100K<n<1M5 likes105 downloads3mo agoHugging Face18ThaiSyntheticQA /ThaiQA-v1 ThaiQA v1 ThaiQA v1 is a Thai Synthetic QA dataset. It was created from synthetic method using open source LLM in Thai language. We used Nvidia Nemotron 4 (340B) to create this dataset. Topics: Technology and Gadgets 100 Travel and Tourism 91 Food and Cooking 99 Sports and Fitness 50 Arts and Entertainment 24 Home and Garden 72 Fashion and Beauty 99 Science and Nature 100 History and Culture 91 Education and Learning 99 Pets and Animals 83 Relationships and Family 78 Personal… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/ThaiQA-v1.texttext-generation10K<n<100K5 likes98 downloads2y agoHugging Face19ThaiLLM /med-app-instruct ThaiLLM Medical Instruction with Tool Calling A synthetic Thai medical instruction-following dataset with tool calling capabilities, designed for training language models to handle healthcare-related queries through a mobile health assistant interface. Dataset Description This dataset contains multi-turn conversations between users and an AI health assistant, featuring both direct responses and tool-augmented interactions. The conversations simulate a realistic Thai… See the full description on the dataset page: https://huggingface.co/datasets/ThaiLLM/med-app-instruct.texttext-generation1M<n<10M0 likes94 downloads6mo agoHugging Face20pythainlp /thaigov-v2-corpus-31032024 ThaiGov V2 Corpus GitHub: https://github.com/PyThaiNLP/thaigov-v2-corpus English Data from Thai government website. https://www.thaigov.go.th This part of PyThaiNLP Project. Compiled by Mr.Wannaphong Phatthiyaphaibun License Dataset is public domain. Data format 1 file, 1 news, which is extracted from 1 url. topic (Blank line) content content content content content (Blank line) ที่มา (URL source) : http://www.thaigov.go.th/news/contents/details/NNN… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaigov-v2-corpus-31032024.texttext-generation10K<n<100K1 likes85 downloads2y agoHugging Face21pythainlp /thai-local-instruction-v2 Thai local instruction v2 Thai local language instruction dataset v2 List: korat (ภาษาโคราช) pattani (ภาษาปักษ์ใต้หรือภาษาใต้) khummuang (ภาษาเหนือหรือภาษาคำเมือง) isan (ภาษาอีสาน) Sources: th.wiktionary.org (CC BY-SA) for khummuang dictionary. isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence. pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences. Created by Wannaphong Phatthiyaphaibun texttext-generation10K<n<100K1 likes84 downloads1y agoHugging Face22pythainlp /thai_wikipedia_clean_20230101 Dataset Card for "thai_wikipedia_clean_20230101" More Information needed Thai Wikipedia Database dumps to plain text for NLP work. This dataset was dump on 1 January 2023 from Thai wikipedia. GitHub: PyThaiNLP / ThaiWiki-clean Notebook for upload to HF: https://github.com/PyThaiNLP/ThaiWiki-clean/blob/main/thai_wikipedia_clean_20230101_hf.ipynb texttext-generation1M<n<10M4 likes81 downloads3y agoHugging Face23pythainlp /thaigov-corpus ThaiGov corpus GitHub: https://github.com/PyThaiNLP/thaigov-corpus English Data from Thai government website. https://www.thaigov.go.th This part of PyThaiNLP Project. Compiled by Mr.Wannaphong Phatthiyaphaibun License Dataset is public domain. Data format 1 file, 1 news, which is extracted from 1 url. topic (Blank line) content content content content content (Blank line) ที่มา (URL source) : http://www.thaigov.go.th/news/contents/details/NNN Thai… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaigov-corpus.texttext-generation10K<n<100K1 likes80 downloads2y agoHugging Face24pythainlp /thai-sent-local-v2 Thai sent local v2 List: korat (ภาษาโคราช) pattani (ภาษาปักษ์ใต้หรือภาษาใต้) khummuang (ภาษาเหนือหรือภาษาคำเมือง) isan (ภาษาอีสาน) Sources: th.wiktionary.org (CC BY-SA) for khummuang dictionary. isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence. pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences. Created by Wannaphong Phatthiyaphaibun texttext-generation1K<n<10K0 likes75 downloads1y agoHugging Face25pythainlp /thai-tnhc2-books Thai TNHC2 Books This dataset collect all books from TNHC2 corpus. We clean the dataset to use text to pretraining model and nlp task. All books: 353 books License: CC-0 TNHC2 Dataset (Original) have many a lots of details (chapter, author's detail and more). The dataset is clean to pretraining model and nlp task. TNHC2 coepus is a Thai old books corpus that all books are copyright expired in Thai law (50 years after the author's death). TNHC2 Dataset (Original):… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-tnhc2-books.texttext-generationn<1K0 likes73 downloads3y agoHugging Face26pythainlp /thai_food_v1.0 Thai Food Recipe dataset v1.0 The Thai Food Recipe dataset is a collection of Thai recipes from old Thai books and social networks. List Book ตำรับอาหาร - เตื้อง สนิทวงศ์, ม.ร.ว., 2426-2510 - Work in process (ยังไม่ครบ) in v2.0 ผัดกะเพรา - “ทีมครัวเนื้อหอม” จ.ลำปาง สูตร "เกี๊ยวกุ้ง" License: cc0-1.0 texttext-generationn<1K9 likes72 downloads3y agoHugging Face27pythainlp /thai-it-books Thai IT books This dataset collects Thai IT books that are the open access books. license: cc-by-3.0 texttext-generationn<1K0 likes69 downloads3y agoHugging Face28pythainlp /thai-oldbooks Thai Old Books dataset This dataset collect books from Vajirayana library. All books are copyright expired in Thai law (50 years after the author's death). All books: 75 books. License: CC-0 News: I created a new dataset named Thai TNHC2 Books that was cleaned from the TNHC2 corpus but this dataset is clear than Thai TNHC2 Books dataset. If you want to train a model. I suggest you mix the two datasets and delete duplicate books. List Books บทละครนอกเรื่องสังข์ทอง ขุนช้างขุนแผน… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-oldbooks.texttext-generationn<1K2 likes69 downloads3y agoHugging Face29wannaphong /KhanomTanLLM-pretrained-dataset-thai-subset KhanomTanLLM pretrained dataset (Thai subset) This daataset collect all raw text for pretraining LLM. (Thai subset) Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0 pythainlp/thai-tnhc2-books pythainlp/thai-constitution-corpus pythainlp/thai-it-books pythainlp/prd_news_3011202 pythainlp/thailand-policy-statements pythainlp/thai-cc-license pythainlp/blognone_news pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.texttext-generation10M<n<100M0 likes68 downloads2y agoHugging Face30parinzee /seed-free-synthetic-instruct-thai-v1 Seed-Free Synthetic Instruct Thai v1 (F+C+D+) This dataset is part of the research paper "Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai" submitted to ACL SRW 2024. It represents the best-performing synthetic dataset (F+C+D+) generated using our novel seed-free framework for low-resource languages, specifically Thai. Dataset Details Size: 5,000 instructions Language: Thai Task: Instruction-tuning for Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/parinzee/seed-free-synthetic-instruct-thai-v1.texttext-generation1K<n<10K3 likes64 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.