CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads5h agoHugging Face02ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face03AdaMLLab /ThaiMix ThaiMix (https://arxiv.org/abs/2512.18834) is a Thai pretraining corpus containing 70 billion tokens across 81 million documents (in the minhash subset). Rather than scraping the web again, ThaiMix combines five publicly available Thai datasets, applies Thai-specific quality filtering, and performs cross-dataset deduplication. Subsets Subset Documents Tokens Description minhash_deduped 81.3M 70.5B Document-level MinHash deduplication matched 10.9M… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/ThaiMix.texttext-generation10M<n<100M1 likes1.2k downloads8mo agoHugging Face04thainamhoang /ViMed-PET-CT ViMed-PET-CT 📅 Update: April 23, 2026 🐛 Bug Fixes: Corrected field mismatches (blank/missing fields) and date/filename inconsistencies. Restored missing metadata for patient 1701 (Dec 2023). ✨ New Feature: Added English translations of reports (/reports_en) using Gemma-4-26B-A4B-it. ℹ️ About the dataset 🍴 Forked and optimized compression of dacthai2807/ViMed-PET, converting .npy and chunked zip files into .npz files. 📝 Better annotation and guideline. 📂… See the full description on the dataset page: https://huggingface.co/datasets/thainamhoang/ViMed-PET-CT.textimage-to-textn<1K0 likes1k downloads5mo agoHugging Face05SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes515 downloads4mo agoHugging Face06AdaMLLab /ThaMix ThaMix (https://arxiv.org/abs/2512.18834) is a Thai pretraining corpus built by combining seven publicly available Thai datasets, applying Thai-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses cross-dataset agreement… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/ThaMix.texttext-generation100M<n<1B1 likes301 downloads5mo agoHugging Face07airesearch /WangchanX-Legal-ThaiCCL-RAG 🏛️ WangchanX-Legal-ThaiCCL-RAG [Technical Report] The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.texttext-generation10K<n<100K13 likes280 downloads2y agoHugging Face08pythainlp /thai-culturax-clean-dataset Thai CulturaX Clean dataset The data is sourced from the Thai subset of CulturaX dataset, which itself is sourced from mC4 and four OSCAR corpora. It has about 8,748,575,684 words (without whitespace) and 16,768,585 lines (97 GB). It was filtered content promoting gambling, adult content, and narcotics. GitHub for clean: https://github.com/wannaphong/thai-filter-website Considerations for Using the Data This dataset is the cleaned version of the CulturaX datasets… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-culturax-clean-dataset.texttext-generation10M<n<100M5 likes272 downloads2y agoHugging Face09Naphon /Thai-English-Corpus Thai-English Corpus Thai-English Corpus is a large-scale bilingual corpus containing Thai and English text collected from publicly available datasets on Hugging Face. The corpus combines educational content, web documents, Wikipedia articles, legal texts, medical articles, financial documents, software-related text, and other general-domain content into a unified format suitable for language model pretraining and NLP research. Each document is stored with its original source… See the full description on the dataset page: https://huggingface.co/datasets/Naphon/Thai-English-Corpus.texttext-generation100M<n<1B0 likes269 downloads3mo agoHugging Face10pythainlp /thaisum Dataset Card for ThaiSum This dataset was forked from thaisum to HF hub. Dataset Summary ThaiSum is a large-scale corpus for Thai text summarization obtained from several online news websites namely Thairath, ThaiPBS, Prachathai, and The Standard. This dataset consists of over 350,000 article and summary pairs written by journalists. Supported Tasks and Leaderboards summarization, language modeling Languages Thai Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaisum.textsummarization100K<n<1M5 likes263 downloads3y agoHugging Face11pythainlp /thai-wiki-dataset-v3 Dataset Card for "thai-wiki-dataset-v3" This dataset collects all Thai Wikimedia project that cleaned all text for Thai language. Example: Wikipedia, Wikiquote, Wikibooks, Wikisource, and Wiktionary. Use cause: RAG, and pretraining model. License: cc-by-sa-3.0 texttext-generation100K<n<1M11 likes247 downloads3y agoHugging Face12freococo /thahabiorg_metadata 📖 Thahabi Books Metadata Dataset This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book and includes bibliographic information such as title, author, category, and source details. 📦 Dataset Structure This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book with full bibliographic and structural information. 📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.tabulartext-generation1M<n<10M0 likes158 downloads3mo agoHugging Face13samuelandaudreymedianetwork /that-backpacker-article-corpus That Backpacker Article Corpus This dataset contains a structured corpus of long-form travel articles published on ThatBackpacker.com, authored primarily by Audrey Bergner as part of the Samuel & Audrey Media Network. The corpus includes 323 article records covering destination guides, multi-day itineraries, hiking, food travel, cultural experiences, city guides, transportation, accommodations, and practical travel planning. It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/that-backpacker-article-corpus.texttext-generation100K<n<1M1 likes157 downloads4mo agoHugging Face14ThatOneShortGuy /SongLyricsDataset contains songs by artists, the names of the songs, the lyrics of the songs, the release date, the cover photo, and the general popularity of the song. imagetext-generation100K<n<1M4 likes150 downloads3y agoHugging Face15pythainlp /thailaw Dataset Card for "thailaw" English Thai Law Dataset (Act of Parliament) Data source from Office of the Council of State, Thailand. https://www.krisdika.go.th/ This part of PyThaiNLP Project. License Dataset is public domain. Download https://github.com/PyThaiNLP/thai-law/releases This hub based on Thailaw v0.2. Thai คลังข้อมูลกฎหมายไทย (พระราชบัญญัติ) ข้อมูลเก็บรวบรวมมาจากเว็บไซต์สำนักงานคณะกรรมการกฤษฎีกา https://www.krisdika.go.th/… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thailaw.texttext-generation10K<n<100K12 likes143 downloads3y agoHugging Face16pythainlp /thai-financial-datasetThis dataset is the cleaned version of the airesearch/CMDF_VISTEC datasets for pretrining model. It is financial domain for Thai language. license: cc-by-4.0 texttext-generation100K<n<1M2 likes141 downloads3y agoHugging Face17adsgpt /marketing-benchmark-of-more-than-10-ai-models Marketing Benchmark of 10+ AI Models A 5,000-question benchmark for evaluating LLMs across six dimensions of modern marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas. Every question is independently authored by the AdsGPT Marketing Bench team. Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.textquestion-answering1K<n<10K5 likes121 downloads4mo agoHugging Face18bzb2023 /Zhihu-KOL-More-Than-100-Upvotes对 https://huggingface.co/datasets/wangrui6/Zhihu-KOL 数据进行了初步整理,保留了100赞及以上的数据。 共271261条。 texttext-generation100K<n<1M16 likes116 downloads1y agoHugging Face19pythainlp /thai-g2p-v4-dataset Thai G2P v4 dataset Thai G2P v4 dataset is a Thai grapheme-to-phoneme dataset that was built from Thai W2P and Wiktionary th-pron transliterator. We split the dataset by using the first 2 characters for the split group. Author: Wannaphong Phatthiyaphaibun GitHub: https://github.com/PyThaiNLP/thai-g2p-v4 sources Thai W2P: https://huggingface.co/datasets/wannaphong/thai-w2p Wiktionary th-pron transliterator: https://github.com/PyThaiNLP/pythainlp/pull/1437 texttext-generation100K<n<1M5 likes105 downloads3mo agoHugging Face20ThaiSyntheticQA /ThaiQA-v1 ThaiQA v1 ThaiQA v1 is a Thai Synthetic QA dataset. It was created from synthetic method using open source LLM in Thai language. We used Nvidia Nemotron 4 (340B) to create this dataset. Topics: Technology and Gadgets 100 Travel and Tourism 91 Food and Cooking 99 Sports and Fitness 50 Arts and Entertainment 24 Home and Garden 72 Fashion and Beauty 99 Science and Nature 100 History and Culture 91 Education and Learning 99 Pets and Animals 83 Relationships and Family 78 Personal… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/ThaiQA-v1.texttext-generation10K<n<100K5 likes98 downloads2y agoHugging Face21ThaiLLM /med-app-instruct ThaiLLM Medical Instruction with Tool Calling A synthetic Thai medical instruction-following dataset with tool calling capabilities, designed for training language models to handle healthcare-related queries through a mobile health assistant interface. Dataset Description This dataset contains multi-turn conversations between users and an AI health assistant, featuring both direct responses and tool-augmented interactions. The conversations simulate a realistic Thai… See the full description on the dataset page: https://huggingface.co/datasets/ThaiLLM/med-app-instruct.texttext-generation1M<n<10M0 likes94 downloads6mo agoHugging Face22pythainlp /thaigov-v2-corpus-31032024 ThaiGov V2 Corpus GitHub: https://github.com/PyThaiNLP/thaigov-v2-corpus English Data from Thai government website. https://www.thaigov.go.th This part of PyThaiNLP Project. Compiled by Mr.Wannaphong Phatthiyaphaibun License Dataset is public domain. Data format 1 file, 1 news, which is extracted from 1 url. topic (Blank line) content content content content content (Blank line) ที่มา (URL source) : http://www.thaigov.go.th/news/contents/details/NNN… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaigov-v2-corpus-31032024.texttext-generation10K<n<100K1 likes85 downloads2y agoHugging Face23pythainlp /thai-local-instruction-v2 Thai local instruction v2 Thai local language instruction dataset v2 List: korat (ภาษาโคราช) pattani (ภาษาปักษ์ใต้หรือภาษาใต้) khummuang (ภาษาเหนือหรือภาษาคำเมือง) isan (ภาษาอีสาน) Sources: th.wiktionary.org (CC BY-SA) for khummuang dictionary. isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence. pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences. Created by Wannaphong Phatthiyaphaibun texttext-generation10K<n<100K1 likes84 downloads1y agoHugging Face24pythainlp /thai_wikipedia_clean_20230101 Dataset Card for "thai_wikipedia_clean_20230101" More Information needed Thai Wikipedia Database dumps to plain text for NLP work. This dataset was dump on 1 January 2023 from Thai wikipedia. GitHub: PyThaiNLP / ThaiWiki-clean Notebook for upload to HF: https://github.com/PyThaiNLP/ThaiWiki-clean/blob/main/thai_wikipedia_clean_20230101_hf.ipynb texttext-generation1M<n<10M4 likes81 downloads3y agoHugging Face25pythainlp /thaigov-corpus ThaiGov corpus GitHub: https://github.com/PyThaiNLP/thaigov-corpus English Data from Thai government website. https://www.thaigov.go.th This part of PyThaiNLP Project. Compiled by Mr.Wannaphong Phatthiyaphaibun License Dataset is public domain. Data format 1 file, 1 news, which is extracted from 1 url. topic (Blank line) content content content content content (Blank line) ที่มา (URL source) : http://www.thaigov.go.th/news/contents/details/NNN Thai… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thaigov-corpus.texttext-generation10K<n<100K1 likes80 downloads2y agoHugging Face26pythainlp /thai-sent-local-v2 Thai sent local v2 List: korat (ภาษาโคราช) pattani (ภาษาปักษ์ใต้หรือภาษาใต้) khummuang (ภาษาเหนือหรือภาษาคำเมือง) isan (ภาษาอีสาน) Sources: th.wiktionary.org (CC BY-SA) for khummuang dictionary. isan.clubs.chula.ac.th (CC BY-SA-NC) for isan dictionary and sentence. pythainlp/thai-local-language-translation-dataset (CC BY-SA) for korat sentences, pattani sentences, and khummuang sentences. Created by Wannaphong Phatthiyaphaibun texttext-generation1K<n<10K0 likes75 downloads1y agoHugging Face27pythainlp /thai-tnhc2-books Thai TNHC2 Books This dataset collect all books from TNHC2 corpus. We clean the dataset to use text to pretraining model and nlp task. All books: 353 books License: CC-0 TNHC2 Dataset (Original) have many a lots of details (chapter, author's detail and more). The dataset is clean to pretraining model and nlp task. TNHC2 coepus is a Thai old books corpus that all books are copyright expired in Thai law (50 years after the author's death). TNHC2 Dataset (Original):… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-tnhc2-books.texttext-generationn<1K0 likes73 downloads3y agoHugging Face28pythainlp /thai_food_v1.0 Thai Food Recipe dataset v1.0 The Thai Food Recipe dataset is a collection of Thai recipes from old Thai books and social networks. List Book ตำรับอาหาร - เตื้อง สนิทวงศ์, ม.ร.ว., 2426-2510 - Work in process (ยังไม่ครบ) in v2.0 ผัดกะเพรา - “ทีมครัวเนื้อหอม” จ.ลำปาง สูตร "เกี๊ยวกุ้ง" License: cc0-1.0 texttext-generationn<1K9 likes72 downloads3y agoHugging Face29pythainlp /thai-it-books Thai IT books This dataset collects Thai IT books that are the open access books. license: cc-by-3.0 texttext-generationn<1K0 likes69 downloads3y agoHugging Face30pythainlp /thai-oldbooks Thai Old Books dataset This dataset collect books from Vajirayana library. All books are copyright expired in Thai law (50 years after the author's death). All books: 75 books. License: CC-0 News: I created a new dataset named Thai TNHC2 Books that was cleaned from the TNHC2 corpus but this dataset is clear than Thai TNHC2 Books dataset. If you want to train a model. I suggest you mix the two datasets and delete duplicate books. List Books บทละครนอกเรื่องสังข์ทอง ขุนช้างขุนแผน… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-oldbooks.texttext-generationn<1K2 likes69 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.