CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads8h agoHugging Face02typhoon-ai /thai_exam Dataset Card for Thai_Exam ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows: ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.tabularquestion-answeringn<1K19 likes2.2k downloads2y agoHugging Face03ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face04wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes800 downloads1mo agoHugging Face05RJTPP /thai_exam-reformattedReformatted version of scb10x/thai_exam Additional Changes: Fix math incorrect answer ถ้า \log_{\frac{1}{4}} 256 + \frac{2\log 625}{\log 5} = 3^a เมื่อ a เป็นจำนวนจริง แล้วคำตอบของ a เท่ากับเท่าใด? a. \log_{3} 2 b. \log_{3} 4 c. \log_{3} \frac{33}{4} d. \log_{3} 10 e. \log_{3} 12 # Original answer: d (\log_{3} 10) # Correction : b (\log_{3} 4) textquestion-answering1K<n<10K0 likes369 downloads1y agoHugging Face06kunato /thai-exam-seacrowdtextn<1K1 likes353 downloads2y agoHugging Face07pythainlp /thai-culturax-clean-dataset Thai CulturaX Clean dataset The data is sourced from the Thai subset of CulturaX dataset, which itself is sourced from mC4 and four OSCAR corpora. It has about 8,748,575,684 words (without whitespace) and 16,768,585 lines (97 GB). It was filtered content promoting gambling, adult content, and narcotics. GitHub for clean: https://github.com/wannaphong/thai-filter-website Considerations for Using the Data This dataset is the cleaned version of the CulturaX datasets… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-culturax-clean-dataset.texttext-generation10M<n<100M5 likes272 downloads2y agoHugging Face08pnpkke /amazon-2023-thai-1m Amazon 2023 Thai 1M / ชุดข้อมูลสินค้า Amazon ภาษาไทย 1 ล้านรายการ ไทย | English below ชุดข้อมูลสินค้าอีคอมเมิร์ซภาษาไทย 1,000,000 รายการ แปลจากชุดข้อมูล Amazon Reviews '23 Extension ของ Google ด้วยโมเดล typhoon-translate-4b เพื่อใช้ทดสอบ/สาธิตระบบค้นหาเชิงความหมาย (semantic search) ภาษาไทย แหล่งที่มา / Source & Attribution ⚠️ สำคัญ ชุดข้อมูลนี้เป็นงานแปล (derivative work) จาก: ชุดข้อมูลต้นทาง google/extended_amazon_2023_dataset (Amazon Reviews '23… See the full description on the dataset page: https://huggingface.co/datasets/pnpkke/amazon-2023-thai-1m.textsentence-similarity1M<n<10M4 likes219 downloads1mo agoHugging Face09EIRTHAIMED /ThaiExamFinetune Multiple-Choice Question Dataset in Thai Details: This dataset is designed for fine-tuning Thai language models, focusing on the Chain-of-Thought (COT) process, which aids in analyzing questions and deriving correct answers step by step. The dataset consists of multiple-choice questions divided into five categories: O-NET: Ordinary National Educational Test IC: Investment Consultant TGAT: Thai General Aptitude Test TPAT: Thai Professional Aptitude Test… See the full description on the dataset page: https://huggingface.co/datasets/EIRTHAIMED/ThaiExam.textquestion-answering1K<n<10K0 likes134 downloads2y agoHugging Face10ThaiSyntheticQA /ThaiQA-v1 ThaiQA v1 ThaiQA v1 is a Thai Synthetic QA dataset. It was created from synthetic method using open source LLM in Thai language. We used Nvidia Nemotron 4 (340B) to create this dataset. Topics: Technology and Gadgets 100 Travel and Tourism 91 Food and Cooking 99 Sports and Fitness 50 Arts and Entertainment 24 Home and Garden 72 Fashion and Beauty 99 Science and Nature 100 History and Culture 91 Education and Learning 99 Pets and Animals 83 Relationships and Family 78 Personal… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/ThaiQA-v1.texttext-generation10K<n<100K5 likes98 downloads2y agoHugging Face11Jnx03 /kanitakorn-v23-thaiexam-clean-20260614 Qwen v23 ThaiExam Clean Mix Audited fallback mix for Kanitakorn. It avoids v20/v21 replay, avoids v13+v17 double replay, uses v1 repair once, includes all normalized worker v2, adds clean worker v3, and keeps small IF/math/identity retainers. Validation Records: 8,214 inspect_generated_jsonl: 3,214 source rows valid, 0 invalid, 0 duplicate prompts inspect_sft_mix: 0 role errors, 0 empty errors Contamination scan: 0 issues with 8,608 benchmark texts loaded… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-v23-thaiexam-clean-20260614.text1K<n<10K0 likes89 downloads3mo agoHugging Face12wannaphong /KhanomTanLLM-pretrained-dataset-thai-subset KhanomTanLLM pretrained dataset (Thai subset) This daataset collect all raw text for pretraining LLM. (Thai subset) Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0 pythainlp/thai-tnhc2-books pythainlp/thai-constitution-corpus pythainlp/thai-it-books pythainlp/prd_news_3011202 pythainlp/thailand-policy-statements pythainlp/thai-cc-license pythainlp/blognone_news pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.texttext-generation10M<n<100M0 likes68 downloads2y agoHugging Face13parinzee /seed-free-synthetic-instruct-thai-v1 Seed-Free Synthetic Instruct Thai v1 (F+C+D+) This dataset is part of the research paper "Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai" submitted to ACL SRW 2024. It represents the best-performing synthetic dataset (F+C+D+) generated using our novel seed-free framework for low-resource languages, specifically Thai. Dataset Details Size: 5,000 instructions Language: Thai Task: Instruction-tuning for Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/parinzee/seed-free-synthetic-instruct-thai-v1.texttext-generation1K<n<10K3 likes64 downloads1y agoHugging Face14Suraponn /thai_instruction_sfttext100K<n<1M1 likes60 downloads10mo agoHugging Face15ZombitX64 /Medical-o1-Reasoning-SFT-Thai Medical-GPT-Reasoning-Thai Dataset Summary This dataset contains medical Q&A data in JSON format, designed for fine-tuning AI models in medical reasoning and response generation.representing a medical question, complex chain-of-thought reasoning, and a concise response. All content is in Thai language. The dataset is derived from a larger medical Q&A collection and has been processed to ensure JSON validity, with multi-line objects combined into single valid entries.… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-Reasoning-SFT-Thai.texttext-generation10K<n<100K2 likes58 downloads1y agoHugging Face16openthaigpt /thai-qa-rag-answer-dataset Thai QA RAG Answer Synthesis Dataset Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined Rows: 9999 rows. Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th) Examples {"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.textquestion-answering1K<n<10K0 likes52 downloads3y agoHugging Face17iapp /Thai-R1-Distill-SFT Thai R1 Distill SFT Thai Reasoning Dataset for Supervised Finetuning Translated by iApp Technology textquestion-answering10K<n<100K3 likes51 downloads2y agoHugging Face18openthaigpt /thai-wiki-summary-dataset Thai Wiki Summary Dataset Rows: 3,000 rows (Cleaned) Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th) Examples {"input":"หน่วยพื้นฐานในการแบ่งเขตแดนในโปแลนด์คือ เทศบาล (กมินา) เมืองก็เป็นเทศบาลด้วยเช่นกัน ทว่ามีตราตั้งให้เป็นเมือง ทั้งเมืองและเทศบาลปกครองโดยนายกเทศมนตรี ทว่าในเทศบาล นายกเทศมนตรีเรียกว่าโวกต์ ( วอยต์ในภาษาโปแลนด์) ส่วนในเมืองเรียกว่าเบอร์มิสตร์ ในเมืองใหญ่ ๆ บางเมืองมีความรับผิดชอบและอำนาจพิเศษ… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-wiki-summary-dataset.textquestion-answering1K<n<10K0 likes46 downloads1y agoHugging Face19ZombitX64 /Wisesight-Sentiment-Thai Dataset Card for Zombitx64 Sentiment Corpus Thai This dataset card describes the "Zombitx64 Sentiment Corpus Thai," a large-scale, manually curated corpus for Thai sentiment analysis and token classification. Dataset Details Dataset Description The Zombitx64 Sentiment Corpus Thai is a collection of Thai-language social media comments and posts, annotated for sentiment at the sentence or token level. The dataset covers a wide range of topics and emotional… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Wisesight-Sentiment-Thai.texttoken-classification100K<n<1M2 likes45 downloads1y agoHugging Face20NaNoBotCo /thai-place-name-romanisation-benchmark Thai romanisation, scored on the words place names are made of A romaniser can do well on dictionary words and still misread a road sign. This dataset scores one engine on both, so the difference is a number rather than an impression. Every distinct pure-Thai word in pythainlp/thai-romanization-dataset is romanised by translit.py, the rule-based RTGS engine behind motdang.net, twice: with its curated exception lexicon and with the rules alone. Results slice… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-place-name-romanisation-benchmark.text100K<n<1M0 likes45 downloads17d agoHugging Face21openthaigpt /thai-qa-multiturn-answer-dataset Thai QA Multiturns Answer Synthesis Dataset Rows: 11,992 rows (Cleaned) Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th) Examples {"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}]", "output": "สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ"} {"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}, {\"assistant\": \"สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ\"}, {\"human\":… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-multiturn-answer-dataset.textquestion-answering10K<n<100K0 likes43 downloads1y agoHugging Face22ThaiLLM-Leaderboard /mt-bench-thai MT-Bench Thai MT-Bench Thai is a dataset for multi-turn benchmarking that covers 9 categories. Writing Roleplay Extraction Reasoning Math Coding STEM Social Science Knowledge III We introduce the final category, Knowledge III, which evaluates understanding of Thai cultural context. Dataset Loading from datasets import load_dataset ds = load_dataset("ThaiLLM-Leaderboard/mt-bench-thai") print(ds) output DatasetDict({ train: Dataset({ features: ['question_id'… See the full description on the dataset page: https://huggingface.co/datasets/ThaiLLM-Leaderboard/mt-bench-thai.texttext-generationn<1K7 likes43 downloads1y agoHugging Face23MRlionman6 /thai-traffic-law-qatextn<1K0 likes41 downloads8d agoHugging Face24NaNoBotCo /thai-rtgs-romanisation-lexicon Thai RTGS romanisation — the exception lexicon 694 Thai forms whose romanisation cannot be derived from the spelling, each with the reading a rule-based romaniser should use instead. This is the companion to a rule engine, not a replacement for one. The rules handle the regular cases; this names the irregulars. Measured: it takes the romaniser behind motdang.net from 77.51% to 79.62% agreement on 153,948 Thai words, and from 83.45% to 84.22% on the 21,425 of them that appear in… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-rtgs-romanisation-lexicon.textn<1K0 likes40 downloads17d agoHugging Face25NaNoBotCo /thai-trade-words-chiang-mai คำบนป้าย — Thai trade words of Chiang Mai and Chiang Rai 1,782 Thai trade terms taken from the tags on business listings in Chiang Mai and Chiang Rai — the words the city uses for what a shop does — each with an RTGS reading and an English gloss. Plus 403 administrative place names (tambon, amphoe, city, province) with their readings. ซ่อมมอเตอร์ไซค์ Som Motoesai motorcycle repair 191 places ตู้น้ำดื่มหยอดเหรียญ Tunam Duem Yotrian coin-op drinking-water… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-trade-words-chiang-mai.text1K<n<10K0 likes36 downloads17d agoHugging Face26StelleX /gs8k_thai_r1_example Additional Information This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes: A mathematical problem statement A detailed step-by-step solution An improvement history showing how the solution was iteratively refined Think in English, Context Thai for improve thai question texttext-generation1K<n<10K0 likes35 downloads2y agoHugging Face27monoboard /thai-land-tax-full-triplets Thai Land & Buildings Tax — Full Triplet Dataset Dataset Name: monoboard/thai-land-tax-full-triplets Language: Thai (th) Tasks: Legal Retrieval • RAG • Contrastive Learning • Triplet Loss • Embedding Training Overview This dataset provides a legally verified retrieval corpus for training Thai-language retrieval models under the Land and Buildings Tax Act (B.E. 2562) and related regulations. Each example includes: A legal question (query) One or more oracle passages (pos)… See the full description on the dataset page: https://huggingface.co/datasets/monoboard/thai-land-tax-full-triplets.text1K<n<10K0 likes33 downloads10mo agoHugging Face28Phettae /thai-multitask-starter Thai Multitask 9.6K ชุดข้อมูลตั้งต้นสำหรับ instruction tuning ภาษาไทย ครอบคลุมงานสนทนา ถาม–ตอบ สรุป แปล จำแนกข้อความ ตรวจแก้ภาษา คณิตศาสตร์ และ structured output ข้อมูลทุกแถวสร้างขึ้นใหม่ด้วยกฎแบบ deterministic ไม่มีการคัดลอกจากเว็บไซต์หรือ ข้อมูลส่วนบุคคลจริง เหมาะสำหรับทดลอง supervised fine-tuning และทดสอบ pipeline แต่ควรเพิ่มข้อมูลที่มนุษย์ตรวจทานและข้อมูลภาษาธรรมชาติก่อนใช้กับระบบจริง จำนวนข้อมูลทั้งหมด 9,599 ตัวอย่าง: train 8,639, validation 480 และ test 480… See the full description on the dataset page: https://huggingface.co/datasets/Phettae/thai-multitask-starter.texttext-generation1K<n<10K0 likes31 downloads1mo agoHugging Face29saksornr /sql-create-context-thai Overview This dataset builds from sql-create-context. @misc{b-mc2_2023_sql-create-context, title = {sql-create-context Dataset}, author = {b-mc2}, year = {2023}, url = {https://huggingface.co/datasets/b-mc2/sql-create-context}, note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.}, } texttext-generation10K<n<100K0 likes30 downloads2y agoHugging Face30Mularstyle /thai-name-spell-v3text1K<n<10K0 likes30 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.