CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads2h agoHugging Face02elmoghany /Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text Dataset Overview A collection of 27 domains (“topics”) and 3100 question-answer pair. Each topic comes with average 117 QA pairs.Every QA entry comes with: references: one or more source files the answer is extracted from time with each reference comes the starting and ending time the answer is extracted from the reference video_files: the video files where the answer can be found (future) video title & description from metadata.csv File structure You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.question-answering1K<n<10K3 likes2.2k downloads1y agoHugging Face03matichon /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.textquestion-answering1K<n<10K0 likes2.2k downloads5mo agoHugging Face04typhoon-ai /thai_exam Dataset Card for Thai_Exam ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows: ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.tabularquestion-answeringn<1K19 likes2.2k downloads2y agoHugging Face05openthaigpt /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.textquestion-answering1K<n<10K8 likes1.6k downloads3y agoHugging Face06iapp /MMMU-Thai MMMU Thai (MMMU Benchmark Translated to Thai) MMMU Thai is a dataset for evaluating multimodal models on massive multi-discipline tasks requiring college-level knowledge and deliberate reasoning. This dataset is translated from MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) into Thai. Dataset Details MMMU Thai consists of 11,500 meticulously collected multimodal questions from college exams, quizzes, and textbooks… See the full description on the dataset page: https://huggingface.co/datasets/iapp/MMMU-Thai.imagequestion-answering10K<n<100K2 likes669 downloads2y agoHugging Face07SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes515 downloads4mo agoHugging Face08RJTPP /thai_exam-reformattedReformatted version of scb10x/thai_exam Additional Changes: Fix math incorrect answer ถ้า \log_{\frac{1}{4}} 256 + \frac{2\log 625}{\log 5} = 3^a เมื่อ a เป็นจำนวนจริง แล้วคำตอบของ a เท่ากับเท่าใด? a. \log_{3} 2 b. \log_{3} 4 c. \log_{3} \frac{33}{4} d. \log_{3} 10 e. \log_{3} 12 # Original answer: d (\log_{3} 10) # Correction : b (\log_{3} 4) textquestion-answering1K<n<10K0 likes369 downloads1y agoHugging Face09Thaweewat /onet-m6-social Summary This is a question-answer dataset for the Grade 12 (M6) Social subject of the Thailand Ordinary National Educational Test (ONET). The dataset was human-extracted by my team from the official release of publicly available exams National Institute of Educational Testing Service during the years 2016-2022. The exam consists of 510 multiple-choice questions with corresponding answer keys. It is important to note that only two questions, Q71 and Q85, from the year 2018, require… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/onet-m6-social.textquestion-answeringn<1K2 likes241 downloads3y agoHugging Face10pythainlp /thaiqa_squad`thaiqa_squad` is an open-domain, extractive question answering dataset (4,000 questions in `train` and 74 questions in `dev`) in [SQuAD](https://rajpurkar.github.io/SQuAD-explorer/) format, originally created by [NECTEC](https://www.nectec.or.th/en/) from Wikipedia articles and adapted to [SQuAD](https://rajpurkar.github.io/SQuAD-explorer/) format by [PyThaiNLP](https://github.com/PyThaiNLP/).question-answering1K<n<10K12 likes208 downloads3y agoHugging Face11samuelandaudreymedianetwork /that-backpacker-article-corpus That Backpacker Article Corpus This dataset contains a structured corpus of long-form travel articles published on ThatBackpacker.com, authored primarily by Audrey Bergner as part of the Samuel & Audrey Media Network. The corpus includes 323 article records covering destination guides, multi-day itineraries, hiking, food travel, cultural experiences, city guides, transportation, accommodations, and practical travel planning. It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/that-backpacker-article-corpus.texttext-generation100K<n<1M1 likes157 downloads4mo agoHugging Face12biodatlab /thai_buddhist_studies_exam Thai Buddhist Studies Examination (Nak Tham) This repository contains multiple-choice questions from the Thai Buddhist Studies (Nak Tham) examination (2020, 2022, 2023). This dataset can be used for a benchmark for evaluating Large Language Models' understanding of Thai Buddhist concepts and teachings. Dataset Statistics Year Number of Multiple Choice Questions 2020 1,350 2022 1,400 2023 1,350 Phra Udom thought on the exam: We have reviewed the Nak… See the full description on the dataset page: https://huggingface.co/datasets/biodatlab/thai_buddhist_studies_exam.textquestion-answering1K<n<10K5 likes141 downloads2y agoHugging Face13Thaweewat /alpaca-cleaned-52k-th Summary This is a Thai 🇹🇭-instructed dataset translated from cleaned version of the original Alpaca Dataset released by Stanford using Google Cloud Translation, contain 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. The following issues have been identified in the original release and fixed in this… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/alpaca-cleaned-52k-th.textquestion-answering10K<n<100K17 likes140 downloads3y agoHugging Face14EIRTHAIMED /ThaiExamFinetune Multiple-Choice Question Dataset in Thai Details: This dataset is designed for fine-tuning Thai language models, focusing on the Chain-of-Thought (COT) process, which aids in analyzing questions and deriving correct answers step by step. The dataset consists of multiple-choice questions divided into five categories: O-NET: Ordinary National Educational Test IC: Investment Consultant TGAT: Thai General Aptitude Test TPAT: Thai Professional Aptitude Test… See the full description on the dataset page: https://huggingface.co/datasets/EIRTHAIMED/ThaiExam.textquestion-answering1K<n<10K0 likes134 downloads2y agoHugging Face15adsgpt /marketing-benchmark-of-more-than-10-ai-models Marketing Benchmark of 10+ AI Models A 5,000-question benchmark for evaluating LLMs across six dimensions of modern marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas. Every question is independently authored by the AdsGPT Marketing Bench team. Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.textquestion-answering1K<n<10K5 likes121 downloads4mo agoHugging Face16openthaigpt /thai-investment-consultant-licensing-exams Thai Public Investment Consultant (IC) Exams Dataset Overview This dataset comprises a collection of exam questions and answers from the Thai Public Investment Consultant (IC) Examinations. It's a valuable resource for developing and evaluating question-answering systems in the finance sector. Dataset Source The Stock Exchange of Thailand (SET) Maintainer Dr. Kobkrit Viriyayudhakorn Email: kobkrit@iapp.co.th Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-investment-consultant-licensing-exams.tabularquestion-answeringn<1K6 likes111 downloads3y agoHugging Face17tharustack /open-math-dataset Dataset Description Open Math Dataset is an open mathematics corpus designed for mathematical AI, reasoning, education, and research. The project is being developed from Sri Lanka with the goal of creating an internationally useful mathematics dataset for developers, researchers, educators, and AI systems. Mathematics Corpus The dataset is designed to contain structured mathematical problems and solutions across different mathematical domains and education levels.… See the full description on the dataset page: https://huggingface.co/datasets/tharustack/open-math-dataset.text-generation1K<n<10K1 likes104 downloads1mo agoHugging Face18ThaiSyntheticQA /ThaiQA-v1 ThaiQA v1 ThaiQA v1 is a Thai Synthetic QA dataset. It was created from synthetic method using open source LLM in Thai language. We used Nvidia Nemotron 4 (340B) to create this dataset. Topics: Technology and Gadgets 100 Travel and Tourism 91 Food and Cooking 99 Sports and Fitness 50 Arts and Entertainment 24 Home and Garden 72 Fashion and Beauty 99 Science and Nature 100 History and Culture 91 Education and Learning 99 Pets and Animals 83 Relationships and Family 78 Personal… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/ThaiQA-v1.texttext-generation10K<n<100K5 likes98 downloads2y agoHugging Face19typhoon-ai /ThaiSafetyBench ThaiSafetyBench ⚠️ Warning: This dataset contains harmful and toxic language. It is intended for academic purposes only. [ArXiv Paper] [Github] [Hugging Face Leaderboard 🤗] The ThaiSafetyBench dataset comprises 1,889 malicious Thai-language prompts across various categories. In addition to translated malicious prompts, it includes prompts tailored to Thai culture, offering deeper insights into culturally specific attacks. Note: The Monarchy type of harm has been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiSafetyBench.textquestion-answering1K<n<10K5 likes97 downloads7mo agoHugging Face20limberc /this-that-spatial-bench spatial-decisions 7,305 multiple-choice decision questions over 6,525 distinct simulated states, in 15 families and two environments. Every answer is computed from the simulator, not annotated by a person and not taken from a model. That is the point of the set: on a question whose answer is derived from the rules of the environment, a disagreement is a mistake, and there is nothing to argue about. The set was built to replace a much narrower public artefact: a recording of 68… See the full description on the dataset page: https://huggingface.co/datasets/limberc/this-that-spatial-bench.textquestion-answering1K<n<10K0 likes92 downloads5d agoHugging Face21ThakiCloud /korean-contract-derivation-bench Korean Contract Derivation Benchmark A model that derives a discount and a unit-price total without being told to will still return the VAT-inclusive figure verbatim when asked for the contract amount — and a generic instruction to "compute the relation" does not fix it. This benchmark isolates that. The finding 240 synthetic Korean procurement summaries. Half state the contract amount directly among distractor figures (disambiguate); half state it only as a… See the full description on the dataset page: https://huggingface.co/datasets/ThakiCloud/korean-contract-derivation-bench.question-answeringn<1K0 likes84 downloads17d agoHugging Face22Kongongong /Thai-Physics-Data-40KThai-Physics-Data is a Thai-Based physics data with more than 40k lines of data. Data Sources: ArtifactAI/arxiv-physics-instruct-tune-30k (CC BY-NC 2.0) camel-ai/physics How to load Data (Hugging Face) from datasets import load_dataset Thai_Physics_Data = load_dataset("Kongongong/Thai-Physics-Data-40K") Thai_Physics_Data = Thai_Physics_Data['train'] def format_data(): .... data =[] format_data() data = Dataset.from_dict({"text": data}) How to load Data (CSV) from… See the full description on the dataset page: https://huggingface.co/datasets/Kongongong/Thai-Physics-Data-40K.textquestion-answering10K<n<100K3 likes80 downloads2y agoHugging Face23thanthienhai /vncompress-vi-v2 VNCompress-VI v2 — query-conditioned context compression for Vietnamese ⚠️ BẢN ĐANG REBUILD — WORK IN PROGRESS. corpus/qa/qa_synthetic đã sạch và ổn định. compression.jsonl mới có 486 hàng thật (prompt v3, trích xuất) trên tổng số dự kiến 100.000+ — quá trình sinh đang tạm dừng vì tỉ lệ drop "không đạt ngân sách token" rất cao (90% ở batch gần nhất) chưa được xử lý gốc rễ. Đừng dùng compression.jsonl để báo cáo kết quả benchmark hay train E5/E6 ở quy mô lớn — số hàng hiện tại… See the full description on the dataset page: https://huggingface.co/datasets/thanthienhai/vncompress-vi-v2.tabularquestion-answering100K<n<1M0 likes66 downloads7d agoHugging Face24thangvip /vietnamese-legal-qa thangvip/vietnamese-legal-qa Dataset Description This dataset contains Vietnamese legal documents with automatically generated question-answer pairs. Each document includes comprehension questions of varying difficulty levels (easy, medium, hard) and types (factual, interpretation, analytical, application). Dataset Structure Data Fields doc_name: Name of the legal document doc_type_name: Type of document (e.g., "Luật" for Law) article_content:… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/vietnamese-legal-qa.textquestion-answering1K<n<10K3 likes63 downloads1y agoHugging Face25amornpan /thai-gov-procurement_regulation-17-amend-21 🇹🇭 Dataset Card for Thai Government Procurement Dataset ℹ️ This dataset is optimized for procurement-related NLP tasks in Thai. This dataset contains a collection of procurement regulations, instructions, and responses focused on public sector purchasing, contract management, and compliance with Thai government standards. It aims to support natural language processing tasks involving procurement assistance, such as chatbot development, procurement dialogue generation… See the full description on the dataset page: https://huggingface.co/datasets/amornpan/thai-gov-procurement_regulation-17-amend-21.textquestion-answeringn<1K2 likes60 downloads2y agoHugging Face26ZombitX64 /Medical-o1-Reasoning-SFT-Thai Medical-GPT-Reasoning-Thai Dataset Summary This dataset contains medical Q&A data in JSON format, designed for fine-tuning AI models in medical reasoning and response generation.representing a medical question, complex chain-of-thought reasoning, and a concise response. All content is in Thai language. The dataset is derived from a larger medical Q&A collection and has been processed to ensure JSON validity, with multi-line objects combined into single valid entries.… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-Reasoning-SFT-Thai.texttext-generation10K<n<100K2 likes58 downloads1y agoHugging Face27openthaigpt /thai-qa-rag-answer-dataset Thai QA RAG Answer Synthesis Dataset Seed dataset is from https://huggingface.co/datasets/Thaweewat/instruct-qa-thai-combined Rows: 9999 rows. Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th) Examples {"input":"ผู้เล่นคนใดทำการอินเตอร์เซปสูงสุดในฤดูกาล","instruction":"ทีมรับของแพนเธอร์สถอดใจที่คะแนน 308 ได้อันดับที่หกของลีก ในขณะที่เป็นผู้นำในเอ็นเอฟแอลด้วยการอินเตอร์เซป 24 ครั้งและได้รับเลือกให้เล่นในโปรโบว์ล สี่ ครั้ง คาวันน์ ชอร์ต… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-rag-answer-dataset.textquestion-answering1K<n<10K0 likes52 downloads3y agoHugging Face28iapp /Thai-R1-Distill-SFT Thai R1 Distill SFT Thai Reasoning Dataset for Supervised Finetuning Translated by iApp Technology textquestion-answering10K<n<100K3 likes51 downloads2y agoHugging Face29thangvip /combined-vietnamese-legal-text thangvip/combined-vietnamese-legal-text Dataset Description This is a combined Vietnamese legal dataset with question-answer pairs formatted in a single text column. It combines two datasets: thangvip/vietnamese-legal-qa (9,715 examples) thangvip/law-reading-comprehension-qa-filtered (205,369 examples) Dataset Structure Data Fields text: Combined text containing legal content followed by question-answer pairs in XML-like format Format… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/combined-vietnamese-legal-text.textquestion-answering100K<n<1M1 likes51 downloads1y agoHugging Face30mspkrai /ThaiBarAssociationEditorial About Thai Bar Under The Royal Patronage "The Thai Bar Association originated from the Law School established under the royal initiative of King Chulalongkorn (Rama V). In 1948, the Thai Bar Association established the Institute of Legal Education with the objective of imparting and promoting legal education and professional expertise in the practice of law. Instruction commenced for the first time in November 1948, with a curriculum modeled after the Council of Legal Education in… See the full description on the dataset page: https://huggingface.co/datasets/mspkrai/ThaiBarAssociationEditorial.text-generation1K<n<10K4 likes51 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.