CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Wanfq /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.tabularquestion-answering1K<n<10K0 likes2.4k downloads2y agoHugging Face02ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face03WangResearchLab /SteeringSafety SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs 🎯 Overview SteeringSafety is a benchmark suite for evaluating representation steering methods across multiple safety perspectives. SteeringSafety provides: 📊 A collection of 17 datasets including 7 perspectives for measuring safety behaviors. 🔧 A modular code framework implementing the taxonomy of training-free steering methods with standardized, interchangeable… See the full description on the dataset page: https://huggingface.co/datasets/WangResearchLab/SteeringSafety.tabulartext-classification10K<n<100K4 likes1.4k downloads10mo agoHugging Face04aisingapore /WangchanLION-Web Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, } We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.texttext-generation10M<n<100M3 likes777 downloads1y agoHugging Face05wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes390 downloads3y agoHugging Face06airesearch /WangchanThaiInstructWangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai (EMNLP'25) WangchanThaiInstruct is a human-authored Thai dataset that improves instruction-following in low-resource settings, capturing cultural and domain-specific nuances across four domains and seven task types. The evaluate code can be found at this github link @inproceedings{limkonchotiwat2025thaiinstruct, title = {WangchanThaiInstruct: An Instruction-Following… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanThaiInstruct.texttext-generation10K<n<100K29 likes390 downloads9mo agoHugging Face07wannaphong /wikipedia-monthly 🚀 Wikipedia Monthly Last updated: March 14, 2026, 21:06 UTC This repository provides monthly, multilingual dumps of Wikipedia, processed and prepared for easy use in NLP projects. 📊 Current Statistics Metric Current Export (March 2026) All Exports (Total) Languages 343 361 Articles 62.8M 62.8M Usage Load any language with a single line of code using 🤗 datasets. latest always refers to the most recent dump, while dated configs refer to… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/wikipedia-monthly.texttext-generation1M<n<10M0 likes350 downloads5mo agoHugging Face08wangzx1210 /OmniEAR OmniEAR Expert Trajectory Dataset Dataset Summary The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.texttext-generation10K<n<100K10 likes301 downloads1y agoHugging Face09airesearch /WangchanX-Legal-ThaiCCL-RAG 🏛️ WangchanX-Legal-ThaiCCL-RAG [Technical Report] The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.texttext-generation10K<n<100K13 likes280 downloads2y agoHugging Face10aisingapore /WangchanLION-Curated Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, } We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Curated.textfill-mask100K<n<1M4 likes239 downloads1y agoHugging Face11wannaphong /KhanomTanLLM-pretrained-dataset KhanomTanLLM pretrained dataset This daataset collect all raw text for pretraining LLM. Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Tokens 53,376,211,711 Tokens English: 31,629,984,243 Tokens Thai: 12,785,565,497 Tokens Code: 8,913,084,300 Toekns Parallel data: 190,310,686 Tokens Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer All subset Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.texttext-generation10M<n<100M1 likes231 downloads2y agoHugging Face12Wannita /PyCoder PyCoder This repository contains the dataset for the paper Syntax-Aware On-the-Fly Code Completion The sample code to run the model can be found in directory: "assets/notebooks/inference.ipynb" in our GitHub: https://github.com/awsm-research/pycoder. PyCoder is an auto code completion model which leverages a Multi-Task Training technique (MTT) to cooperatively learn the code prediction task and the type prediction task. For the type prediction task, we propose to leverage the… See the full description on the dataset page: https://huggingface.co/datasets/Wannita/PyCoder.texttext-generation100K<n<1M0 likes163 downloads3y agoHugging Face13wangxj123 /LumiRAG LumiRAG: A Unified Multimodal RAG Large Model Bridging Text and Image Retrieval Dataset Description This is the dataset used in the paper "LumiRAG: A Unified Multimodal RAG Large Model Bridging Text and Image Retrieval". Languages English Data Fields { "question" "answer" "imagepath" "reward_method" "data_source" } imagequestion-answering100K<n<1M0 likes161 downloads7mo agoHugging Face14rose-e-wang /bridgeTLDR: This dataset is a real-world math tutoring dataset from the NAACL 2024 paper ``Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes''. The dataset targets scenarios where the student makes a math mistake. c_h is the conversation history c_r is the original tutor's response c_r_ is the experienced teacher's response Optionally, there is other interesting metadata from our Bridge method: e is the student error type that the experienced… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/bridge.texttext-generationn<1K5 likes134 downloads2y agoHugging Face15wanhaoliu /PolyReal PolyReal: A Benchmark for Real-World Polymer Science Workflows Evaluating Multimodal Large Language Models Across the Full Lifecycle of Polymer Science Contents data/train.jsonl: Hugging Face viewer-friendly training split. PolyReal.json: main dataset file. ref/: referenced images and CSV files used by Path entries in the dataset. 🔥 Overview PolyReal is a multimodal benchmark for real-world polymer science workflows.It is… See the full description on the dataset page: https://huggingface.co/datasets/wanhaoliu/PolyReal.imagevisual-question-answeringn<1K4 likes121 downloads3mo agoHugging Face16wannabeyourfriend-hf /mind2dialogue Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States This dataset contains synthetic conversations between a simulated user and an Oracle assistant. Both use the same evolving user state. Paper · Code · Project page from datasets import load_dataset dataset = load_dataset( "wannabeyourfriend-hf/mind2dialogue", "conversations", split="train" ) Citation @misc{wang2026mind2dialoguetraininghumanawarelanguage… See the full description on the dataset page: https://huggingface.co/datasets/wannabeyourfriend-hf/mind2dialogue.texttext-generation10K<n<100K1 likes117 downloads8d agoHugging Face17wangekxy /classical-tcm-canon Classical Chinese Medicine Canon — 中医经典文本数据集 (v1) A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon), 难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics — assembled from public-domain source works. Useful for LLM training, RAG, and search over classical Chinese medicine / 中医药 literature. Summary 115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.tabulartext-generationn<1K0 likes104 downloads3mo agoHugging Face18trjxter /Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB is a W&B-only synthetic reasoning dataset generated with Kimi-K2.6 through Weights & Biases Inference. This dataset is the pure W&B-generated subset from a larger planned 8,000-example Kimi reasoning distillation run. Generation stopped when the W&B quota limit was reached, and the completed accepted rows were audited, cleaned, and exported as a standalone dataset. This release contains 3,303 accepted W&B-generated rows… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Reasoning-3300x-WandB.texttext-generation1K<n<10K7 likes77 downloads4mo agoHugging Face19wangbing1416 /MSMS-LIMO-v2-SFT A Multi-Source Multi-Solution Long CoT SFT Dataset from LIMO-v2 texttext-generation100K<n<1M1 likes73 downloads9mo agoHugging Face20wannaphong /KhanomTanLLM-pretrained-dataset-thai-subset KhanomTanLLM pretrained dataset (Thai subset) This daataset collect all raw text for pretraining LLM. (Thai subset) Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0 pythainlp/thai-tnhc2-books pythainlp/thai-constitution-corpus pythainlp/thai-it-books pythainlp/prd_news_3011202 pythainlp/thailand-policy-statements pythainlp/thai-cc-license pythainlp/blognone_news pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.texttext-generation10M<n<100M0 likes68 downloads2y agoHugging Face21wangbing1416 /LonsRex-Misinfomation-SFTtexttext-generation100K<n<1M1 likes66 downloads4mo agoHugging Face22wannaphong /typhoon-s-instruct-post-training Typhoon-S Instruct Post-Training Dataset Summary This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths. The dataset follows a two-part mixture philosophy: Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-instruct-post-training.texttext-generation100K<n<1M0 likes56 downloads6mo agoHugging Face23wannaphong /typhoon-s-sovereign-capability-dataset Typhoon-S Training Assets Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project. Datasets NitiBench (Legal Domain) nitibench_train_rl.parquet - RL training set (8,211 examples) nitibench_train_pretrain.parquet - Pretrain set (3,648 examples) nitibench_train_sft.parquet - SFT set (3,648 examples) nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.texttext-generation10K<n<100K0 likes52 downloads6mo agoHugging Face24Wannita /PyCoder-Type PyCoder This repository contains the dataset for the paper Syntax-Aware On-the-Fly Code Completion The sample code to run the model can be found in directory: "assets/notebooks/inference.ipynb" in our GitHub: https://github.com/awsm-research/pycoder. PyCoder is an auto code completion model which leverages a Multi-Task Training technique (MTT) to cooperatively learn the code prediction task and the type prediction task. For the type prediction task, we propose to leverage the… See the full description on the dataset page: https://huggingface.co/datasets/Wannita/PyCoder-Type.texttext-generation100K<n<1M0 likes48 downloads3y agoHugging Face25wangmingyuan /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例… See the full description on the dataset page: https://huggingface.co/datasets/wangmingyuan/HundredCV-Chat.texttext-generation10K<n<100K0 likes48 downloads26d agoHugging Face26airesearch /wangchanx-seed-free-synthetic-instruct-thai-120k Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k Dataset Summary This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.tabulartext-generation100K<n<1M3 likes46 downloads2y agoHugging Face27wangzihaogithub /job-educational-parser-dataset-08-0-0805 Job Educational Parser Dataset 招聘领域的岗位与学历要求数据集。 输入:岗位描述 -> 输出:学历要求 Splits train: 19w_0701.csv (约 19 万条) test: 2w_0716.csv (约 2 万条) validation: 4w_0708.csv (约 4 万条) 每条数据至少包含字段: user: 职位描述 assistant: 要求的学历(如 "博士、硕士、本科"),遵循从高到低 由 @wangzihaogithub 创建。 tabulartext-generation100K<n<1M0 likes46 downloads1y agoHugging Face28wanwan1212 /Futurex-Past FutureX-Past 📜 Overview This repository contains a dataset of past questions from the FutureX benchmark. FutureX is a live, dynamic benchmark designed to evaluate the future prediction capabilities of Large Language Model (LLM) agents. It features a fully automated pipeline that generates new questions about upcoming real-world events, deploys agents to predict their outcomes, and scores the results automatically. For more information on the live benchmark, please refer… See the full description on the dataset page: https://huggingface.co/datasets/wanwan1212/Futurex-Past.tabularquestion-answeringn<1K0 likes46 downloads8mo agoHugging Face29wannaphong /thai_synthetic_mathtexttext-generationn<1K0 likes45 downloads2mo agoHugging Face30wannaphong /clean_thaiwikipediatexttext-generation100K<n<1M0 likes44 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.