CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Wanfq /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.tabularquestion-answering1K<n<10K0 likes2.4k downloads2y agoHugging Face02ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face03WangResearchLab /SteeringSafety SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs 🎯 Overview SteeringSafety is a benchmark suite for evaluating representation steering methods across multiple safety perspectives. SteeringSafety provides: 📊 A collection of 17 datasets including 7 perspectives for measuring safety behaviors. 🔧 A modular code framework implementing the taxonomy of training-free steering methods with standardized, interchangeable… See the full description on the dataset page: https://huggingface.co/datasets/WangResearchLab/SteeringSafety.tabulartext-classification10K<n<100K4 likes1.2k downloads10mo agoHugging Face04aisingapore /WangchanLION-Web Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, } We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.texttext-generation10M<n<100M3 likes736 downloads1y agoHugging Face05airesearch /WangchanThaiInstructWangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai (EMNLP'25) WangchanThaiInstruct is a human-authored Thai dataset that improves instruction-following in low-resource settings, capturing cultural and domain-specific nuances across four domains and seven task types. The evaluate code can be found at this github link @inproceedings{limkonchotiwat2025thaiinstruct, title = {WangchanThaiInstruct: An Instruction-Following… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanThaiInstruct.texttext-generation10K<n<100K29 likes402 downloads9mo agoHugging Face06wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes381 downloads3y agoHugging Face07wannaphong /wikipedia-monthly 🚀 Wikipedia Monthly Last updated: March 14, 2026, 21:06 UTC This repository provides monthly, multilingual dumps of Wikipedia, processed and prepared for easy use in NLP projects. 📊 Current Statistics Metric Current Export (March 2026) All Exports (Total) Languages 343 361 Articles 62.8M 62.8M Usage Load any language with a single line of code using 🤗 datasets. latest always refers to the most recent dump, while dated configs refer to… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/wikipedia-monthly.texttext-generation1M<n<10M0 likes324 downloads5mo agoHugging Face08wangzx1210 /OmniEAR OmniEAR Expert Trajectory Dataset Dataset Summary The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.texttext-generation10K<n<100K10 likes290 downloads1y agoHugging Face09airesearch /WangchanX-Legal-ThaiCCL-RAG 🏛️ WangchanX-Legal-ThaiCCL-RAG [Technical Report] The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.texttext-generation10K<n<100K13 likes289 downloads2y agoHugging Face10wannaphong /KhanomTanLLM-pretrained-dataset KhanomTanLLM pretrained dataset This daataset collect all raw text for pretraining LLM. Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Tokens 53,376,211,711 Tokens English: 31,629,984,243 Tokens Thai: 12,785,565,497 Tokens Code: 8,913,084,300 Toekns Parallel data: 190,310,686 Tokens Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer All subset Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.texttext-generation10M<n<100M1 likes235 downloads2y agoHugging Face11aisingapore /WangchanLION-Curated Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, } We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Curated.textfill-mask100K<n<1M4 likes232 downloads1y agoHugging Face12Wannita /PyCoder PyCoder This repository contains the dataset for the paper Syntax-Aware On-the-Fly Code Completion The sample code to run the model can be found in directory: "assets/notebooks/inference.ipynb" in our GitHub: https://github.com/awsm-research/pycoder. PyCoder is an auto code completion model which leverages a Multi-Task Training technique (MTT) to cooperatively learn the code prediction task and the type prediction task. For the type prediction task, we propose to leverage the… See the full description on the dataset page: https://huggingface.co/datasets/Wannita/PyCoder.texttext-generation100K<n<1M0 likes166 downloads3y agoHugging Face13opendatalab /WanJuan-Korean 💡 Introduction WanJuan-Korean(万卷丝路-韩语) corpus, with a volume exceeding 280GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Korean.text-generation2 likes161 downloads1y agoHugging Face14wangxj123 /LumiRAG LumiRAG: A Unified Multimodal RAG Large Model Bridging Text and Image Retrieval Dataset Description This is the dataset used in the paper "LumiRAG: A Unified Multimodal RAG Large Model Bridging Text and Image Retrieval". Languages English Data Fields { "question" "answer" "imagepath" "reward_method" "data_source" } imagequestion-answering100K<n<1M0 likes155 downloads7mo agoHugging Face15rose-e-wang /bridgeTLDR: This dataset is a real-world math tutoring dataset from the NAACL 2024 paper ``Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes''. The dataset targets scenarios where the student makes a math mistake. c_h is the conversation history c_r is the original tutor's response c_r_ is the experienced teacher's response Optionally, there is other interesting metadata from our Bridge method: e is the student error type that the experienced… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/bridge.texttext-generationn<1K5 likes128 downloads2y agoHugging Face16wangrongsheng /step3.5-flash-sft Step-3.5-Flash-SFT Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models. This repository keeps the full training interface in one place: json/: canonical raw training data tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment compiled/: tokenizer-specific compiled shards for StepTronOSS training Data Format Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/wangrongsheng/step3.5-flash-sft.text-generation0 likes128 downloads6mo agoHugging Face17wannabeyourfriend-hf /mind2dialogue Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States This dataset contains synthetic conversations between a simulated user and an Oracle assistant. Both use the same evolving user state. Paper · Code · Project page from datasets import load_dataset dataset = load_dataset( "wannabeyourfriend-hf/mind2dialogue", "conversations", split="train" ) Citation @misc{wang2026mind2dialoguetraininghumanawarelanguage… See the full description on the dataset page: https://huggingface.co/datasets/wannabeyourfriend-hf/mind2dialogue.texttext-generation10K<n<100K2 likes124 downloads7d agoHugging Face18opendatalab /WanJuan-Thai 💡 Introduction WanJuan-Thai (万卷丝路-泰语) corpus, with a volume exceeding 155GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Thai.text-generation3 likes123 downloads1y agoHugging Face19wanhaoliu /PolyReal PolyReal: A Benchmark for Real-World Polymer Science Workflows Evaluating Multimodal Large Language Models Across the Full Lifecycle of Polymer Science Contents data/train.jsonl: Hugging Face viewer-friendly training split. PolyReal.json: main dataset file. ref/: referenced images and CSV files used by Path entries in the dataset. 🔥 Overview PolyReal is a multimodal benchmark for real-world polymer science workflows.It is… See the full description on the dataset page: https://huggingface.co/datasets/wanhaoliu/PolyReal.imagevisual-question-answeringn<1K4 likes115 downloads3mo agoHugging Face20wangekxy /classical-tcm-canon Classical Chinese Medicine Canon — 中医经典文本数据集 (v1) A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon), 难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics — assembled from public-domain source works. Useful for LLM training, RAG, and search over classical Chinese medicine / 中医药 literature. Summary 115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.tabulartext-generationn<1K0 likes113 downloads3mo agoHugging Face21opendatalab /WanJuan-BaiHua OpenDataLab近期拟发布万卷·百华大规模专业领域数据集,主要面向金融、能源、文化教育、政务、通信、交通运输、医疗健康、汽车、烟草、计算机等领域大模型训练,提供高质量、精细处理、领域分类的多模态专用语料。 在此我们将面向社区征集需求,若有相关行业数据集的需求,请填写以下调查问卷,我们会根据社区反馈来决定不同领域数据集的发布顺序,欢迎大家提供想法! 调查问卷链接:https://www.wjx.cn/vm/mxip3eE.aspx# 数据集样例详见:https://opendatalab.com/OpenDataLab/WanJuan-BaiHua text-generation4 likes98 downloads1y agoHugging Face22wangrongsheng /HealthCoreBench-preview imagetext-generation10M<n<100M0 likes87 downloads2mo agoHugging Face23trjxter /Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB is a W&B-only synthetic reasoning dataset generated with Kimi-K2.6 through Weights & Biases Inference. This dataset is the pure W&B-generated subset from a larger planned 8,000-example Kimi reasoning distillation run. Generation stopped when the W&B quota limit was reached, and the completed accepted rows were audited, cleaned, and exported as a standalone dataset. This release contains 3,303 accepted W&B-generated rows… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Reasoning-3300x-WandB.texttext-generation1K<n<10K7 likes80 downloads4mo agoHugging Face24opendatalab /WanJuan-Arabic 💡 Introduction WanJuan-Arabic(万卷丝路-阿拉伯语) corpus, with a volume exceeding 220GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Arabic.text-generation2 likes73 downloads1y agoHugging Face25wannaphong /typhoon-s-instruct-post-training Typhoon-S Instruct Post-Training Dataset Summary This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths. The dataset follows a two-part mixture philosophy: Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-instruct-post-training.texttext-generation100K<n<1M0 likes73 downloads6mo agoHugging Face26wangbing1416 /LonsRex-Misinfomation-SFTtexttext-generation100K<n<1M1 likes70 downloads4mo agoHugging Face27wangbing1416 /MSMS-LIMO-v2-SFT A Multi-Source Multi-Solution Long CoT SFT Dataset from LIMO-v2 texttext-generation100K<n<1M1 likes69 downloads9mo agoHugging Face28wannaphong /typhoon-s-sovereign-capability-dataset Typhoon-S Training Assets Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project. Datasets NitiBench (Legal Domain) nitibench_train_rl.parquet - RL training set (8,211 examples) nitibench_train_pretrain.parquet - Pretrain set (3,648 examples) nitibench_train_sft.parquet - SFT set (3,648 examples) nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.texttext-generation10K<n<100K0 likes67 downloads6mo agoHugging Face29wannaphong /KhanomTanLLM-pretrained-dataset-thai-subset KhanomTanLLM pretrained dataset (Thai subset) This daataset collect all raw text for pretraining LLM. (Thai subset) Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0 pythainlp/thai-tnhc2-books pythainlp/thai-constitution-corpus pythainlp/thai-it-books pythainlp/prd_news_3011202 pythainlp/thailand-policy-statements pythainlp/thai-cc-license pythainlp/blognone_news pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.texttext-generation10M<n<100M0 likes64 downloads2y agoHugging Face30Wannita /PyCoder-Type PyCoder This repository contains the dataset for the paper Syntax-Aware On-the-Fly Code Completion The sample code to run the model can be found in directory: "assets/notebooks/inference.ipynb" in our GitHub: https://github.com/awsm-research/pycoder. PyCoder is an auto code completion model which leverages a Multi-Task Training technique (MTT) to cooperatively learn the code prediction task and the type prediction task. For the type prediction task, we propose to leverage the… See the full description on the dataset page: https://huggingface.co/datasets/Wannita/PyCoder-Type.texttext-generation100K<n<1M0 likes50 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.