CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face02aisingapore /WangchanLION-Web Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, } We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.texttext-generation10M<n<100M3 likes777 downloads1y agoHugging Face03wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes390 downloads3y agoHugging Face04wangzx1210 /OmniEAR OmniEAR Expert Trajectory Dataset Dataset Summary The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.texttext-generation10K<n<100K10 likes301 downloads1y agoHugging Face05wannaphong /KhanomTanLLM-pretrained-dataset KhanomTanLLM pretrained dataset This daataset collect all raw text for pretraining LLM. Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Tokens 53,376,211,711 Tokens English: 31,629,984,243 Tokens Thai: 12,785,565,497 Tokens Code: 8,913,084,300 Toekns Parallel data: 190,310,686 Tokens Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer All subset Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.texttext-generation10M<n<100M1 likes231 downloads2y agoHugging Face06rose-e-wang /bridgeTLDR: This dataset is a real-world math tutoring dataset from the NAACL 2024 paper ``Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes''. The dataset targets scenarios where the student makes a math mistake. c_h is the conversation history c_r is the original tutor's response c_r_ is the experienced teacher's response Optionally, there is other interesting metadata from our Bridge method: e is the student error type that the experienced… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/bridge.texttext-generationn<1K5 likes134 downloads2y agoHugging Face07wanhaoliu /PolyReal PolyReal: A Benchmark for Real-World Polymer Science Workflows Evaluating Multimodal Large Language Models Across the Full Lifecycle of Polymer Science Contents data/train.jsonl: Hugging Face viewer-friendly training split. PolyReal.json: main dataset file. ref/: referenced images and CSV files used by Path entries in the dataset. 🔥 Overview PolyReal is a multimodal benchmark for real-world polymer science workflows.It is… See the full description on the dataset page: https://huggingface.co/datasets/wanhaoliu/PolyReal.imagevisual-question-answeringn<1K4 likes121 downloads3mo agoHugging Face08trjxter /Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB Kimi-K2.6-Reasoning-3300x-WandB is a W&B-only synthetic reasoning dataset generated with Kimi-K2.6 through Weights & Biases Inference. This dataset is the pure W&B-generated subset from a larger planned 8,000-example Kimi reasoning distillation run. Generation stopped when the W&B quota limit was reached, and the completed accepted rows were audited, cleaned, and exported as a standalone dataset. This release contains 3,303 accepted W&B-generated rows… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Reasoning-3300x-WandB.texttext-generation1K<n<10K7 likes77 downloads4mo agoHugging Face09wangbing1416 /MSMS-LIMO-v2-SFT A Multi-Source Multi-Solution Long CoT SFT Dataset from LIMO-v2 texttext-generation100K<n<1M1 likes73 downloads9mo agoHugging Face10wannaphong /KhanomTanLLM-pretrained-dataset-thai-subset KhanomTanLLM pretrained dataset (Thai subset) This daataset collect all raw text for pretraining LLM. (Thai subset) Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0 pythainlp/thai-tnhc2-books pythainlp/thai-constitution-corpus pythainlp/thai-it-books pythainlp/prd_news_3011202 pythainlp/thailand-policy-statements pythainlp/thai-cc-license pythainlp/blognone_news pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.texttext-generation10M<n<100M0 likes68 downloads2y agoHugging Face11wangbing1416 /LonsRex-Misinfomation-SFTtexttext-generation100K<n<1M1 likes66 downloads4mo agoHugging Face12wangmingyuan /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例… See the full description on the dataset page: https://huggingface.co/datasets/wangmingyuan/HundredCV-Chat.texttext-generation10K<n<100K0 likes48 downloads26d agoHugging Face13wangbing1416 /MSMS-AceReason-20K-SFT A Multi-Source Multi-Solution Long CoT SFT Dataset from 20K AceReason Questions texttext-generation100K<n<1M1 likes38 downloads9mo agoHugging Face14sci-m-wang /Anna-CPsyCounD Dataset Card for AnnaAgent Virtual Seeker Dataset Dataset Overview Repository: AnnaAgent GitHubPurpose: Supports the AI psychotherapy agent framework AnnaAgent by providing multidimensional configuration for virtual seekersLanguage: ChineseData Source: Synthetic data constructed using GPT-4o based on CPsyCounD datasetKey Features: Contains long-term memory (historical counseling records) and short-term memory (current session state) Supports dynamic emotion evolution… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/Anna-CPsyCounD.texttext-generation1K<n<10K0 likes31 downloads6mo agoHugging Face15wannaphong /gsm8k_distilled Additional Information This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes: A mathematical problem statement A detailed step-by-step solution texttext-generation1K<n<10K0 likes29 downloads7mo agoHugging Face16arvindcr4 /tinker-rl-bench-wandb TinkerRL-Bench W&B Run Archive Full export of every Weights & Biases run under the arvindcr4-pes-university entity, covering the experiments reported in our NeurIPS submission "A Unified Benchmark for RL Post-Training of Language Models" (repo). Contents File Rows Description runs.jsonl 334 One record per run: project, run_id, run_name, state, config, summary, tags, url, runtime history.jsonl 9,255 Per-step metric history (step, reward, loss, accuracy, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/arvindcr4/tinker-rl-bench-wandb.tabulartext-generation1K<n<10K0 likes28 downloads5mo agoHugging Face17wannipaa007 /claude-opus-4.6-4.7-reasoning-8.7k Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/wannipaa007/claude-opus-4.6-4.7-reasoning-8.7k.texttext-generation10K<n<100K0 likes27 downloads4mo agoHugging Face18wangekxy /classical-chinese-punctuation Classical Chinese Punctuation Restoration · 文言文断句·标点 💰 (Commercial Dataset) This is a commercial dataset. A free 200-record sample is provided below (sample.jsonl); the full 5.3M-pair corpus is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — built from public-domain classical works. The task Restore punctuation and sentence segmentation (句读) to unpunctuated Classical Chinese — a… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-punctuation.texttext-generationn<1K0 likes25 downloads3mo agoHugging Face19wangekxy /classical-chinese-variant-collation Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset) This is a commercial dataset. A free 50-work preview sample is provided below (sample.jsonl, texts truncated); the full set with complete aligned texts is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — both members are public-domain classical works. What this is A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.tabulartext-generationn<1K0 likes20 downloads3mo agoHugging Face20dibao-research /wanli-dibao-corpus Wanli Dibao Corpus / 萬曆邸鈔校訂語料庫 Dataset Description Summary English: The Wanli Dibao Corpus is a structured, proofread digital corpus of the Wanli Dichao (萬曆邸鈔), a collection of manuscript copies of official gazettes (dibao 邸報) from the Wanli reign (1573–1620) of the Ming dynasty. The dibao system was the primary channel of official communication in imperial China, transmitting memorials, edicts, personnel appointments, and policy decisions from the capital to… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/wanli-dibao-corpus.tabulartext-generation10K<n<100K1 likes17 downloads6mo agoHugging Face21wannaphong /WanJuanSiLu-sft WanJuanSiLu sft subset This dataset is built from opendatalab/WanJuanSiLu-Multimodal-5Languages in sft subset. 180,000 SFT data Languages: Arabic, Russian, Korean, Vietnamese, and Thai 🤖Featured instructions for fine-tuning SFT data: Cultural adversarial samples: Contains culturally relevant question-answer pairs designed by local residents to detect cultural bias in models Hybrid quality inspection process: Rules + model scoring to filter translation data and… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/WanJuanSiLu-sft.texttext-generation100K<n<1M0 likes13 downloads1y agoHugging Face22guilty1987 /wangweitexttext-generationn<1K0 likes10 downloads2y agoHugging Face23subhanghuf /Fiqih_Wanita_Dataset Dataset Fiqih Wanita 10K Dataset fine-tuning chatbot ustadzah fiqih wanita berbahasa Indonesia.Berisi 10.000 pasangan tanya-jawab berdasarkan kitab Uyunul Masa-il Linnisa' madzhab Syafi'i. Topik Haidl, Nifas, Istihadloh, Darah Fasad, Thoharoh, Wudhu, Mandi Wajib, Larangan saat haidl, Qodlo sholat & puasa, Haji & Umroh, Ramadhan, Keputihan, KB, Kehamilan, Mitos, Relasi suami-istri, Dalil Al-Quran & Hadits. Format { "messages": [ {"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/subhanghuf/Fiqih_Wanita_Dataset.texttext-generation10K<n<100K0 likes7 downloads5mo agoHugging Face24wannaphong /dolma_web1_rawgated Raw data for Thai Dolma Commoncrawl Fineweb2 texttext-generation100M<n<1B0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.