CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ChilleD /SVAMPtexttext-generation1K<n<10K23 likes86k downloads2y agoHugging Face02opencsg /chinese-fineweb-edu This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.texttext-generation10M<n<100M117 likes18k downloads10mo agoHugging Face03chilomax /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/fineweb-edu.tabulartext-generation1B<n<10B0 likes4.4k downloads3mo agoHugging Face04Morton-Li /ChineseWebText2.0-HighQuality 📘 ChineseWebText2.0-HighQuality Overview ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License). This subset retains only samples with: quality_score ≥ 0.9 toxicity.score ≤ 0.01 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, and quality-sensitive downstream tasks. This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.texttext-generation100M<n<1B4 likes3.4k downloads7mo agoHugging Face05opencsg /chinese-fineweb-edu-v2 This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset V2 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu Dataset V2 is a comprehensive upgrade of the original Chinese Fineweb Edu, designed and optimized for natural language processing (NLP) tasks in the education sector. This high-quality Chinese pretraining dataset has… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu-v2.tabulartext-generation100M<n<1B75 likes2.4k downloads10mo agoHugging Face06opencsg /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/smoltalk-chinese.tabulartext-generation10K<n<100K52 likes1.7k downloads10mo agoHugging Face07b3x0m /Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier. Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make. (͡ ° ͜ʖ ͡ °) tags: socks, garter belt, foot fetish, ntr, netori..... Thanks Moleys/Numeron for the dataset donation. texttext-classification100M<n<1B247 likes1.5k downloads2y agoHugging Face08jorgeortizfuentes /universal_spanish_chilean_corpus Universal Chilean Spanish Corpus Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal. Los textos en español multidialectal provienen del spanish books. Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus Name Count Source books 87967 spanish books mc4 8706681 from mc4 (.cl domains) in chilean-spanish-corpus twitter 27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.texttext-generation10M<n<100M8 likes1.5k downloads3y agoHugging Face09Mxode /Chinese-Instruct-Lite 中文指令微调数据集 - Lite 版本 💻 Github Repo [!TIP] 这不是 Chinese-Instruct 的子集,而是一个全新的简化数据集。 如果您想要一个可以真实使用、而不仅仅适用于学习的数据集,欢迎访问:Mxode/Chinese-Instruct 如果您想要一个更加简单易收敛、主题集中的数据集,可以访问:Mxode/I_Wonder_Why-Chinese 具体构成 本数据集包含如下 5 个子集,总数据量 10M+。 code:代码主题的指令数据集,数据量 1.2M+。 math:数学主题的指令数据集,数据量 1.7M+。 general:通用指令数据集,主题广泛,与 code 和 math 指令不重复,数据量 5.1M+。 math(reasoning):数学推理数据集,指令采样自 math 子集,可通过 id 关联,数据量 1.2M+。code(reasoning):代码推理数据集,指令采样自 code 子集,可通过 id 关联,数据量 700K+。 如何使用… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Instruct-Lite.textquestion-answering10M<n<100M13 likes895 downloads1y agoHugging Face10erhwenkuo /c4-chinese-zhtw Dataset Card for "c4-chinese-zhtw" 內容 Common Crawl 是一個非營利組織,負責抓取網路並向公眾免費提供其檔案和資料集。Common Crawl 的網路檔案包含自 2008 年以來收集的 PB 級資料。它一般每月完成一次抓取。 Common Crawl 的爬蟲程式遵守 nofollow 和 robots.txt 政策。用於處理 Common Crawl 資料集的開源程式碼是公開可用的。 這個繁中的數據來是來自 Common Crawl 2023-14 的 data archive 下載并進行清理 。 這是 jed351 準備的版本,託管在這個位址: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered 支援的任務 C4主要用於預訓練語言模型(pretrain language model)。 範例 一個樣本的範例: {… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/c4-chinese-zhtw.texttext-generation1M<n<10M12 likes714 downloads3y agoHugging Face11lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes604 downloads7mo agoHugging Face12lchakkei /OpenOrca-Traditional-Chinese🐋 OpenOrca-Chinese 数据集!🐋 感謝 Open-Orca/OpenOrca 資料集的發布,為廣大NLP研究人員和開發者帶來了寶貴的資源! 這是一個對 Open-Orca/OpenOrca 資料集中文翻譯的版本,翻譯引擎為 Google 翻譯,希望能為中文 LLM 研究做出一點點貢獻。 Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/lchakkei/OpenOrca-Traditional-Chinese.texttext-classification1M<n<10M11 likes560 downloads3y agoHugging Face13Mxode /I_Wonder_Why-Chinese 🧐 十万个为什么 - 中文百科开放问答数据集 💻 Github Repo 这是一个中文百科开放问答数据集,共分为 3 个子集:general、preference、reasoning。这个数据集可适用于 SFT 指令微调、DPO 类强化学习、R1 类推理蒸馏任务。 [!tip] [2025/05/09] 发布了一个新的中文指令数据集 Chinese-Instruct-Lite,包含代码、数学、通用多场景,同样包含一般指令微调数据与推理数据,数据总量 10M+ [2025/05/05] 更新:数据集扩增,现在指令由 600K+ 增加到 1.2M+ 了! 数据集详情 所有的子集共享相同的指令(prompt),共计 1.2M+,每一条指令都有自己独有的 12 位 id。这意味着你可以根据 id 交叉混合使用不同的子集。 由于指令相同,因此所有子集的数据量都是一致的,均为 1.2M+。 general:这个子集适用于 SFT 指令微调,形式是最简单的 prompt-response 格式。… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/I_Wonder_Why-Chinese.texttext-generation1M<n<10M24 likes552 downloads1y agoHugging Face14DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes538 downloads4mo agoHugging Face15senry5433 /china-effective-laws-regulations 全国现行法律法规合集 现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。 数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。 这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。 快照日期:2026-08-26 效力说明 本数据集 以现行有效法律法规为主体: 效力 status 法规份数 说明 有效 17,649 现行有效,默认应使用这一部分 尚未生效 7 已公布、施行日晚于快照日 失效 45 文件名含「失效」,多为已到期的全国人大常委会试点授权决定 使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.tabularquestion-answering100K<n<1M0 likes457 downloads1mo agoHugging Face16erhwenkuo /clean_passages_80m-chinese-zhtw Dataset Card for "clean_passages_80m-chinese-zhtw" 包含8千萬餘萬(88328203)個中文段落,不包含任何字母、數字。文字長度大部分介於 50~200 個字。 原始資料集是用於訓練GENIUS模型中文版。論文參考引用: @article{guo2022genius, title={GENIUS: Sketch-based Language Model Pre-training via Extreme and Selective Masking for Text Generation and Augmentation}, author={Guo, Biyang and Gong, Yeyun and Shen, Yelong and Han, Songqiao and Huang, Hailiang and Duan, Nan and Chen, Weizhu}, journal={arXiv preprint arXiv:2211.10330}, year={2022} }… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/clean_passages_80m-chinese-zhtw.texttext-generation10M<n<100M2 likes454 downloads3y agoHugging Face17yys /OpenOrca-Chinese🐋 OpenOrca-Chinese 数据集!🐋 感谢 Open-Orca/OpenOrca 数据集的发布,给广大NLP研究人员和开发者带来了宝贵的资源! 这是一个对 Open-Orca/OpenOrca 数据集中文翻译的版本,翻译引擎为 Google 翻译,希望能给中文 LLM 研究做出一点点贡献。 Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing… See the full description on the dataset page: https://huggingface.co/datasets/yys/OpenOrca-Chinese.texttext-classification1M<n<10M104 likes445 downloads3y agoHugging Face18BBBBBBBBBBBQ /TC260-Chinese-Safety-Prompts TC260 Chinese Safety Prompts V1 Public research dataset containing synthetic Chinese safety-testing prompts. Records have different quality tiers; the full dataset must not be described as human-verified or Gold data. 这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据 由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径 扫描、精确去重和四字shingle近似去重。 本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。 类别名称和映射用于研究性实现,不构成法律、监管或合规结论。 数据规模 原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。 A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.tabulartext-generation1K<n<10K1 likes357 downloads2mo agoHugging Face19ChipHolmes /securecode-web-archive SecureCode Web: Traditional Web & Application Security Dataset Production-grade web security vulnerability dataset with complete incident grounding, 4-turn conversational structure, and comprehensive operational guidance Paper | GitHub | Dataset | Model Collection | Blog Post What's new in v2.6 v2.6 restores proper Express.js coverage for the topics whose examples were removed in v2.5.1 (they had shared one reused answer). 29 new, genuinely distinct Express.js… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/securecode-web-archive.texttext-generation1K<n<10K0 likes307 downloads2mo agoHugging Face20chiffonng /en-vocab-mnemonicsThis dataset is produced as a part of my Capstone (see Github). The dataset uses roughly 500 examples from Balepur et. al, 2024. Shallow-Encoding Mnemonics Deep-Encoding Mnemonics Homophonic: olfactory sounds like "old factory." Etymology: preposterous - pre (before) + post (after) + erous, which implies absurd. Chunking: obsequious sounds like "ob-se-ki-ass. Obedient servants kiss your ass Morphology: Suffixes "ate" are usually verbs. Prefix "ab" means from, away. Keyword:… See the full description on the dataset page: https://huggingface.co/datasets/chiffonng/en-vocab-mnemonics.texttext-classification1K<n<10K0 likes306 downloads2y agoHugging Face21TianHongZXY /CHIMERA CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning CHIMERA is a compact but high-difficulty synthetic reasoning datasetwith long Chain-of-Thought (CoT) trajectories and broad STEM coverage, designed for reasoning post-training. All examples are fully LLM-generated and automatically verified without human annotation. Total: 9,225 problems Subjects: 8 Topics: 1,179 🔥 Why CHIMERA? Recent reasoning advances rely heavily on high-quality… See the full description on the dataset page: https://huggingface.co/datasets/TianHongZXY/CHIMERA.texttext-generation1K<n<10K24 likes285 downloads5mo agoHugging Face22chiffonng /en-vocab-en-mnemonics-cottexttext-generation1K<n<10K1 likes281 downloads1y agoHugging Face23zake7749 /chinese-writing-benchmark Zhiyin: Exploring the Frontier of Chinese LLM Writing Website • GitHub • Hugging Face Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks. Benchmark Overview Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5. Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-benchmark.texttext-generation1K<n<10K1 likes260 downloads7mo agoHugging Face24alphagocc /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different folders… See the full description on the dataset page: https://huggingface.co/datasets/alphagocc/Fineweb-Edu-Chinese-V2.1.texttext-generation10M<n<100M0 likes254 downloads9mo agoHugging Face25chilomax /SWE-smith-trajectories SWE-smith Trajectories Code • Paper • Site This dataset contains the 5017 trajectories we fine-tuned Qwen 2.5 Coder Instruct on, leading to SWE-agent-LM-32B, a coding LM agent that achieve 40.2% on SWE-bench Verified (no verifiers or multiple rollouts, just 1 attempt per instance). Trajectories were generated by running SWE-agent + Claude 3.7 Sonnet on task instances from the SWE-smith dataset. texttext-generation10K<n<100K0 likes229 downloads3mo agoHugging Face26erhwenkuo /openorca-chinese-zhtw Dataset Card for "openorca-chinese-zhtw" Dataset Summary The OpenOrca dataset is a collection of augmented FLAN Collection data. Currently ~1M GPT-4 completions, and ~3.2M GPT-3.5 completions. It is tabularized in alignment with the distributions presented in the ORCA paper and currently represents a partial completion of the full intended dataset, with ongoing generation to expand its scope. The data is primarily used for training and evaluation in the field of natural… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/openorca-chinese-zhtw.texttext-classification1M<n<10M3 likes228 downloads3y agoHugging Face27chillies /ielts-writing-task2-essays 📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9) Dataset Summary This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.). Each record includes: essay_id: Unique identifier on Writing9 topic: Topic category (e.g. Art, Business and Companies, Cities) question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.tabulartext-classification1K<n<10K3 likes223 downloads2mo agoHugging Face28sunorme /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/smoltalk-chinese.tabulartext-generation10K<n<100K0 likes213 downloads6mo agoHugging Face29AmanPriyanshu /reasoning-sft-CHIMERA reasoning-sft-CHIMERA Converted version of TianHongZXY/CHIMERA, filtered and reformatted for SFT/reasoning training. Both subsets (Qwen3-235B-2507 and Qwen3.5-397B) are included. No content was modified or regenerated, just reformatted the columns into a standard messages format. Filtering Kept only rows with correctness == True Randomly dropped 50% of Mathematics rows to reduce math dominance Both subsets combined into a single file Format Each row has three… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-CHIMERA.texttext-generation10K<n<100K0 likes204 downloads7mo agoHugging Face30zake7749 /chinese-writing-bench-judgements-gpt-5.4 Zhiyin: Exploring the Frontier of Chinese LLM Writing Website • GitHub • Hugging Face Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks. Benchmark Overview Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5. Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements-gpt-5.4.tabulartext-generation1K<n<10K0 likes176 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.