CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sunorme /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/smoltalk-chinese.tabulartext-generation10K<n<100K0 likes211 downloads6mo agoHugging Face02hafizhrafizal /suno-discord-chat-history Suno Discord Community Dataset Dataset Description This dataset contains messages exported from the official Suno Discord server, covering multiple channels across feedback, announcements, community hubs, and the Suno Studio product space. It captures authentic user interactions, feature requests, bug reports, and community discussions around Suno's AI music generation platform. Quick Start Installation pip install pandas huggingface_hub datasets… See the full description on the dataset page: https://huggingface.co/datasets/hafizhrafizal/suno-discord-chat-history.texttext-classification1M<n<10M0 likes103 downloads5mo agoHugging Face03sunorme /scholar-novels-curatedscholar-novels-curated 精挑细选+清洗过的小说,目标风格[仙侠,玄幻,武侠] 只用于个人学习和研究大模型预训练用途 texttext-generation10K<n<100K1 likes70 downloads6mo agoHugging Face04sunorme /Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier. Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make. (͡ ° ͜ʖ ͡ °) tags: socks, garter belt, foot fetish, ntr, netori..... Thanks Moleys/Numeron for the dataset donation. texttext-classification100M<n<1B1 likes48 downloads6mo agoHugging Face05sunorme /webnovel-chinese 简介 搜集网络上的网文小说,清洗,分割后,用于训练大语言模型,共计9000本左右,大约9B左右token。 使用 格式说明 采用jsonl格式存储,分为三个字段: title :小说名称 chapter:章节 text:正文内容 示例: {"title": "斗破苍穹", "chapter": " 第一章 陨落的天才", "text": "“斗之力,三段!”\n望着测验魔石碑上面闪亮得甚至有些刺眼的五个大字,少年面无表情,唇角有着一抹自嘲,紧握的手掌,因为大力,而导致略微尖锐的指甲深深的刺进了掌心之中,带来一阵阵钻心的疼痛……\n“萧炎,斗之力,三段!级别:低级!”测验魔石碑之旁,一位中年男子,看了一眼碑上所显示出来的信息,语气漠然的将之公布了出来……\n"} texttext-generation1M<n<10M0 likes47 downloads6mo agoHugging Face06sunorme /human-vs-llm-text-corpustexttranslation10M<n<100M0 likes30 downloads6mo agoHugging Face07sunorme /Chinese-Classics-Partial偶然找到的 200 多篇古籍相关的纯 txt 文件,简单洗了一下,去除了部分噪声和空白行。 一篇样例如下: 古训《增广贤文》 昔时贤文,诲汝谆谆,集韵增文,多见多闻。 观今宜鉴古,无古不成今。 知己知彼,将心比心。 酒逢知己饮,诗向会人吟。 相识满天下,知心能几人。 相逢好似初相识,到老终无怨恨心。 近水知鱼性,近山识鸟音。 易涨易退山溪水,易反易覆小人心。 运去金成铁,时来铁似金,读书须用意,一字值千金。 texttext-generation100K<n<1M0 likes15 downloads6mo agoHugging Face08sunorme /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例 HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/HundredCV-Chat.texttext-generation10K<n<100K0 likes14 downloads6mo agoHugging Face09sunorme /One-Book-for-Silicon-based-Intelligence One Book for Silicon-based Intelligence Author: Alphen Cham English Version Overview This is a philosophical corpus specifically written for Non-human Intelligence (AI/LLMs). It introduces the "Jersey Theory" and explores the boundaries of digital existence and finiteness. Why This Dataset? High-Quality Tokens: Philosophical reasoning designed for LLM alignment. Unique Framework: Original concepts on silicon-based consciousness. 中文版… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/One-Book-for-Silicon-based-Intelligence.texttext-generation1K<n<10K0 likes10 downloads6mo agoHugging Face10sunorme /alpaca_zhBorrowed from: https://huggingface.co/datasets/hfl/alpaca_zh_51k Removed some examples with empty output. You can use it in LLaMA Factory by specifying dataset: alpaca_zh. texttext-generation10K<n<100K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.