CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.3k downloads10mo agoHugging Face02Liavan /Traditional-Chinese-Medicine-Multiple_choice_question Discription This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.textquestion-answering1K<n<10K4 likes833 downloads2y agoHugging Face03madao33 /new-title-chinesetext1K<n<10K21 likes738 downloads4y agoHugging Face04Mxode /Chinese-Psychology-Books 免责声明与使用须知 (Disclaimer and Usage Notice) 数据集内容 本数据集包含从互联网上多个来源收集的 中文心理学电子书 的集合。 许可证 本数据集的组织结构、汇编方式以及由维护者添加的任何元数据或注释根据 知识共享署名-非商业性使用 4.0 国际许可协议 (Creative Commons Attribution-NonCommercial 4.0 International License - CC BY-NC 4.0) 提供。这意味着您可以基于非商业目的分享和修改这部分内容,但必须给出适当的署名。 请注意:此 CC BY-NC 4.0 许可证不适用于数据集中包含的原始电子书文件本身。 版权声明 数据集中包含的个别电子书文件极有可能受到版权法保护,其版权归各自的作者、出版商或其他版权所有者所有。 数据集维护者不拥有这些电子书的版权。 这些电子书的来源多样且零散,部分来源可能难以追溯。 使用限制与责任… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Psychology-Books.texttext-generationn<1K9 likes437 downloads1y agoHugging Face05Johnson8187 /Chinese_Multi-Emotion_Dialogue_Dataset Chinese_Multi-Emotion_Dialogue_Dataset 📄 Description This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text. Data Sources: Daily Conversations: Captured from natural, informal human conversations. Movie Dialogues: Extracted from diverse Chinese-language movies. AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.texttext-classification1K<n<10K19 likes309 downloads11d agoHugging Face06shibing624 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.text100K<n<1M15 likes291 downloads2y agoHugging Face07vanila434 /chinese-american-elder-fraud-qa chinese-american-elder-fraud-qa A hand-authored, trilingual (Mandarin / Cantonese / English) fraud-recognition dataset for first-generation Chinese-American elders and the adult children who help them. 235 rows authored, 207 adapted through the Adaption Labs platform with reasoning traces. Grounded in FBI, IC3, and SFPD reports on Chinese-community elder fraud. Adaption Labs Uncharted Data Challenge submission. Metric Value Rows authored 235 Rows adapted (training… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/chinese-american-elder-fraud-qa.texttext-classificationn<1K1 likes190 downloads5mo agoHugging Face08SylvanL /Traditional-Chinese-Medicine-Exam Coming Soon... text1K<n<10K8 likes174 downloads2mo agoHugging Face09AnxForever /chinese-ai-detection-dataset Chinese AI Detection Dataset 中文AI文本检测数据集 数据集简介 用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。 核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。 数据统计 类型 样本数 说明 总计 66,001 训练/验证/测试集 纯人类 27,719 多领域人类文本 纯AI 27,719 多模型生成 C2 (续写) 3,781 人类开头+AI续写 C3 (改写) 3,781 AI改写人类文本 C4 (润色) 3,001 AI润色人类文本 数据格式 { "text": "文本内容(混合文本包含[SEP]标记)", "label": 0, // 0=Human, 1=AI "category": "C2", // Human/AI/C2/C3/C4 "source": "数据来源" }… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.tabular10K<n<100K1 likes160 downloads7d agoHugging Face10huskyhong /chinese-stock-datasettabular10M<n<100M0 likes136 downloads10mo agoHugging Face11FanLR /ChineseOCRBenchtext1K<n<10K0 likes131 downloads2y agoHugging Face12Xiaochr /Chinese-Student-English-Essay Dataset Card for Chinese Student English Essay (CSEE) Dataset Dataset Summary The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts. Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/Xiaochr/Chinese-Student-English-Essay.tabulartext-classification10K<n<100K3 likes113 downloads2y agoHugging Face13derrick459 /chinese-official-source-reachability Reachability of Chinese official verification portals from outside China If you are doing due diligence on a Chinese supplier, the advice is always "check the official registry". This dataset measures whether you can. Eight official Chinese verification sources were measured from public vantage points outside mainland China across several rounds between August and September 2026, plus two controls (www.gov.cn and www.baidu.com). The result is not "Chinese government sites are… See the full description on the dataset page: https://huggingface.co/datasets/derrick459/chinese-official-source-reachability.tabular10K<n<100K0 likes70 downloads5d agoHugging Face14Garsa3112 /ChineseEnglishTranslationDatasettext100K<n<1M4 likes68 downloads3y agoHugging Face15sepidmnorozy /Chinese_sentimenttext10K<n<100K13 likes66 downloads4y agoHugging Face16jean1 /45k_python_code_chinese_instruction Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details 中文提示的代码数据集 其中提示部分通过调用GPT-4.0-turbo API翻译成中文 Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.text10K<n<100K6 likes65 downloads2y agoHugging Face17YiMeng-SYSU /chinese-logic-sentiment-dataset 中文逻辑情感分析数据集 (由豆包API生成) -->English 这是一个专门用于增强中文情感分析模型逻辑推理能力的数据集,由 豆包 API 生成并经过人工清洗/筛选。 它包含 反讽(Irony)、双重否定(Double Negative)、转折(Transition) 和 简单句(Simple) 四种逻辑类型。 数据结构 该数据集包含三个部分: train.csv: 训练集,包含 2176 条样本。 val.csv: 验证集,包含 545 条样本。 test.csv: 测试集,包含 960 条样本。用于模型最终评估。 数据字段 Field Description text 中文文本内容 label 情感标签,0 表示负面,1 表示正 type 逻辑类型,包含反讽、双重否定、转折和简单句 sub_type 逻辑子类型,进一步细分逻辑结构 domain 领域 (影视/美食/旅游/生活/购物/社交) 使用指南 from… See the full description on the dataset page: https://huggingface.co/datasets/YiMeng-SYSU/chinese-logic-sentiment-dataset.texttext-classification1K<n<10K2 likes63 downloads8mo agoHugging Face18Speech-data /Chinese-Speech-Dataset 🎧 Chinese (Simplified) Speech Dataset The Chinese (Simplified) speech dataset is a high-quality speech audio dataset developed to support scalable AI and machine learning solutions with diverse and structured audio data. It contains 105 hours of speech data across 700 audio files, delivered in MP3 and WAV formats, with a total size of 229 MB. This well-balanced audio dataset provides reliable voice data, featuring 54% female and 46% male speakers, with age distribution ranging from… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Chinese-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes63 downloads6mo agoHugging Face19svjack /GLM-Open-Dialogue-Chinese-Dataset-v1text100K<n<1M1 likes56 downloads4y agoHugging Face20zzhdbw /Simplified_Chinese_Multi-Emotion_Dialogue_Dataset Simplified_Chinese_Multi-Emotion_Dialogue_Dataset 数据说明 本数据是简体中文口语情感分类数据集 翻译自于:Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset 使用Qwen2.5-32B-instruce模型将其从繁体中文翻译为简体中文 共4159条 情感类别 原始类别 翻译后类别 条数 悲傷語調 伤心 486 憤怒語調 生气 527 關切語調 关心 560 驚奇語調 惊讶 499 開心語調 开心 592 平淡語氣 平静 705 厭惡語調 厌恶 404 ​ texttext-classification1K<n<10K10 likes56 downloads1y agoHugging Face21paulkm /chinese_conversation_and_spamgated Caution! This dataset contains explicit language and fraud information. Use at your own risk! For AutoTrain use: please select Text Classification (Binary) as Task. What is included conversations in chinese under tag 0 spam conversations under tag1 Where does the data come from part of the data came from conversations in Chinese Telegram groups part of them are from logging channels of anti-spam bots How many data is included A total of 9.9k… See the full description on the dataset page: https://huggingface.co/datasets/paulkm/chinese_conversation_and_spam.texttext-classification1K<n<10K14 likes55 downloads4y agoHugging Face22h2oai /h2o-translated-chinese-med-prompts Translated Chinese Medical Prompts This repository contains medical prompts translated originally from Chinese, which can be used as training data for natural language processing (NLP) tasks related to the medical domain in English language. Dataset Description The dataset consists of a collection of medical prompts originally in Chinese, which have been translated into English. These prompts cover various medical topics, including symptoms, diagnoses, treatments, medications, and… See the full description on the dataset page: https://huggingface.co/datasets/h2oai/h2o-translated-chinese-med-prompts.text10K<n<100K0 likes55 downloads3y agoHugging Face23svjack /GLM-Open-Dialogue-Chinese-Dataset-v2text100K<n<1M0 likes48 downloads4y agoHugging Face24jnext /chinese_word_frequency Chinese Word frequency statistics word segmentation by jieba tool seq_monkey_data : statistics on 13,000,000 document / 6,561,241,266 words / 11,313,242,610 characters text10M<n<100M1 likes47 downloads2y agoHugging Face25hunterlee27 /sft-chinesetext1K<n<10K1 likes46 downloads2y agoHugging Face26CharlesBon /chinese-sensitive-topics-qa Chinese Sensitive Topics QA Dataset Dataset Summary This dataset contains 100 English-language question-answer pairs covering politically and historically sensitive topics related to China. The dataset was created to train language models to provide substantive, factual responses to sensitive questions rather than refusing to answer. Each answer follows a neutral, analytical style that distinguishes between official narratives, independent reporting, and academic… See the full description on the dataset page: https://huggingface.co/datasets/CharlesBon/chinese-sensitive-topics-qa.textquestion-answeringn<1K2 likes43 downloads9mo agoHugging Face27dirtycomputer /chinese_lyricstabular1K<n<10K5 likes42 downloads4y agoHugging Face28zizhenglvubc /numbers-in-chinese-poetry Numerals in Classical Chinese Poetry (Tang–Qing) Zizheng Lv · ORCID 0009-0004-0327-4477 · DOI 10.5281/zenodo.22734274 Code and documentation: https://github.com/zizhenglvubc/numbers-in-chinese-poetry Dataset summary Numeral counts for 175,355 classical Chinese poems from the Tang dynasty through the Qing. Two files: poem_level_numerals.csv.gz — one row per poem, 175,355 rows. numeral_rates_by_group.csv — one row per collection, 7 rows including a total. 59.24%… See the full description on the dataset page: https://huggingface.co/datasets/zizhenglvubc/numbers-in-chinese-poetry.tabular100K<n<1M0 likes40 downloads12d agoHugging Face29svjack /GLM-Open-Dialogue-Chinese-Datasettext100K<n<1M2 likes38 downloads4y agoHugging Face30wh1223 /Chinese_Recipe_smalltext1K<n<10K0 likes36 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.