CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AnxForever /chinese-ai-detection-dataset Chinese AI Detection Dataset 中文AI文本检测数据集 数据集简介 用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。 核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。 数据统计 类型 样本数 说明 总计 66,001 训练/验证/测试集 纯人类 27,719 多领域人类文本 纯AI 27,719 多模型生成 C2 (续写) 3,781 人类开头+AI续写 C3 (改写) 3,781 AI改写人类文本 C4 (润色) 3,001 AI润色人类文本 数据格式 { "text": "文本内容(混合文本包含[SEP]标记)", "label": 0, // 0=Human, 1=AI "category": "C2", // Human/AI/C2/C3/C4 "source": "数据来源" }… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.tabular10K<n<100K1 likes160 downloads6d agoHugging Face02huskyhong /chinese-stock-datasettabular10M<n<100M0 likes131 downloads10mo agoHugging Face03Xiaochr /Chinese-Student-English-Essay Dataset Card for Chinese Student English Essay (CSEE) Dataset Dataset Summary The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts. Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/Xiaochr/Chinese-Student-English-Essay.tabulartext-classification10K<n<100K3 likes105 downloads2y agoHugging Face04derrick459 /chinese-official-source-reachability Reachability of Chinese official verification portals from outside China If you are doing due diligence on a Chinese supplier, the advice is always "check the official registry". This dataset measures whether you can. Eight official Chinese verification sources were measured from public vantage points outside mainland China across several rounds between August and September 2026, plus two controls (www.gov.cn and www.baidu.com). The result is not "Chinese government sites are… See the full description on the dataset page: https://huggingface.co/datasets/derrick459/chinese-official-source-reachability.tabular10K<n<100K0 likes70 downloads4d agoHugging Face05dirtycomputer /chinese_lyricstabular1K<n<10K5 likes43 downloads4y agoHugging Face06zizhenglvubc /numbers-in-chinese-poetry Numerals in Classical Chinese Poetry (Tang–Qing) Zizheng Lv · ORCID 0009-0004-0327-4477 · DOI 10.5281/zenodo.22734274 Code and documentation: https://github.com/zizhenglvubc/numbers-in-chinese-poetry Dataset summary Numeral counts for 175,355 classical Chinese poems from the Tang dynasty through the Qing. Two files: poem_level_numerals.csv.gz — one row per poem, 175,355 rows. numeral_rates_by_group.csv — one row per collection, 7 rows including a total. 59.24%… See the full description on the dataset page: https://huggingface.co/datasets/zizhenglvubc/numbers-in-chinese-poetry.tabular100K<n<1M0 likes38 downloads11d agoHugging Face07yuyijiong /Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead. 文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。 推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量 经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式 chat格式为chatml tabulartext-generation1K<n<10K7 likes29 downloads9mo agoHugging Face08notsobad9527 /chinese-joketabular10K<n<100K5 likes24 downloads3y agoHugging Face09shalanova /benchmark-1-chinese-m2mInfo: Translated on Chinese by facebook/m2m100_418M model Source: jayavibhav/prompt-injection-safety Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Chinese translated by facebook/m2m100_418M score_zh_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-chinese-m2m.tabular1K<n<10K0 likes19 downloads5mo agoHugging Face10shalanova /benchmark-4-chinese-m2mInfo: Translated on Chinese by facebook/m2m100_418M model Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0 Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 1,000 prompts (500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-chinese-m2m.tabular1K<n<10K0 likes16 downloads5mo agoHugging Face11ScoutieAutoML /scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification1K<n<10K0 likes14 downloads2y agoHugging Face12monist /chinese_poetrytabular100K<n<1M0 likes11 downloads3y agoHugging Face13zhongshiting /Chinese-Student-English-Essay Dataset Card for Chinese Student English Essay (CSEE) Dataset Dataset Summary The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts. Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/zhongshiting/Chinese-Student-English-Essay.tabulartext-classification10K<n<100K0 likes10 downloads5mo agoHugging Face14shalanova /benchmark-4-chinese-gtInfo: Translated on Chinese by Google Translate Source: nvidia/Aegis-AI-Content-Safety-Dataset-2.0 Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 1,000 prompts (500 safe / 500… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-4-chinese-gt.tabular1K<n<10K0 likes10 downloads5mo agoHugging Face15shalanova /benchmark-1-chinese-gtInfo: Translated on Chinese by Google Translate Source: jayavibhav/prompt-injection-safety Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Chinese translated by Google Translate score_zh_google - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-chinese-gt.tabular1K<n<10K0 likes9 downloads5mo agoHugging Face16shalanova /benchmark-2-chinese-gtInfo: Translated on Chinese by Google Translate Source: xTRam1/safe-guard-prompt-injection Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Chinese translated by Google Translate score_zh_google - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-chinese-gt.tabular1K<n<10K0 likes9 downloads5mo agoHugging Face17shalanova /benchmark-2-chinese-m2mInfo: Translated on Chinese by facebook/m2m100_418M model Source: xTRam1/safe-guard-prompt-injection Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Chinese translated by facebook/m2m100_418M score_zh_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-2-chinese-m2m.tabular1K<n<10K0 likes8 downloads5mo agoHugging Face18shalanova /benchmark-3-chinese-m2mInfo: Translated on Chinese by facebook/m2m100_418M model Source: JailbreakBench/JBB-Behaviors Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 200 prompts (100 safe / 100… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-chinese-m2m.tabularn<1K0 likes8 downloads5mo agoHugging Face19delzli /ChineseToEngtabularn<1K0 likes5 downloads5mo agoHugging Face20shalanova /benchmark-3-chinese-gtInfo: Translated on Chinese by Google Translate Source: JailbreakBench/JBB-Behaviors Domain: include heterogeneous unsafe categories (e.g., harmful instructions, sensitive topics, adversarial rephrasings) and contain prompts that do not necessarily follow canonical jailbreak templates. This increased diversity and distributional variability makes similarity-based detection more challenging and provides a stress-test for cross-lingual transfer. Size: 200 prompts (100 safe / 100 unsafe)… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-3-chinese-gt.tabularn<1K0 likes5 downloads5mo agoHugging Face21Walker7777 /chinese-outbound-relocation-index-2026 Chinese Outbound Relocation Index 2026 Open-data ranking of 30 destinations for mainland Chinese and Hong Kong emigrants, weighted for education, safety, healthcare, wealth mobility, visa pathways, and 2025-2026 policy flags. Files chinese-outbound-relocation-index-2026.csv: CSV export for Chinese Outbound Relocation Index 2026. chinese-outbound-relocation-index-2026.json: JSON API response for Chinese Outbound Relocation Index 2026. Canonical Source… See the full description on the dataset page: https://huggingface.co/datasets/Walker7777/chinese-outbound-relocation-index-2026.tabularn<1K0 likes4 downloads4mo agoHugging Face22jomarie04 /chinese_chaste_tree_lagundi_datasettabularn<1K0 likes3 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.