CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DannHiroaki /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes12k downloads8mo agoHugging Face02liuhangbiao /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes3.8k downloads6mo agoHugging Face03LAMDA-NeSy /ChinaTravel ChinaTravel Query Dataset This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). ChinaTravel is an open-ended travel-planning benchmark with compositional constraint validation for language agents. See the paper, Hugging Face paper page, code, and bilingual sandbox database (ModelScope mirror) for the complete benchmark resources. Introduction For a given query, a language agent uses the sandbox tools to collect information and… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel.tabulartext-generation1K<n<10K14 likes1.4k downloads11d agoHugging Face04BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.4k downloads10mo agoHugging Face05Baoruixi /chimera-bench CHIMERA-Bench v1.0 A unified benchmark for epitope-specific antibody CDR sequence-structure co-design. Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop) Code: github.com/mansoorbaloch/chimera-bench Dataset Summary Property Value Complexes 2,922 PDB structures 2,721 Pre-computed features 2,941 .pt files Splits 3 (epitope-group, antigen-fold, temporal) Numbering schemes IMGT, Chothia Contact… See the full description on the dataset page: https://huggingface.co/datasets/Baoruixi/chimera-bench.tabularother1K<n<10K0 likes947 downloads2mo agoHugging Face06Liavan /Traditional-Chinese-Medicine-Multiple_choice_question Discription This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.textquestion-answering1K<n<10K4 likes816 downloads2y agoHugging Face07madao33 /new-title-chinesetext1K<n<10K21 likes750 downloads4y agoHugging Face08mansoorbaloch /chimera-bench CHIMERA-Bench v1.0 A unified benchmark for epitope-specific antibody CDR sequence-structure co-design. Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop) Code: github.com/mansoorbaloch/chimera-bench Dataset Summary Property Value Complexes 2,922 PDB structures 2,721 Pre-computed features 2,941 .pt files Splits 3 (epitope-group, antigen-fold, temporal) Numbering schemesIMGT, Chothia Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.tabularother1K<n<10K0 likes668 downloads4mo agoHugging Face09fdemelo /ipa-childes-split IPA-CHILDES split This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular, the following changes have been implemented: column processed_gloss dropped as it duplicates information of gloss up to punctuation column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+) column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.tabular10M<n<100M0 likes606 downloads1y agoHugging Face10Mxode /Chinese-Psychology-Books 免责声明与使用须知 (Disclaimer and Usage Notice) 数据集内容 本数据集包含从互联网上多个来源收集的 中文心理学电子书 的集合。 许可证 本数据集的组织结构、汇编方式以及由维护者添加的任何元数据或注释根据 知识共享署名-非商业性使用 4.0 国际许可协议 (Creative Commons Attribution-NonCommercial 4.0 International License - CC BY-NC 4.0) 提供。这意味着您可以基于非商业目的分享和修改这部分内容,但必须给出适当的署名。 请注意:此 CC BY-NC 4.0 许可证不适用于数据集中包含的原始电子书文件本身。 版权声明 数据集中包含的个别电子书文件极有可能受到版权法保护,其版权归各自的作者、出版商或其他版权所有者所有。 数据集维护者不拥有这些电子书的版权。 这些电子书的来源多样且零散,部分来源可能难以追溯。 使用限制与责任… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Psychology-Books.texttext-generationn<1K9 likes451 downloads1y agoHugging Face11kjhq /China-Stock-Symbols-and-Metadata China Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in China.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/China-Stock-Symbols-and-Metadata.text1K<n<10K0 likes444 downloads1y agoHugging Face12Johnson8187 /Chinese_Multi-Emotion_Dialogue_Dataset Chinese_Multi-Emotion_Dialogue_Dataset 📄 Description This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text. Data Sources: Daily Conversations: Captured from natural, informal human conversations. Movie Dialogues: Extracted from diverse Chinese-language movies. AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.texttext-classification1K<n<10K19 likes367 downloads12d agoHugging Face13LeoBorai /chile-seismological-records Chile's Seismological Records tabular10K<n<100K0 likes325 downloads5h agoHugging Face14shibing624 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.text100K<n<1M15 likes300 downloads2y agoHugging Face15busy-pig /ChinaPaint CCPP Evaluating and Benchmarking Classical Chinese Poetry-to-Painting for Multimodal Large Language Models CCPP: Classical Chinese Poetry-to-Painting Project A comprehensive project supporting the research on Classical Chinese Poetry-to-Painting (CCPP) generation and evaluation, including benchmark datasets, human painting references, model outputs, and auxiliary scripts. This project serves as the official code & data repository for the corresponding academic… See the full description on the dataset page: https://huggingface.co/datasets/busy-pig/ChinaPaint.image10K<n<100K1 likes294 downloads3d agoHugging Face16HHS-Official /health-conditions-among-children-under-age-18-by-s Health conditions among children under age 18, by selected characteristics: United States Description NOTE: On October 19, 2021, estimates for 2016–2018 by health insurance status were revised to correct errors. Changes are highlighted and tagged at https://www.cdc.gov/nchs/data/hus/2019/012-508.pdf Data on health conditions among children under age 18, by selected population characteristics. Please refer to the PDF or Excel version of this table in the HUS 2019 Data… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/health-conditions-among-children-under-age-18-by-s.tabular1K<n<10K0 likes291 downloads1y agoHugging Face17phonemetransformers /IPA-CHILDES IPA-CHILDES Dataset This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here. Description Key Columns The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.tabular10M<n<100M7 likes280 downloads1y agoHugging Face18inria-chile /latamqa_mcq_es-la LatamQA LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_mcq_es-la.textmultiple-choice10K<n<100K0 likes218 downloads3mo agoHugging Face19vanila434 /chinese-american-elder-fraud-qa chinese-american-elder-fraud-qa A hand-authored, trilingual (Mandarin / Cantonese / English) fraud-recognition dataset for first-generation Chinese-American elders and the adult children who help them. 235 rows authored, 207 adapted through the Adaption Labs platform with reasoning traces. Grounded in FBI, IC3, and SFPD reports on Chinese-community elder fraud. Adaption Labs Uncharted Data Challenge submission. Metric Value Rows authored 235 Rows adapted (training… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/chinese-american-elder-fraud-qa.texttext-classificationn<1K1 likes209 downloads5mo agoHugging Face20chillies /IELTS-writing-task-2-evaluationtext10K<n<100K39 likes206 downloads3y agoHugging Face21chimcis /searcless-chess-10mtext10K<n<100K1 likes204 downloads11mo agoHugging Face22SylvanL /Traditional-Chinese-Medicine-Exam Coming Soon... text1K<n<10K8 likes174 downloads2mo agoHugging Face23AnxForever /chinese-ai-detection-dataset Chinese AI Detection Dataset 中文AI文本检测数据集 数据集简介 用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。 核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。 数据统计 类型 样本数 说明 总计 66,001 训练/验证/测试集 纯人类 27,719 多领域人类文本 纯AI 27,719 多模型生成 C2 (续写) 3,781 人类开头+AI续写 C3 (改写) 3,781 AI改写人类文本 C4 (润色) 3,001 AI润色人类文本 数据格式 { "text": "文本内容(混合文本包含[SEP]标记)", "label": 0, // 0=Human, 1=AI "category": "C2", // Human/AI/C2/C3/C4 "source": "数据来源" }… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.tabular10K<n<100K1 likes170 downloads8d agoHugging Face24chitradrishti /reddew reddew Reddit Download and Datasets image1M<n<10M3 likes156 downloads2y agoHugging Face25mihai-chindris /policy-rag-corpus-metadata Policy RAG Corpus Metadata (No Raw Data) This repository is a metadata-only companion for the Policy RAG project built for the Quantic MSSE AI Engineering program. It does not include the actual PDF files. The source PDFs are hosted in the companion GitHub repository. What this repo includes metadata.csv: structured metadata for 11 policy documents (filename, title, category, page count, source type, description) Citation and provenance notes for reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/mihai-chindris/policy-rag-corpus-metadata.textquestion-answeringn<1K1 likes151 downloads3mo agoHugging Face26uralstech /AIDE-Chip-15K-gem5-Sims AIDE-Chip 15K gem5 Simulation Dataset AIDE-Chip-15K-gem5-Sims is a structured dataset of approximately 15,000 validated RISC-V gem5 simulations covering cache hierarchy design-space exploration (DSE) for single-core processors. The dataset was generated using gem5's Syscall Emulation (SE) mode and six representative workloads, spanning compute-bound, memory-bound, and irregular access patterns. Each sample maps cache configuration parameters to IPC and L2 miss rate, enabling… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/AIDE-Chip-15K-gem5-Sims.tabulartabular-regression10K<n<100K0 likes144 downloads8mo agoHugging Face27huskyhong /chinese-stock-datasettabular10M<n<100M0 likes138 downloads10mo agoHugging Face28chiapudding /kaggle-financial-sentimenttext1K<n<10K4 likes137 downloads3y agoHugging Face29FanLR /ChineseOCRBenchtext1K<n<10K0 likes137 downloads2y agoHugging Face30chimbiwide /sciqa-thinking sciqa-thinking Randomly extracted 3000 rows from sciq and prompting Qwen3-14b to generate the intermediate reasoning traces, we created this dataset. This should be used for LLM post-training, especially RL. textquestion-answering1K<n<10K0 likes124 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.