CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01neigezhu /china-a-share-1min-ohlcv China A-Share Equities 1-Minute OHLCV Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports. Dataset summary This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.tabulartime-series-forecasting100K<n<1M12 likes64k downloads1mo agoHugging Face02actava /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.documenttext-generationn<1K61 likes6.3k downloads4mo agoHugging Face03Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.9k downloads2mo agoHugging Face04chilomax /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/fineweb-edu.tabulartext-generation1B<n<10B0 likes4.4k downloads3mo agoHugging Face05opencsg /chinese-fineweb-edu-v2 This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset V2 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu Dataset V2 is a comprehensive upgrade of the original Chinese Fineweb Edu, designed and optimized for natural language processing (NLP) tasks in the education sector. This high-quality Chinese pretraining dataset has… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu-v2.tabulartext-generation100M<n<1B75 likes2.4k downloads10mo agoHugging Face06vessel888 /china-a-share-1min-ohlcv China A-Share Equities 1-Minute OHLCV Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports. Dataset summary This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/china-a-share-1min-ohlcv.tabulartime-series-forecasting100K<n<1M1 likes2.4k downloads16d agoHugging Face07BAAI /IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine IndustryCorpus2: Health & Medicine This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.tabular10M<n<100M11 likes2.3k downloads1mo agoHugging Face08chibifire /kaggle-womens-ecom-clothing-reviews Women's Clothing E-Commerce Reviews (wide-row repack) Repack of Kaggle dataset nicapotato/womens-ecommerce-clothing-reviews (CC0-1.0) into a single wide parquet with ZStandard level 9 compression. Schema: one row per review — row_id (int64), clothing_id (int32), age (int32), rating (int8), recommended_ind (int8), positive_feedback_count (int32), title (string, nullable), review_text (string, nullable), division_name (string, nullable), department_name (string, nullable)… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/kaggle-womens-ecom-clothing-reviews.tabulartext-classification10K<n<100K0 likes1.7k downloads23d agoHugging Face09opencsg /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/smoltalk-chinese.tabulartext-generation10K<n<100K52 likes1.7k downloads10mo agoHugging Face10Magpie-Align /Magpie-Qwen2-Pro-200K-Chinese Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2-Pro-200K-Chinese.tabularquestion-answering100K<n<1M84 likes1.5k downloads2y agoHugging Face11LAMDA-NeSy /ChinaTravel ChinaTravel Query Dataset This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). ChinaTravel is an open-ended travel-planning benchmark with compositional constraint validation for language agents. See the paper, Hugging Face paper page, code, and bilingual sandbox database (ModelScope mirror) for the complete benchmark resources. Introduction For a given query, a language agent uses the sandbox tools to collect information and… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel.tabulartext-generation1K<n<10K14 likes1.4k downloads11d agoHugging Face12china-ai-law-challenge /cail2018 Dataset Card for CAIL 2018 Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/china-ai-law-challenge/cail2018.tabularother1M<n<10M31 likes1.2k downloads3y agoHugging Face13Baoruixi /chimera-bench CHIMERA-Bench v1.0 A unified benchmark for epitope-specific antibody CDR sequence-structure co-design. Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop) Code: github.com/mansoorbaloch/chimera-bench Dataset Summary Property Value Complexes 2,922 PDB structures 2,721 Pre-computed features 2,941 .pt files Splits 3 (epitope-group, antigen-fold, temporal) Numbering schemes IMGT, Chothia Contact… See the full description on the dataset page: https://huggingface.co/datasets/Baoruixi/chimera-bench.tabularother1K<n<10K0 likes947 downloads2mo agoHugging Face14Congliu /Chinese-DeepSeek-R1-Distill-data-110k 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face&nbsp;&nbsp; | &nbsp;&nbsp;🤖 ModelScope &nbsp;&nbsp; | &nbsp;&nbsp;🚀 Github &nbsp;&nbsp; | &nbsp;&nbsp;📑 Blog 注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。 该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.tabulartext-generation100K<n<1M789 likes835 downloads2y agoHugging Face15ryan-koch /US-Child-Care-ProvidersPer-state child care provider licensing data. Each U.S. state is exposed as a separate dataset configuration: pick a state in the dataset viewer, or pass its name as the config argument, e.g. load_dataset("<repo>", "alabama"). US Child Care Providers The goal of this dataset is to collect data on child care providers in each of the states within the United States. The data collection sources are public themselves with this effort representing bundling them together to make the… See the full description on the dataset page: https://huggingface.co/datasets/ryan-koch/US-Child-Care-Providers.tabular100K<n<1M1 likes750 downloads4d agoHugging Face16mansoorbaloch /chimera-bench CHIMERA-Bench v1.0 A unified benchmark for epitope-specific antibody CDR sequence-structure co-design. Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop) Code: github.com/mansoorbaloch/chimera-bench Dataset Summary Property Value Complexes 2,922 PDB structures 2,721 Pre-computed features 2,941 .pt files Splits 3 (epitope-group, antigen-fold, temporal) Numbering schemesIMGT, Chothia Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.tabularother1K<n<10K0 likes668 downloads4mo agoHugging Face17neigezhu /china-etf-1min-ohlcv China Exchange-Traded Funds 1-Minute OHLCV Minute-level OHLCV bars for selected exchange-listed Chinese ETFs. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports. Dataset summary This snapshot contains 4,058,681 rows for 7 instruments across selected exchange-listed Chinese ETFs. It covers 2015-01-05 09:30:00 through 2026-08-20 15:00:00. Prices are unadjusted. Volume is stored in shares, turnover in… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-etf-1min-ohlcv.tabulartime-series-forecasting1K<n<10K2 likes623 downloads1mo agoHugging Face18thomaslee1818 /llava-onevision-qwen35-chinese-statstabular1M<n<10M0 likes609 downloads2mo agoHugging Face19fdemelo /ipa-childes-split IPA-CHILDES split This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular, the following changes have been implemented: column processed_gloss dropped as it duplicates information of gloss up to punctuation column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+) column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.tabular10M<n<100M0 likes606 downloads1y agoHugging Face20lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes604 downloads7mo agoHugging Face21BitRobot /G1_WBT_Dex1_Building-Children-Table Data Structure Observations observation.state.ee_state (12) End-effector states of the robot. Computed via forward kinematics (FK) from the root link to the left and right end-effectors. Includes the contribution of the waist. Represented as concatenated poses of both end-effectors. observation.state.hand_state (2) Finger states for both hands. Dex1 Hand (range: 5.5 – 0.0, open → close) Per hand: Open/close… See the full description on the dataset page: https://huggingface.co/datasets/BitRobot/G1_WBT_Dex1_Building-Children-Table.tabular1M<n<10M3 likes572 downloads3mo agoHugging Face22Scicom-intl /Malaysian-Chinese-Emilia Malaysian-Chinese-Emilia Use https://github.com/mesolitica/Emilia to pseudo-label Malaysian Chinese audio. Total rows: 605169 Total hours: 1857.611445057867 hours Permutation for Voice Conversion Also we already calculated speaker permutation to prepare for voice conversion. tabular10M<n<100M1 likes553 downloads8mo agoHugging Face23DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes538 downloads4mo agoHugging Face24bdager /CHIRLAgated Dataset Card for CHIRLA CHIRLA (Comprehensive High-resolution Identification and Re-identification for Large-scale Analysis) is a long-term, multi-camera person Re-Identification (Re-ID) and tracking dataset. It spans 7 months, 7 cameras, 22 identities, and ~1M identity-annotated bounding boxes across ~596k frames, captured in connected indoor environments. Dataset Details Dataset Description CHIRLA targets long-term appearance change (e.g., clothing changes… See the full description on the dataset page: https://huggingface.co/datasets/bdager/CHIRLA.image10K<n<100K8 likes465 downloads1y agoHugging Face25senry5433 /china-effective-laws-regulations 全国现行法律法规合集 现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。 数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。 这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。 快照日期:2026-08-26 效力说明 本数据集 以现行有效法律法规为主体: 效力 status 法规份数 说明 有效 17,649 现行有效,默认应使用这一部分 尚未生效 7 已公布、施行日晚于快照日 失效 45 文件名含「失效」,多为已到期的全国人大常委会试点授权决定 使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.tabularquestion-answering100K<n<1M0 likes457 downloads1mo agoHugging Face26MonumentalSystems /chinchilla-master-corpus-v1 Chinchilla Master Corpus v1 This dataset contains parquet shards for a curated text corpus with train and validation splits. Files data/train-*.parquet data/validation-*.parquet _dedup.sqlite Notable columns text, domain, doremi_domain, doremi_weight, source_id, source_ref, source_type, split_source, content_type, subject, reading_level, complexity, curriculum_stage, difficulty, flesch, mtld, quality_pass, quality_reason, word_count, char_count… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/chinchilla-master-corpus-v1.tabular1M<n<10M0 likes429 downloads5mo agoHugging Face27simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes429 downloads53m agoHugging Face28LAMDA-NeSy /ChinaTravel-Sandbox ChinaTravel Sandbox Environment Database This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). English | 简体中文 Release version: 2026.08.2 English This dataset contains the bilingual static sandbox used by ChinaTravel. It is a companion to the ChinaTravel query dataset and an artifact of the ChinaTravel paper. The raw ZIP snapshots preserve the exact directory layout expected by the ChinaTravel evaluator. Viewer-friendly Parquet… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel-Sandbox.tabular10K<n<100K0 likes421 downloads11d agoHugging Face29yufan /SFT_Chinese_Generaltabular1M<n<10M7 likes418 downloads2y agoHugging Face30Congliu /Chinese-DeepSeek-R1-Distill-data-110k-SFT 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face   |   🤖 ModelScope    |   🚀 Github    |   📑 Blog 注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下: Math:共计36568个样本, Exam:共计2432个样本, STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.tabulartext-generation100K<n<1M225 likes409 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.