CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ChilleD /MultiArithtextn<1K17 likes103k downloads3y agoHugging Face02gksriharsha /chitralekha Chitralekha Dataset Details Dataset Version Some of the fonts do not have proper letters/rendering of different telugu letter combinations. Those have been removed as much as I can find them. If there are any other mistakes that you notice, please raise an issue and I will try my best to look into it Dataset Description This extensive dataset, hosted on Huggingface, is a comprehensive resource for Optical Character Recognition (OCR) in the Telugu… See the full description on the dataset page: https://huggingface.co/datasets/gksriharsha/chitralekha.imageimage-to-text10M<n<100M5 likes87k downloads2y agoHugging Face03ChilleD /SVAMPtexttext-generation1K<n<10K23 likes86k downloads2y agoHugging Face04neigezhu /china-a-share-1min-ohlcv China A-Share Equities 1-Minute OHLCV Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports. Dataset summary This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.tabulartime-series-forecasting100K<n<1M12 likes64k downloads1mo agoHugging Face05opencsg /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.1.text-generation10B<n<100B80 likes57k downloads8mo agoHugging Face06venvoo /china-a-share-l2-level2-limit-order-book-tick-datagated China A-Share Level-2 Archive 2017–2026 · Quotes, orders and trades · Parquet A historical archive of Chinese exchange Level-2 data, supplied through a vendor export. It includes ten-level quote snapshots, individual order messages and trade-stream records. The files cover A-share stocks and non-stock instruments such as ETFs and bonds. “Full-market” describes the export's scope, not a guarantee that every instrument or message is present. 中国证券市场 Level-2… See the full description on the dataset page: https://huggingface.co/datasets/venvoo/china-a-share-l2-level2-limit-order-book-tick-data.n>1T115 likes30k downloads5d agoHugging Face07opencsg /chinese-fineweb-edu This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.texttext-generation10M<n<100M117 likes18k downloads10mo agoHugging Face08opencsg /Fineweb-Edu-Chinese-V2.2 Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train) [[中文]] | [[English]] OpenCSG Community | 👾 GitHub | 📖 Technical Report Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain. This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.2.text-generation10B<n<100B83 likes17k downloads8mo agoHugging Face09ChilleD /StrategyQAtext1K<n<10K6 likes16k downloads3y agoHugging Face10FreedomIntelligence /MMLU_ChineseChinese version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT. 2 likes13k downloads3y agoHugging Face11DannHiroaki /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes12k downloads8mo agoHugging Face12PediaMedAI /ChildGaitDecoding Children's Gait Behavior ECCV 2026 Yifan Shen1,2,*, Boyi Li1,*, Meihuan Huang2,3,4,*, Yuanzhe Liu1,*, Xu Cao1,2,*,§, Jinyang Jin1, Zhengyuan Li1, Anglin Liu5, Junho Kim1, Jingyuan Zhu2, Fangzhou Lan2, Jianguo Cao2,3, Jintai Chen5, Ismini Lourentzou1, James M. Rehg1,† 1 University of Illinois Urbana-Champaign &nbsp; 2 PediaMed AI &nbsp; 3 Shenzhen Children's Hospital 4 Hong Kong Polytechnic University… See the full description on the dataset page: https://huggingface.co/datasets/PediaMedAI/ChildGait.video-classification1K<n<10K4 likes11k downloads1mo agoHugging Face13Reza2kn /telegram-audiobook-chizzled Telegram Persian Audiobook Chizzled 1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.audion<1K4 likes6.7k downloads14d agoHugging Face14actava /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.documenttext-generationn<1K61 likes6.3k downloads4mo agoHugging Face15M3LEO /china Decompression If you need to decompress the files, please see the main README at the github repo. If you want to use them directly from the parquet files, the original .tif/.nc files were read into the rows as binary file data sources China AOI We provide '.csv' files with predefined train (blue), test (orange) and validation (green) splits that can be used for repeatability and comparability of experiments. 60% of tiles are allocated for training, 20% for validation… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO/china.text1M<n<10M1 likes6.2k downloads2y agoHugging Face16silk-road /alpaca-data-gpt4-chinesetexttext-generation10K<n<100K104 likes5.5k downloads3y agoHugging Face17CASIA-LM /ChineseWebText ChineseWebText: Large-Scale High-quality Chinese Web Text Extracted with Effective Evaluation Model This directory contains the ChineseWebText dataset, and the EvalWeb tool-chain to process CommonCrawl Data. Our EvalWeb tool is publicly available on github https://github.com/CASIA-LM/ChineseWebText. ChineseWebText Dataset Overview We release the latest and largest Chinese dataset ChineseWebText, which consists of 1.42 TB data and each text is assigned a… See the full description on the dataset page: https://huggingface.co/datasets/CASIA-LM/ChineseWebText.text1K<n<10K45 likes5.5k downloads3y agoHugging Face18ChiefJang /visual_robust_robocasa_x1 likes5.2k downloads24d agoHugging Face19chinh02 /UIBenchKit2 likes5.1k downloads6mo agoHugging Face20Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.9k downloads2mo agoHugging Face21chilomax /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/fineweb-edu.tabulartext-generation1B<n<10B0 likes4.4k downloads3mo agoHugging Face22beyond /chinese_clean_passages_80m chinese_clean_passages_80m 包含8千余万(88328203)个纯净中文段落,不包含任何字母、数字。Containing more than 80 million pure & clean Chinese passages, without any letters/digits/special tokens. 文本长度大部分介于50~200个汉字之间。The passage length is approximately 50~200 Chinese characters. 通过datasets.load_dataset()下载数据,会产生38个大小约340M的数据包,共约12GB,所以请确保有足够空间。Downloading the dataset will result in 38 data shards each of which is about 340M and 12GB in total. Make sure there's enough space in your device:) >>> passage_dataset… See the full description on the dataset page: https://huggingface.co/datasets/beyond/chinese_clean_passages_80m.text10M<n<100M30 likes4.3k downloads4y agoHugging Face23liuhangbiao /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes3.8k downloads6mo agoHugging Face24lainka0o0 /chinese-novel-nonH-collect Dataset Card for Dataset Name license: cc0-1.0 task_categories: - text-classification - summarization language: - zh tags: - art size_categories: - 100M<n<1B texttext-classification100M<n<1B6 likes3.7k downloads1y agoHugging Face25skkwowee /chimera-cs2 Chimera CS2 Dataset Labeled Counter-Strike 2 screenshots for vision-language model training. Each sample has: A CS2 gameplay screenshot Ground truth JSON with game_state, analysis, and advice Structure screenshots/ # PNG/JPG images labels/ # Matching JSON files (same stem name) manifest.jsonl # Data provenance tracking Usage from datasets import load_dataset ds = load_dataset("skkwowee/chimera-cs2") Stats Labels: 5309… See the full description on the dataset page: https://huggingface.co/datasets/skkwowee/chimera-cs2.image-to-text1K<n<10K0 likes3.6k downloads2mo agoHugging Face26RainPPR /china-textbook-2021-hf China Textbook 2021 中国教育部审定的中小学教科书 PDF 合集。 简介 本数据集包含从小学到高中各年级的教育部审定教科书 PDF。原始数据来自 TapXWorld/ChinaTextbook,由本项目自动合并碎片化文件后托管在 Hugging Face Datasets 上。 处理代码:https://github.com/rainewhk/china-textbook-2021-hf 数据结构 教材按学段、学科、版本、年级分层存放: 小学/ ├── ... 初中/ ├── ... 初中(五•四学制)/ ├── ... 小学(五•四学制)/ ├── ... 高中/ └── ... 文件名为 义务教育教科书·<学科> <年级> <册>册.pdf 或类似格式。 处理说明 上游 PDF 存在碎片化存储(.1, .2 等分片),本项目通过 Rust 合并器 自动合并为完整文件。 部分教材同时存在完整 PDF 和 _merge_folder… See the full description on the dataset page: https://huggingface.co/datasets/RainPPR/china-textbook-2021-hf.document1K<n<10K3 likes3.5k downloads3mo agoHugging Face27chiayewken /flan-v2 Dataset Card for "flan-v2" More Information needed text10M<n<100M4 likes3.5k downloads3y agoHugging Face28Morton-Li /ChineseWebText2.0-HighQuality 📘 ChineseWebText2.0-HighQuality Overview ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License). This subset retains only samples with: quality_score ≥ 0.9 toxicity.score ≤ 0.01 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, and quality-sensitive downstream tasks. This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.texttext-generation100M<n<1B4 likes3.4k downloads7mo agoHugging Face29CASIA-LM /ChineseWebText2.0 ChineseWebText 2.0: Large-Scale High-quality Chinese Web Text with Multi-dimensional and fine-grained information This directory contains the ChineseWebText2.0 dataset, and a new tool-chain called MDFG-tool for constructing large-scale and high-quality Chinese datasets with multi-dimensional and fine-grained information. Our ChineseWebText2.0 code is publicly available on github (here). ChineseWebText2.0 Dataset Overview We have released the latest… See the full description on the dataset page: https://huggingface.co/datasets/CASIA-LM/ChineseWebText2.0.text1K<n<10K34 likes3.1k downloads2y agoHugging Face30shareAI /ShareGPT-Chinese-English-90k ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss) Features: Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.question-answering10K<n<100K288 likes2.9k downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.