CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ChilleD /MultiArithtextn<1K17 likes103k downloads3y agoHugging Face02gksriharsha /chitralekha Chitralekha Dataset Details Dataset Version Some of the fonts do not have proper letters/rendering of different telugu letter combinations. Those have been removed as much as I can find them. If there are any other mistakes that you notice, please raise an issue and I will try my best to look into it Dataset Description This extensive dataset, hosted on Huggingface, is a comprehensive resource for Optical Character Recognition (OCR) in the Telugu… See the full description on the dataset page: https://huggingface.co/datasets/gksriharsha/chitralekha.imageimage-to-text10M<n<100M5 likes87k downloads2y agoHugging Face03ChilleD /SVAMPtexttext-generation1K<n<10K23 likes86k downloads2y agoHugging Face04neigezhu /china-a-share-1min-ohlcv China A-Share Equities 1-Minute OHLCV Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports. Dataset summary This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.tabulartime-series-forecasting100K<n<1M12 likes64k downloads1mo agoHugging Face05opencsg /chinese-fineweb-edu This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.texttext-generation10M<n<100M117 likes18k downloads10mo agoHugging Face06ChilleD /StrategyQAtext1K<n<10K6 likes16k downloads3y agoHugging Face07DannHiroaki /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes12k downloads8mo agoHugging Face08Reza2kn /telegram-audiobook-chizzled Telegram Persian Audiobook Chizzled 1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.audion<1K4 likes6.7k downloads14d agoHugging Face09actava /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.documenttext-generationn<1K61 likes6.3k downloads4mo agoHugging Face10M3LEO /china Decompression If you need to decompress the files, please see the main README at the github repo. If you want to use them directly from the parquet files, the original .tif/.nc files were read into the rows as binary file data sources China AOI We provide '.csv' files with predefined train (blue), test (orange) and validation (green) splits that can be used for repeatability and comparability of experiments. 60% of tiles are allocated for training, 20% for validation… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO/china.text1M<n<10M1 likes6.2k downloads2y agoHugging Face11silk-road /alpaca-data-gpt4-chinesetexttext-generation10K<n<100K104 likes5.5k downloads3y agoHugging Face12CASIA-LM /ChineseWebText ChineseWebText: Large-Scale High-quality Chinese Web Text Extracted with Effective Evaluation Model This directory contains the ChineseWebText dataset, and the EvalWeb tool-chain to process CommonCrawl Data. Our EvalWeb tool is publicly available on github https://github.com/CASIA-LM/ChineseWebText. ChineseWebText Dataset Overview We release the latest and largest Chinese dataset ChineseWebText, which consists of 1.42 TB data and each text is assigned a… See the full description on the dataset page: https://huggingface.co/datasets/CASIA-LM/ChineseWebText.text1K<n<10K45 likes5.5k downloads3y agoHugging Face13Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.9k downloads2mo agoHugging Face14chilomax /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/fineweb-edu.tabulartext-generation1B<n<10B0 likes4.4k downloads3mo agoHugging Face15beyond /chinese_clean_passages_80m chinese_clean_passages_80m 包含8千余万(88328203)个纯净中文段落,不包含任何字母、数字。Containing more than 80 million pure & clean Chinese passages, without any letters/digits/special tokens. 文本长度大部分介于50~200个汉字之间。The passage length is approximately 50~200 Chinese characters. 通过datasets.load_dataset()下载数据,会产生38个大小约340M的数据包,共约12GB,所以请确保有足够空间。Downloading the dataset will result in 38 data shards each of which is about 340M and 12GB in total. Make sure there's enough space in your device:) >>> passage_dataset… See the full description on the dataset page: https://huggingface.co/datasets/beyond/chinese_clean_passages_80m.text10M<n<100M30 likes4.3k downloads4y agoHugging Face16liuhangbiao /China-Building-Footprints-CMAB-Mirror Origin Data @misc{Zhang2025CMAB, author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying}, title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}}, year = {2025}, month = apr, publisher = {figshare}, doi = {10.6084/m9.figshare.27992417}, url = {https://doi.org/10.6084/m9.figshare.27992417}, howpublished = {dataset} } Paper @article{Zhang2025SciData, author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.geospatialn<1K0 likes3.8k downloads6mo agoHugging Face17lainka0o0 /chinese-novel-nonH-collect Dataset Card for Dataset Name license: cc0-1.0 task_categories: - text-classification - summarization language: - zh tags: - art size_categories: - 100M<n<1B texttext-classification100M<n<1B6 likes3.7k downloads1y agoHugging Face18chiayewken /flan-v2 Dataset Card for "flan-v2" More Information needed text10M<n<100M4 likes3.5k downloads3y agoHugging Face19Morton-Li /ChineseWebText2.0-HighQuality 📘 ChineseWebText2.0-HighQuality Overview ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License). This subset retains only samples with: quality_score ≥ 0.9 toxicity.score ≤ 0.01 The goal is to provide a cleaner and more reliable dataset suitable for language model pre-training, instruction tuning, and quality-sensitive downstream tasks. This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.texttext-generation100M<n<1B4 likes3.4k downloads7mo agoHugging Face20CASIA-LM /ChineseWebText2.0 ChineseWebText 2.0: Large-Scale High-quality Chinese Web Text with Multi-dimensional and fine-grained information This directory contains the ChineseWebText2.0 dataset, and a new tool-chain called MDFG-tool for constructing large-scale and high-quality Chinese datasets with multi-dimensional and fine-grained information. Our ChineseWebText2.0 code is publicly available on github (here). ChineseWebText2.0 Dataset Overview We have released the latest… See the full description on the dataset page: https://huggingface.co/datasets/CASIA-LM/ChineseWebText2.0.text1K<n<10K34 likes3.1k downloads2y agoHugging Face21chiuratto-AIgourakis /sounio-code-examples Sounio Curated Code Examples Curated compile-clean .sio examples for training and evaluating code models on Sounio, a self-hosted systems and scientific programming language for epistemic computing, uncertainty propagation, and algebraic effects. This directory is the Cx-1 expansion lane for chiuratto-AIgourakis/sounio-code-examples. Current batch Examples: 5,000 Metadata files: 5,000 Compiler gate: bin/souc check pass rate 5,000/5,000 Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.texttext-generation1K<n<10K0 likes2.9k downloads4mo agoHugging Face22jed351 /Traditional-Chinese-Common-Crawl-Filtered Traditional Chinese C4 Dataset Summary Data obtained from 2013~2025 Common Crawl. Downloaded and processed using code based on another project attempting to recreate the C4 dataset. The resultant dataset contains both simplified and traditional Chinese, which could be found here. It was then filtered using a modified list of simplified Chinese characters to obtain this traditional Chinese dataset. Unfortunately, I don't have enough funding to run a deduplication across… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Traditional-Chinese-Common-Crawl-Filtered.text100M<n<1B26 likes2.8k downloads1y agoHugging Face23jed351 /Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese. The hash based cleaned dataset can be found here. Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow) text100M<n<1B0 likes2.5k downloads1y agoHugging Face24OpenStellarTeam /Chinese-SimpleQA Overview 🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper • 📊 Leaderboard Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, our benchmark covers 6 major topics with 99 diverse subtopics. Please visit our website or check our paper for more details.… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SimpleQA.textquestion-answering1K<n<10K38 likes2.5k downloads2y agoHugging Face25opencsg /chinese-fineweb-edu-v2 This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset V2 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu Dataset V2 is a comprehensive upgrade of the original Chinese Fineweb Edu, designed and optimized for natural language processing (NLP) tasks in the education sector. This high-quality Chinese pretraining dataset has… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu-v2.tabulartext-generation100M<n<1B75 likes2.4k downloads10mo agoHugging Face26vessel888 /china-a-share-1min-ohlcv China A-Share Equities 1-Minute OHLCV Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports. Dataset summary This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/china-a-share-1min-ohlcv.tabulartime-series-forecasting100K<n<1M1 likes2.4k downloads16d agoHugging Face27hails /agieval-gaokao-chinese Dataset Card for "agieval-gaokao-chinese" Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub. This dataset contains the contents of the Gaokao Chinese subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 . Citation: @misc{zhong2023agieval, title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-chinese.textn<1K2 likes2.4k downloads3y agoHugging Face28BAAI /IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine IndustryCorpus2: Health & Medicine This repository contains the IndustryCorpus2: Health & Medicine domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_medicine_health_psychology_traditional_chinese_medicine.tabular10M<n<100M11 likes2.3k downloads1mo agoHugging Face29shjwudp /chinese-c4 Introduction Chinese-C4 is a clean Chinese internet dataset based on Common Crawl. The dataset is 46.29GB and has undergone multiple cleaning strategies, including Chinese filtering, heuristic cleaning based on punctuation, line-based hashing for deduplication, and repetition removal. The dataset is open source and free for commercial use, and you are welcome to use the data and the cleaning strategies provided and contribute your cleaning strategies. You can find the cleaning… See the full description on the dataset page: https://huggingface.co/datasets/shjwudp/chinese-c4.text1M<n<10M35 likes2.3k downloads3y agoHugging Face30chiayewken /competition_mathtext10K<n<100K0 likes2.1k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.