datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiArithchitralekha
Chitralekha
Dataset Details
Dataset Version
Some of the fonts do not have proper letters/rendering of different telugu letter combinations. Those have been removed as much as I can find them. If there are any other mistakes that you notice, please raise an issue and I will try my best to look into it
Dataset Description
This extensive dataset, hosted on Huggingface, is a comprehensive resource for Optical Character Recognition (OCR) in the Telugu… See the full description on the dataset page: https://huggingface.co/datasets/gksriharsha/chitralekha.SVAMPchina-a-share-1min-ohlcv
China A-Share Equities 1-Minute OHLCV
Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports.
Dataset summary
This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.Fineweb-Edu-Chinese-V2.1
Chinese Fineweb Edu Dataset V2.1 [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.1.china-a-share-l2-level2-limit-order-book-tick-data
China A-Share Level-2 Archive
2017–2026 · Quotes, orders and trades · Parquet
A historical archive of Chinese exchange Level-2 data, supplied through a vendor export.
It includes ten-level quote snapshots, individual order messages and trade-stream records.
The files cover A-share stocks and non-stock instruments such as ETFs and bonds.
“Full-market” describes the export's scope, not a guarantee that every instrument or message is present.
中国证券市场 Level-2… See the full description on the dataset page: https://huggingface.co/datasets/venvoo/china-a-share-l2-level2-limit-order-book-tick-data.chinese-fineweb-edu
This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 !
Chinese Fineweb Edu Dataset [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.Fineweb-Edu-Chinese-V2.2
Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train)
[[中文]] | [[English]]
OpenCSG Community | 👾 GitHub | 📖 Technical Report
Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs
Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain.
This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.2.StrategyQAMMLU_ChineseChinese version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.ChildGaitDecoding Children's Gait Behavior
ECCV 2026
Yifan Shen1,2,*,
Boyi Li1,*,
Meihuan Huang2,3,4,*,
Yuanzhe Liu1,*,
Xu Cao1,2,*,§,
Jinyang Jin1,
Zhengyuan Li1,
Anglin Liu5,
Junho Kim1,
Jingyuan Zhu2,
Fangzhou Lan2,
Jianguo Cao2,3,
Jintai Chen5,
Ismini Lourentzou1,
James M. Rehg1,†
1 University of Illinois Urbana-Champaign
2 PediaMed AI
3 Shenzhen Children's Hospital
4 Hong Kong Polytechnic University… See the full description on the dataset page: https://huggingface.co/datasets/PediaMedAI/ChildGait.telegram-audiobook-chizzled
Telegram Persian Audiobook Chizzled
1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release
This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.china
Decompression
If you need to decompress the files, please see the main README at the github repo.
If you want to use them directly from the parquet files, the original .tif/.nc files were read into the rows as binary file data sources
China AOI
We provide '.csv' files with predefined train (blue), test (orange) and validation (green) splits that can be used for repeatability and comparability of experiments.
60% of tiles are allocated for training, 20% for validation… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO/china.alpaca-data-gpt4-chineseChineseWebText
ChineseWebText: Large-Scale High-quality Chinese Web Text Extracted with Effective Evaluation Model
This directory contains the ChineseWebText dataset, and the EvalWeb tool-chain to process CommonCrawl Data. Our EvalWeb tool is publicly available on github https://github.com/CASIA-LM/ChineseWebText.
ChineseWebText
Dataset Overview
We release the latest and largest Chinese dataset ChineseWebText, which consists of 1.42 TB data and each text is assigned a… See the full description on the dataset page: https://huggingface.co/datasets/CASIA-LM/ChineseWebText.visual_robust_robocasa_xUIBenchKitpersian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.fineweb-edu
📚 FineWeb-Edu
1.3 trillion tokens of the finest educational data the 🌐 web has to offer
Paper: https://arxiv.org/abs/2406.17557
What is it?
📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version.
To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/fineweb-edu.chinese_clean_passages_80m
chinese_clean_passages_80m
包含8千余万(88328203)个纯净中文段落,不包含任何字母、数字。Containing more than 80 million pure & clean Chinese passages, without any letters/digits/special tokens.
文本长度大部分介于50~200个汉字之间。The passage length is approximately 50~200 Chinese characters.
通过datasets.load_dataset()下载数据,会产生38个大小约340M的数据包,共约12GB,所以请确保有足够空间。Downloading the dataset will result in 38 data shards each of which is about 340M and 12GB in total. Make sure there's enough space in your device:)
>>>
passage_dataset… See the full description on the dataset page: https://huggingface.co/datasets/beyond/chinese_clean_passages_80m.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.chinese-novel-nonH-collect
Dataset Card for Dataset Name
license: cc0-1.0
task_categories:
- text-classification
- summarization
language:
- zh
tags:
- art
size_categories:
- 100M<n<1B
chimera-cs2
Chimera CS2 Dataset
Labeled Counter-Strike 2 screenshots for vision-language model training.
Each sample has:
A CS2 gameplay screenshot
Ground truth JSON with game_state, analysis, and advice
Structure
screenshots/ # PNG/JPG images
labels/ # Matching JSON files (same stem name)
manifest.jsonl # Data provenance tracking
Usage
from datasets import load_dataset
ds = load_dataset("skkwowee/chimera-cs2")
Stats
Labels: 5309… See the full description on the dataset page: https://huggingface.co/datasets/skkwowee/chimera-cs2.china-textbook-2021-hf
China Textbook 2021
中国教育部审定的中小学教科书 PDF 合集。
简介
本数据集包含从小学到高中各年级的教育部审定教科书 PDF。原始数据来自 TapXWorld/ChinaTextbook,由本项目自动合并碎片化文件后托管在 Hugging Face Datasets 上。
处理代码:https://github.com/rainewhk/china-textbook-2021-hf
数据结构
教材按学段、学科、版本、年级分层存放:
小学/
├── ...
初中/
├── ...
初中(五•四学制)/
├── ...
小学(五•四学制)/
├── ...
高中/
└── ...
文件名为 义务教育教科书·<学科> <年级> <册>册.pdf 或类似格式。
处理说明
上游 PDF 存在碎片化存储(.1, .2 等分片),本项目通过 Rust 合并器 自动合并为完整文件。
部分教材同时存在完整 PDF 和 _merge_folder… See the full description on the dataset page: https://huggingface.co/datasets/RainPPR/china-textbook-2021-hf.flan-v2
Dataset Card for "flan-v2"
More Information needed
ChineseWebText2.0-HighQuality
📘 ChineseWebText2.0-HighQuality
Overview
ChineseWebText2.0-HighQuality is a high-quality filtered subset of the original
CASIA-LM/ChineseWebText2.0 dataset (Apache-2.0 License).
This subset retains only samples with:
quality_score ≥ 0.9
toxicity.score ≤ 0.01
The goal is to provide a cleaner and more reliable dataset suitable for
language model pre-training, instruction tuning, and quality-sensitive downstream tasks.
This work is independent and not affiliated with the… See the full description on the dataset page: https://huggingface.co/datasets/Morton-Li/ChineseWebText2.0-HighQuality.ChineseWebText2.0
ChineseWebText 2.0: Large-Scale High-quality Chinese Web Text with Multi-dimensional and fine-grained information
This directory contains the ChineseWebText2.0 dataset, and a new tool-chain called MDFG-tool for constructing large-scale and high-quality Chinese datasets with multi-dimensional and fine-grained information. Our ChineseWebText2.0 code is publicly available on github (here).
ChineseWebText2.0
Dataset Overview
We have released the latest… See the full description on the dataset page: https://huggingface.co/datasets/CASIA-LM/ChineseWebText2.0.ShareGPT-Chinese-English-90k
ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset
A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss)
Features:
Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.
