china
Datasets
All datasets matching “china”china-a-share-1min-ohlcv
China A-Share Equities 1-Minute OHLCV
Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports.
Dataset summary
This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.china-a-share-l2-level2-limit-order-book-tick-data
China A-Share Level-2 Archive
2017–2026 · Quotes, orders and trades · Parquet
A historical archive of Chinese exchange Level-2 data, supplied through a vendor export.
It includes ten-level quote snapshots, individual order messages and trade-stream records.
The files cover A-share stocks and non-stock instruments such as ETFs and bonds.
“Full-market” describes the export's scope, not a guarantee that every instrument or message is present.
中国证券市场 Level-2… See the full description on the dataset page: https://huggingface.co/datasets/venvoo/china-a-share-l2-level2-limit-order-book-tick-data.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.china
Decompression
If you need to decompress the files, please see the main README at the github repo.
If you want to use them directly from the parquet files, the original .tif/.nc files were read into the rows as binary file data sources
China AOI
We provide '.csv' files with predefined train (blue), test (orange) and validation (green) splits that can be used for repeatability and comparability of experiments.
60% of tiles are allocated for training, 20% for validation… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO/china.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.china-textbook-2021-hf
China Textbook 2021
中国教育部审定的中小学教科书 PDF 合集。
简介
本数据集包含从小学到高中各年级的教育部审定教科书 PDF。原始数据来自 TapXWorld/ChinaTextbook,由本项目自动合并碎片化文件后托管在 Hugging Face Datasets 上。
处理代码:https://github.com/rainewhk/china-textbook-2021-hf
数据结构
教材按学段、学科、版本、年级分层存放:
小学/
├── ...
初中/
├── ...
初中(五•四学制)/
├── ...
小学(五•四学制)/
├── ...
高中/
└── ...
文件名为 义务教育教科书·<学科> <年级> <册>册.pdf 或类似格式。
处理说明
上游 PDF 存在碎片化存储(.1, .2 等分片),本项目通过 Rust 合并器 自动合并为完整文件。
部分教材同时存在完整 PDF 和 _merge_folder… See the full description on the dataset page: https://huggingface.co/datasets/RainPPR/china-textbook-2021-hf.
