datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
china-a-share-1min-ohlcv
China A-Share Equities 1-Minute OHLCV
Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports.
Dataset summary
This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.china
Decompression
If you need to decompress the files, please see the main README at the github repo.
If you want to use them directly from the parquet files, the original .tif/.nc files were read into the rows as binary file data sources
China AOI
We provide '.csv' files with predefined train (blue), test (orange) and validation (green) splits that can be used for repeatability and comparability of experiments.
60% of tiles are allocated for training, 20% for validation… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO/china.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.china-a-share-1min-ohlcv
China A-Share Equities 1-Minute OHLCV
Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports.
Dataset summary
This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/china-a-share-1min-ohlcv.ChinaTravel
ChinaTravel Query Dataset
This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).
ChinaTravel is an open-ended travel-planning benchmark with compositional
constraint validation for language agents. See the
paper,
Hugging Face paper page,
code, and
bilingual sandbox database
(ModelScope mirror)
for the complete benchmark resources.
Introduction
For a given query, a language agent uses the sandbox tools to collect
information and… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel.cail2018
Dataset Card for CAIL 2018
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/china-ai-law-challenge/cail2018.china-etf-1min-ohlcv
China Exchange-Traded Funds 1-Minute OHLCV
Minute-level OHLCV bars for selected exchange-listed Chinese ETFs. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports.
Dataset summary
This snapshot contains 4,058,681 rows for 7 instruments across selected exchange-listed Chinese ETFs. It covers 2015-01-05 09:30:00 through 2026-08-20 15:00:00. Prices are unadjusted. Volume is stored in shares, turnover in… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-etf-1min-ohlcv.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.China-Stock-Symbols-and-Metadata
China Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in China.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/China-Stock-Symbols-and-Metadata.ChinaTravel-Sandbox
ChinaTravel Sandbox Environment Database
This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).
English | 简体中文
Release version: 2026.08.2
English
This dataset contains the bilingual static sandbox used by
ChinaTravel. It is a companion to
the ChinaTravel query dataset
and an artifact of the
ChinaTravel paper.
The raw ZIP snapshots preserve the exact directory layout expected by the
ChinaTravel evaluator. Viewer-friendly Parquet… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel-Sandbox.Tick-by-Tick-Orders-China
Tick-by-Tick Orders and Trades — China A-Shares (2025)
Full-depth Level-2 message data for every listed stock, fund and ETF on the
Shanghai (.SH) and Shenzhen (.SZ) exchanges, 2025.
This repository holds one calendar year. The remaining years live in sibling
repositories, one per year:
Year
Repository
2023
alphat01/Tick-by-Tick-Orders-China
2024
alphat02/Tick-by-Tick-Orders-China
2025
alphat03/Tick-by-Tick-Orders-China
2026
alphat04/Tick-by-Tick-Orders-China… See the full description on the dataset page: https://huggingface.co/datasets/alphat03/Tick-by-Tick-Orders-China.Tick-by-Tick-Orders-China
Tick-by-Tick Orders and Trades — China A-Shares (2023)
Full-depth Level-2 message data for every listed stock, fund and ETF on the
Shanghai (.SH) and Shenzhen (.SZ) exchanges, 2023.
This repository holds one calendar year. The remaining years live in sibling
repositories, one per year:
Year
Repository
2023
alphat01/Tick-by-Tick-Orders-China
2024
alphat02/Tick-by-Tick-Orders-China
2025
alphat03/Tick-by-Tick-Orders-China
2026
alphat04/Tick-by-Tick-Orders-China… See the full description on the dataset page: https://huggingface.co/datasets/alphat01/Tick-by-Tick-Orders-China.Tick-by-Tick-Orders-China
Tick-by-Tick Orders and Trades — China A-Shares (2024)
Full-depth Level-2 message data for every listed stock, fund and ETF on the
Shanghai (.SH) and Shenzhen (.SZ) exchanges, 2024.
This repository holds one calendar year. The remaining years live in sibling
repositories, one per year:
Year
Repository
2023
alphat01/Tick-by-Tick-Orders-China
2024
alphat02/Tick-by-Tick-Orders-China
2025
alphat03/Tick-by-Tick-Orders-China
2026
alphat04/Tick-by-Tick-Orders-China… See the full description on the dataset page: https://huggingface.co/datasets/alphat02/Tick-by-Tick-Orders-China.smoltalk-chinese-QwQ-Distrill
smoltalk-chinese-QwQ-Distrill [中文] [English]
📖Technical Report
smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.ChinaPaint
CCPP
Evaluating and Benchmarking Classical Chinese Poetry-to-Painting for Multimodal Large Language Models
CCPP: Classical Chinese Poetry-to-Painting Project
A comprehensive project supporting the research on Classical Chinese Poetry-to-Painting (CCPP) generation and evaluation, including benchmark datasets, human painting references, model outputs, and auxiliary scripts. This project serves as the official code & data repository for the corresponding academic… See the full description on the dataset page: https://huggingface.co/datasets/busy-pig/ChinaPaint.ipfs_china_laws_ir
China legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_china_laws (revision 416a1b3eac0d8f0708dd7871a72dfdf171f8b0ba) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of China prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_china_laws_ir.china-city-render
VisNavMap City Offline Render Snapshots
This dataset contains one offline-render cache per city, derived from the
matching city-level OpenStreetMap PBF from the Hugging Face dataset
86Cao/china-city-osm-pbf.
It lets VisNavMap render
global and local map observations without parsing or querying the source PBF
during rollout.
Data provenance
The data-generation chain is:
Hugging Face: 86Cao/china-city-osm-pbf
│
│ download city-level… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/china-city-render.ChinaOpen1k-T2V
ChinaOpen-1k T2V retrieval
Multilingual video retrieval built from
ChinaOpen-1k:
1,092 Bilibili videos with human-written Chinese captions and their
English translations. Both language configs share the same videos, so
the two subsets form a controlled comparison.
Built by scripts/data/chinaopen_retrieval/create_data.py in
mteb.
Please cite the original dataset:
@inproceedings{chen2023chinaopen,
title = {ChinaOpen: A Dataset for Open-world Multimodal Learning},
author =… See the full description on the dataset page: https://huggingface.co/datasets/shriyasudhakar/ChinaOpen1k-T2V.Tick-by-Tick-Orders-China
Tick-by-Tick Orders and Trades — China A-Shares (2026)
Full-depth Level-2 message data for every listed stock, fund and ETF on the
Shanghai (.SH) and Shenzhen (.SZ) exchanges, 2026.
This repository holds one calendar year. The remaining years live in sibling
repositories, one per year:
Year
Repository
2023
alphat01/Tick-by-Tick-Orders-China
2024
alphat02/Tick-by-Tick-Orders-China
2025
alphat03/Tick-by-Tick-Orders-China
2026
alphat04/Tick-by-Tick-Orders-China… See the full description on the dataset page: https://huggingface.co/datasets/alphat04/Tick-by-Tick-Orders-China.ChinaOpen1k-V2T
ChinaOpen-1k V2T retrieval
Multilingual video retrieval built from
ChinaOpen-1k:
1,092 Bilibili videos with human-written Chinese captions and their
English translations. Both language configs share the same videos, so
the two subsets form a controlled comparison.
Built by scripts/data/chinaopen_retrieval/create_data.py in
mteb.
Please cite the original dataset:
@inproceedings{chen2023chinaopen,
title = {ChinaOpen: A Dataset for Open-world Multimodal Learning},
author =… See the full description on the dataset page: https://huggingface.co/datasets/shriyasudhakar/ChinaOpen1k-V2T.AQuA-RATA jsonlines dataset of 98000 prompt-completion pairs for algebra questions.
The prompt has a question and the Completion has the answer with rationale.
Originally taken from https://www.deepmind.com/open-source/aqua-rat for finetuning GPT-3 but use it for jobs of your choice.
For questions, open a discussion on community.
Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources
Torah Codes Religion Texts Sources
Data Tree
── arabs
│ ├── astrological_stelar_magic.txt
│ └── Holy-Quran-English.txt
├── ars
│ ├── ars_magna_ramon_llull.txt
│ └── lemegeton_book_solomon.txt
├── asimov
│ ├── foundation.txt
│ └── prelude_to_foundation.txt
├── budist
│ ├── bardo_todhol_book_of_deads_tibet_libro_tibetano_de_los_muertos.txt
│ ├── rig_veda.txt
│ └── TheTeachingofBuddha.txt
├── cathars
├── china
│ ├── arte_de_la_guerra_art_of_war.txt
│… See the full description on the dataset page: https://huggingface.co/datasets/torahCodes/Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources.Halal-Food-In-China
Halal Food In China RAG Corpus 🕌🍜
Dataset Overview
The Halal Food In China RAG Corpus is an expertly curated, highly structured, and multi-format dataset targeting the intersection of Chinese culinary traditions, Hui Muslim history, and Islamic dietary laws (Halal). It acts as an authoritative ground-truth database to mitigate Large Language Model (LLM) hallucinations regarding minority Islamic culture in China.
[!TIP]
Human Readers: Looking for the full… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Halal-Food-In-China.china-etf-1min-ohlcv
China Exchange-Traded Funds 1-Minute OHLCV
Minute-level OHLCV bars for selected exchange-listed Chinese ETFs. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports.
Dataset summary
This snapshot contains 4,058,681 rows for 7 instruments across selected exchange-listed Chinese ETFs. It covers 2015-01-05 09:30:00 through 2026-08-20 15:00:00. Prices are unadjusted. Volume is stored in shares, turnover in… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/china-etf-1min-ohlcv.China-Halal-Restaurant
China Halal Restaurant Dataset (RAG Optimized) 🕌
This is a rigorously formatted Chinese Halal Restaurant corpus containing 201 authentic articles and travel guides. It is explicitly optimized for Retrieval-Augmented Generation (RAG) and pure text indexing. The data was explicitly designed to pass Hugging Face's Dataset Viewer standards natively by using optimal Parquet partitioning.
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/China-Halal-Restaurant.ipfs_china_laws
China In-Force National Laws (NPC 现行有效法律目录)
Research snapshot of in-force national 法律 listed in the NPC
现行有效法律目录
(cutoff 2026-08-15, advertised 301). Bodies come from official NPC 权威发布 /
法律文件 HTML, NPC PDFs, ministry .gov.cn (CMA/NRA/PBC/MOF/SPC 公报), and gov.cn
pages; when live fetch failed, Internet Archive Wayback snapshots of the same
official npc.gov.cn / gov.cn URLs (labeled archive-of-official).
flk.npc.gov.cn was not crawled (robots.txt Disallow: /). pkulaw, lawinfochina… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_china_laws.camera-calib-and-scene-alignment-datachina-city-graphml
VisNavMap City GraphML Snapshots
This dataset contains the city-level navigation graph cache derived from the
city-level OpenStreetMap PBF files released in the Hugging Face dataset
86Cao/china-city-osm-pbf.
It is intended to make
VisNavMap training, rollout, and evaluation reproducible without rebuilding a
graph from a PBF at runtime.
Data provenance
The data-generation chain is:
Hugging Face: 86Cao/china-city-osm-pbf
│
│ download city-level… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/china-city-graphml.
