datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChinaTravel
ChinaTravel Query Dataset
This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).
ChinaTravel is an open-ended travel-planning benchmark with compositional
constraint validation for language agents. See the
paper,
Hugging Face paper page,
code, and
bilingual sandbox database
(ModelScope mirror)
for the complete benchmark resources.
Introduction
For a given query, a language agent uses the sandbox tools to collect
information and… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel.chimera-bench
CHIMERA-Bench v1.0
A unified benchmark for epitope-specific antibody CDR sequence-structure co-design.
Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop)
Code: github.com/mansoorbaloch/chimera-bench
Dataset Summary
Property
Value
Complexes
2,922
PDB structures
2,721
Pre-computed features
2,941 .pt files
Splits
3 (epitope-group, antigen-fold, temporal)
Numbering schemes
IMGT, Chothia
Contact… See the full description on the dataset page: https://huggingface.co/datasets/Baoruixi/chimera-bench.chimera-bench
CHIMERA-Bench v1.0
A unified benchmark for epitope-specific antibody CDR sequence-structure co-design.
Paper: CHIMERA-Bench: A Benchmark Dataset for Epitope-Specific Antibody Design (ICLR 2026 GEM Workshop)
Code: github.com/mansoorbaloch/chimera-bench
Dataset Summary
Property
Value
Complexes
2,922
PDB structures
2,721
Pre-computed features
2,941 .pt files
Splits
3 (epitope-group, antigen-fold, temporal)
Numbering schemesIMGT, Chothia
Contact… See the full description on the dataset page: https://huggingface.co/datasets/mansoorbaloch/chimera-bench.ipa-childes-split
IPA-CHILDES split
This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular,
the following changes have been implemented:
column processed_gloss dropped as it duplicates information of gloss up to punctuation
column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+)
column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package
columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.chile-seismological-records
Chile's Seismological Records
ChinaPaint
CCPP
Evaluating and Benchmarking Classical Chinese Poetry-to-Painting for Multimodal Large Language Models
CCPP: Classical Chinese Poetry-to-Painting Project
A comprehensive project supporting the research on Classical Chinese Poetry-to-Painting (CCPP) generation and evaluation, including benchmark datasets, human painting references, model outputs, and auxiliary scripts. This project serves as the official code & data repository for the corresponding academic… See the full description on the dataset page: https://huggingface.co/datasets/busy-pig/ChinaPaint.health-conditions-among-children-under-age-18-by-s
Health conditions among children under age 18, by selected characteristics: United States
Description
NOTE: On October 19, 2021, estimates for 2016–2018 by health insurance status were revised to correct errors. Changes are highlighted and tagged at https://www.cdc.gov/nchs/data/hus/2019/012-508.pdf
Data on health conditions among children under age 18, by selected population characteristics. Please refer to the PDF or Excel version of this table in the HUS 2019 Data… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/health-conditions-among-children-under-age-18-by-s.IPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.chinese-ai-detection-dataset
Chinese AI Detection Dataset
中文AI文本检测数据集
数据集简介
用于训练中文AI生成文本检测模型的综合数据集,包含纯人类、纯AI以及混合文本(人类+AI)。
核心特色:使用[SEP]标记显式标注混合文本的人类/AI边界。
数据统计
类型
样本数
说明
总计
66,001
训练/验证/测试集
纯人类
27,719
多领域人类文本
纯AI
27,719
多模型生成
C2 (续写)
3,781
人类开头+AI续写
C3 (改写)
3,781
AI改写人类文本
C4 (润色)
3,001
AI润色人类文本
数据格式
{
"text": "文本内容(混合文本包含[SEP]标记)",
"label": 0, // 0=Human, 1=AI
"category": "C2", // Human/AI/C2/C3/C4
"source": "数据来源"
}… See the full description on the dataset page: https://huggingface.co/datasets/AnxForever/chinese-ai-detection-dataset.reddew
reddew
Reddit Download and Datasets
AIDE-Chip-15K-gem5-Sims
AIDE-Chip 15K gem5 Simulation Dataset
AIDE-Chip-15K-gem5-Sims is a structured dataset of approximately 15,000 validated RISC-V gem5 simulations covering cache hierarchy design-space exploration (DSE) for single-core processors.
The dataset was generated using gem5's Syscall Emulation (SE) mode and six representative workloads, spanning compute-bound, memory-bound, and irregular access patterns. Each sample maps cache configuration parameters to IPC and L2 miss rate, enabling… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/AIDE-Chip-15K-gem5-Sims.chinese-stock-datasetChinese-Student-English-Essay
Dataset Card for Chinese Student English Essay (CSEE) Dataset
Dataset Summary
The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts.
Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/Xiaochr/Chinese-Student-English-Essay.IPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/superBigPigeon/IPA-CHILDES.COLING-2025-CHIPSAL
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-CHIPSAL.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.chinese-official-source-reachability
Reachability of Chinese official verification portals from outside China
If you are doing due diligence on a Chinese supplier, the advice is always "check the official registry". This dataset measures whether you can.
Eight official Chinese verification sources were measured from public vantage points outside mainland China across several rounds between August and September 2026, plus two controls (www.gov.cn and www.baidu.com).
The result is not "Chinese government sites are… See the full description on the dataset page: https://huggingface.co/datasets/derrick459/chinese-official-source-reachability.china-myeloma-clinical-trials
China Multiple Myeloma Clinical Trials — Open Dataset
Multiple myeloma clinical trials registered in China, curated from official NMPA / CDE filings by the China Myeloma Digital Network (CMDN), an independent non-profit patient advocacy organisation.
This is a mirror. The citable version of record lives at doi.org/10.5281/zenodo.22690814; the source repository is chinamyeloma/china-myeloma-clinical-trials; the documentation is at chinamyeloma.org.
Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/chinamyeloma/china-myeloma-clinical-trials.9-sao-chieu-menh
9 sao chiếu mệnh (Cửu Diệu)
The nine presiding stars (Cuu Dieu)
1. Mô tả · Description
Chín sao kèm phân loại, diễn giải và các tuổi mụ tương ứng, tách riêng nam và nữ.
The nine stars with their classification, reading, and the lunar ages they fall on, listed separately for men and women.
Số dòng · Rows: 9
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa ·… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/9-sao-chieu-menh.12-dia-chi
12 Địa Chi
The twelve earthly branches
1. Mô tả · Description
Mười hai Địa Chi kèm con giáp, ngũ hành và khung giờ hai tiếng tương ứng.
The twelve earthly branches with their zodiac animal, five-element attribution and two-hour window.
Số dòng · Rows: 12
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa · Meaning
id
string
Định danh ổn định của dòng, không… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/12-dia-chi.22-an-chinh-tarot
22 lá Ẩn Chính Tarot
The 22 Major Arcana of the Tarot
1. Mô tả · Description
Bộ Major Arcana kèm tên Việt và Anh, từ khoá, nghĩa xuôi và nghĩa ngược.
The Major Arcana with Vietnamese and English names, keywords, and upright and reversed meanings.
Số dòng · Rows: 22
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa · Meaning
id
string
Định danh ổn định của… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/22-an-chinh-tarot.RingABell-NudityThe files contain inversing prompts for nudity generated by Ring-A-Bell.
Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Chima207/Goodreads-Books.separate-chip-enrollment-by-month-and-state-histor
Separate CHIP Enrollment by Month and State – Historic CAA/Unwinding Period
Description
This historic dataset with total enrollment in separate CHIP programs by month and state was created to fulfill reporting requirements under section 1902(tt)(1) of the Social Security Act, which was added by section 5131(b) of subtitle D of title V of division FF of the Consolidated Appropriations Act, 2023 (P.L. 117-328) (CAA, 2023). For each month from April 1, 2023, through June 30… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/separate-chip-enrollment-by-month-and-state-histor.numbers-in-chinese-poetry
Numerals in Classical Chinese Poetry (Tang–Qing)
Zizheng Lv · ORCID 0009-0004-0327-4477 · DOI 10.5281/zenodo.22734274
Code and documentation: https://github.com/zizhenglvubc/numbers-in-chinese-poetry
Dataset summary
Numeral counts for 175,355 classical Chinese poems from the Tang dynasty through
the Qing. Two files:
poem_level_numerals.csv.gz — one row per poem, 175,355 rows.
numeral_rates_by_group.csv — one row per collection, 7 rows including a total.
59.24%… See the full description on the dataset page: https://huggingface.co/datasets/zizhenglvubc/numbers-in-chinese-poetry.Chicago-Crime-Datasetchinese_lyricsdata_for_China-4-factor-model-reccurence
中国股票风格因子构造与复现
本项目是一次机器学习量化投资课程作业,核心目标是基于中国 A 股月频数据,按照 LSY (2019) 风格构造并分析一组风格因子。项目最终在 hwfinal.ipynb 中完成数据清洗、因子构造、统计汇总和可视化,并生成可用于后续研究的月度因子序列。
当前仓库已经包含整理后的主数据文件 completedata.csv 和完整实验 notebook hwfinal.ipynb,适合直接阅读思路、复现实验主流程,或在此基础上继续扩展。
项目内容
hwfinal.ipynb:项目主 notebook,包含数据预处理、因子构造、统计表输出与累计收益绘图。
completedata.csv:整理后的月频股票面板数据,是后半段因子构造脚本的直接输入。
研究目标
项目主要完成以下任务:
构造月度股票特征,包括异常换手率、盈利变量、无风险利率、月收益率与市值等。
基于滞后信息构造中国市场风格因子。
输出因子统计结果、相关系数矩阵和累计收益曲线。
比较 2000-01 ~ 2016-12 与扩展样本… See the full description on the dataset page: https://huggingface.co/datasets/transiencee/data_for_China-4-factor-model-reccurence.suitable-childhood-6292c3
suitable-childhood-6292c3
Synthetic products test data: 60 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/sugja07/suitable-childhood-6292c3.texas-licensed-childcare-centers
Texas Licensed Childcare Operations Database — Free Sample (200 rows)
Every licensed childcare operation in Texas — centers and licensed homes — with
contact info, capacity, licensing history, and inspection record. Compiled from
Texas Health and Human Services public licensing records (updated 2026-09-15).
This is a 200-row free sample. The full database has 14,306 operations
statewide — available at https://deepdata.gumroad.com/l/texas-childcare-database. (Quarterly refreshes… See the full description on the dataset page: https://huggingface.co/datasets/deepdataa/texas-licensed-childcare-centers.
