datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
boxrr-23
BOXRR-23: Berkeley Open Extended Reality Recording Dataset 2023
This is a copy of the official Berkeley Open Extended Reality Recording Dataset 2023 (BOXRR-23). Please visit the project website for more information.
In users/ you find one tarball for each user (which you can untar with tar xvf <path/to/user.tar>), which includes all replays of that user. Each replay is stored in a dedicated file in the XROR format.
Metadata
The entire dataset is around 5 TB large… See the full description on the dataset page: https://huggingface.co/datasets/cschell/boxrr-23.cs_csfd-movie-reviews
Dataset Card for CSFD movie reviews (Czech)
Dataset Description
The dataset contains user reviews from Czech/Slovak movie databse website https://csfd.cz.
Each review contains text, rating, date, and basic information about the movie (or TV series).
The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced - each rating has approximately the same frequency.
Dataset Features
Each sample contains:
review_id: unique string identifier… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_csfd-movie-reviews.glove.6B.100d.txtglove.6B.100d.txt for practice
DianJin-CSC-Data
Qwen DianJin Platform |
Github |
ModelScope |
Paper
📢 Introduction
Effective customer support requires not only accurate problem-solving but also structured and empathetic communication aligned with professional standards. However, existing dialogue datasets often lack strategic guidance, and realworld service data is difficult to access and annotate. To address this, we introduce the task of Customer Support Conversation (CSC)… See the full description on the dataset page: https://huggingface.co/datasets/DianJin/DianJin-CSC-Data.CSC
Dataset Card for CSC
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.
中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。
Original Dataset Summary
test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.cs_czech-named-entity-corpus_2.0
Dataset Card for Czech Named Entity Corpus 2.0
Dataset Description
The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test)
Dataset Features
Each sample contains:
text: source sentence
entities: list of selected entities. Each entity contains:
category_id: string identifier of the entity category
category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.csc
Dataset for CSC
中文纠错数据集
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
共计 120w 条数据,以下是数据来源
数据集
语料
链接
SIGHAN+Wang271K 拼写纠错数据集
SIGHAN+Wang271K(27万条)
https://huggingface.co/datasets/shibing624/CSC
ECSpell 拼写纠错数据集
包含法律、医疗、金融等领域
https://github.com/Aopolin-Lv/ECSpell
CGED 语法纠错数据集
仅包含了2016和2021年的数据集… See the full description on the dataset page: https://huggingface.co/datasets/Weaxs/csc.cscommcsc_eval_public
csc_eval_public
一、测评数据说明
1.1 测评数据来源
1.gen_de3.json(5545): '的地得'纠错, 由人民日报/学习强国/chinese-poetry等高质量数据人工生成;
2.lemon_v2.tet.json(1053): relm论文提出的数据, 多领域拼写纠错数据集(7个领域), ; 包括game(GAM), encyclopedia (ENC), contract (COT), medical care(MEC), car (CAR), novel (NOV), and news (NEW)等领域;
3.acc_rmrb.tet.json(4636): 来自NER-199801(人民日报高质量语料);
4.acc_xxqg.tet.json(5000): 来自学习强国网站的高质量语料;
5.gen_passage.tet.json(10000): 源数据为qwen生成的好词好句, 由几乎所有的开源数据汇总的混淆词典生成;… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_eval_public.csc_dataCSC数据:W271K:279,816 条,Medical:39,303 条,Lemon:22,259 条,ECSpell:6,688 条,CSCD:35,001 条。完整项目代码:https://github.com/TW-NLP/ChineseErrorCorrector
csc_clean_wang271k
csc_eval_public
一、测评数据说明
1.1 数据清洗
余-馀: 替换为馀-余
other - 馀: 替换为余
覆-复: 替换为复-覆
other-覆: # 答疆/回覆/反覆
# 覆审
他-她:不纠
她-他:不纠
人名不纠: 识别人名并丢弃
的得地: 建议丢弃(标注得不准)
# # 的 - 地
# # 的 - 得
# # 它 - 他
# # 哪 - 那
# # 改-大小改: 余-馀 覆-复 借-藉 功-工 琅-瑯 震-振 百-白 也-叶 经-禁(经不起-禁不起)
# # 部分不变(人名): 小-晓 一-逸 佳-家 得-地(马哈得) 红-虹 民-明
# # 匹配上但是不改的: 惟-唯 象-像 查-察 立-利 止-只 建-健 他-它 地-的 定-订 带-戴 力-利 成-城 点-店
# # 匹配上但是不改的: 作-做 得-的 场-厂 身-生 有-由 种-重 理-里
# # 空白没匹配上: 今-在 年-今 前-目 当-在 目-在 者-是
# # 外国人名等:其-齐 课-科 博-波… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_clean_wang271k.test-repopara_crawl_cscscsc-wireless-latency-synthetic-100k
CSC Wireless Latency Synthetic Dataset (100k)
This synthetic dataset provides 100,000 prompt-completion pairs designed for training and evaluating PHY/MAC cross-layer optimization models in hybrid Li-Fi/RF wireless networks.
Official Core Implementation & Runtime
To parse, simulate, or process this dataset according to the official protocol specifications, please utilize the official runtime library:
Core Protocol Library (npm):… See the full description on the dataset page: https://huggingface.co/datasets/csc-architecture/csc-wireless-latency-synthetic-100k.DianJin-CSC-Data
Qwen DianJin Platform |
Github |
ModelScope |
Paper
📢 Introduction
Effective customer support requires not only accurate problem-solving but also structured and empathetic communication aligned with professional standards. However, existing dialogue datasets often lack strategic guidance, and realworld service data is difficult to access and annotate. To address this, we introduce the task of Customer Support Conversation (CSC)… See the full description on the dataset page: https://huggingface.co/datasets/navilable/DianJin-CSC-Data.CSC-gpt4
Dataset Card for Chinese Spelling Correction(gpt4 fixed version)
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC-gpt4.csc-decision-intelligence-dataset
CSC Decision Intelligence Dataset
Deterministic decision intelligence seeds and cryptographic verification samples for multi-dimensional evaluation protocols.
Dataset Description
This dataset provides deterministic baseline seeds used by the CSC Protocol (@csc-protocol/core) to evaluate institutional, corporate, and healthcare entities under autonomous AI governance rules.
Supported Domains
Healthcare (medical): Facility operational efficiency… See the full description on the dataset page: https://huggingface.co/datasets/csc-architecture/csc-decision-intelligence-dataset.cs_czech-court-decisions-ner
Dataset Card for Czech Court Decisions NER
Dataset Description
Czech Court Decisions NER is a dataset of 300 court decisions published by The Supreme Court of the Czech Republic and the Constitutional Court of the Czech Republic.
In the documents, 4 types of named entities are selected.
Dataset Features
Each sample contains:
filename: file name in the original dataset
text: court decision document in plain text
entities: list of selected entities. Each entity… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-court-decisions-ner.CSCAN-CodeQL
CSCAN-CodeQL
CodeQL vulnerability detection
Attribution
Author: Euisuh JeongAffiliation: Qatar Computing Research Institute (QCRI), Hamad Bin Khalifa UniversityLicense: MIT
Citation
@dataset{cscan_codeql,
author={Jeong, Euisuh},
year={2026},
title={CSCAN-CodeQL},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/datasets/euisuh/CSCAN-CodeQL}}
}
kate-pd
Welcome to the KATE-PD Dataset
Welcome to the home page of Kahramanmaraş Türkiye Earthquake-Post Disaster Dataset (KATE-PD). If you are reading this README, you are probably visiting one of the following places to learn more about KATE-PD Dataset and the associated study "A Rapir Damage Assessment using Remote Sensing: Türkiye 2023 Post-Earthquake Dataset: KATE-PD", to be presented in the IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2025).
Code Ocean Capsule… See the full description on the dataset page: https://huggingface.co/datasets/CSCRS/kate-pd.axia-csc-corpus
Axia — Chandra Source Catalog corpus
astromindinc/axia-csc-corpus — 51,450 X-ray sources from the
Chandra Source Catalog 2.1, each carrying:
Per-photon event lists in two forms: an event_list pruned to a single
8 h window in 0.5-8 keV (the input shape the Axia fine-tuned model trained
on), and an original_event_list containing the full unpruned observation
(the input to the model-free spectrum-snapshot / light-curve pipeline).
A 64-d learned embedding (pca_64d) suitable for… See the full description on the dataset page: https://huggingface.co/datasets/astromindinc/axia-csc-corpus.cs-combined-002
Dataset Card for "cs-combined-002"
More Information needed
CSC
Dataset Card for CSC
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.
中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。
Original Dataset Summary
test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/henryy1990/CSC.csc_public_de3
csc_public_de3数据集
数据来源
1.由人民日报/学习强国/chinese-poetry等高质量数据人工生成;
2.来自人民日报高质量语料;
3.来自学习强国网站的高质量语料;
4.源数据为qwen生成的好词好句;
5.古诗词chinese-poetry; 文言文garychowcmu/daizhigev20;
数据简介
该数据主要为'的地得'纠错;
其中训练数据130753条, 验证数据5545条, 测试数据5545条;
句子平均长度为36, 最长句子长度为414, 最短为5, 95%的为89, 75%的为46, 60%的为34;
每个句子中字的平均错误数为2;
数据详情
################################################################################################################################
train.json
130753… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_public_de3.cscl_parse1This is cscl datasets.
kl3m-data-dotgov-www.csce.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.csce.gov.sonic-seasoning
Sonic Seasoning
Sonic Seasoning is a multi-source corpus of short music/sound clips, each perceptually rated for how strongly it evokes the five basic tastes — sweet, bitter, salty, sour, spicy — with additional temperature and emotion annotations on a subset. Every rating is normalized to [0, 1], and every row points to a single uniformly-encoded .wav file. It is the benchmark corpus for taste-from-audio prediction as a content-based music-information-retrieval task.
📄… See the full description on the dataset page: https://huggingface.co/datasets/csc-unipd/sonic-seasoning.0320_cosmetic2_dscs_corpora_parliament_processedkl3m-filter-data-dotgov-www.csce.gov
