datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xr-motion-dataset-catalogue
XR Motion Dataset Catalogue
Overview
The XR Motion Dataset Catalogue, accompanying our paper "Navigating the Kinematic Maze: A Comprehensive Guide to XR Motion Dataset Standards," standardizes and simplifies access to Extended Reality (XR) motion datasets. The catalogue represents our initiative to streamline the usage of kinematic data in XR research by aligning various datasets to a consistent format and structure.
Dataset Specifications
All datasets in this… See the full description on the dataset page: https://huggingface.co/datasets/cschell/xr-motion-dataset-catalogue.boxrr-23
BOXRR-23: Berkeley Open Extended Reality Recording Dataset 2023
This is a copy of the official Berkeley Open Extended Reality Recording Dataset 2023 (BOXRR-23). Please visit the project website for more information.
In users/ you find one tarball for each user (which you can untar with tar xvf <path/to/user.tar>), which includes all replays of that user. Each replay is stored in a dedicated file in the XROR format.
Metadata
The entire dataset is around 5 TB large… See the full description on the dataset page: https://huggingface.co/datasets/cschell/boxrr-23.csclircs_csfd-movie-reviews
Dataset Card for CSFD movie reviews (Czech)
Dataset Description
The dataset contains user reviews from Czech/Slovak movie databse website https://csfd.cz.
Each review contains text, rating, date, and basic information about the movie (or TV series).
The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced - each rating has approximately the same frequency.
Dataset Features
Each sample contains:
review_id: unique string identifier… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_csfd-movie-reviews.glove.6B.100d.txtglove.6B.100d.txt for practice
who-is-alyx-testDianJin-CSC-Data
Qwen DianJin Platform |
Github |
ModelScope |
Paper
📢 Introduction
Effective customer support requires not only accurate problem-solving but also structured and empathetic communication aligned with professional standards. However, existing dialogue datasets often lack strategic guidance, and realworld service data is difficult to access and annotate. To address this, we introduce the task of Customer Support Conversation (CSC)… See the full description on the dataset page: https://huggingface.co/datasets/DianJin/DianJin-CSC-Data.CSC722_SP26_GROUP1_CHANGE_DETECTIONCSC
Dataset Card for CSC
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.
中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。
Original Dataset Summary
test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.MS-OCT-CSC-ERM-MS-ROVscitylens
How to add files using Git LFS
"This is incomplete, but I was having difficulty adding files using Git LFS or Huggingface Xet, so these are some of the steps I took
You essentially need to track files using Git LFS. You do this by running a command like 'git lfs track file_name'
Make sure you have git-lfs (or git-xet?) installed first
Also if you already set up the repository without file tracking you may need to run 'git lfs migrate import --everything' to rewrite the repo… See the full description on the dataset page: https://huggingface.co/datasets/CUNY-Hunter-CSCI-49900-Group-6/citylens.cs_czech-named-entity-corpus_2.0
Dataset Card for Czech Named Entity Corpus 2.0
Dataset Description
The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test)
Dataset Features
Each sample contains:
text: source sentence
entities: list of selected entities. Each entity contains:
category_id: string identifier of the entity category
category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.csc
Dataset for CSC
中文纠错数据集
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
共计 120w 条数据,以下是数据来源
数据集
语料
链接
SIGHAN+Wang271K 拼写纠错数据集
SIGHAN+Wang271K(27万条)
https://huggingface.co/datasets/shibing624/CSC
ECSpell 拼写纠错数据集
包含法律、医疗、金融等领域
https://github.com/Aopolin-Lv/ECSpell
CGED 语法纠错数据集
仅包含了2016和2021年的数据集… See the full description on the dataset page: https://huggingface.co/datasets/Weaxs/csc.cscommFranka_place_container_plateThis dataset was created using LeRobot.
Dataset Description
Franka Dual-Arm Place Container Plate Dataset
Move the container to the plate.
Hardware
Robot: 2× Franka Emika Panda (7-DOF dual-arm setup)
Cameras: 3× Intel RealSense D435 (256×256 RGB, front + wrist views)
State Space (38 dimensions)
Left Arm (19 dimensions):
tcp_pose (6): End-effector pose [x, y, z, roll, pitch, yaw] in meters and radians
tcp_vel (6): End-effector velocity [vx, vy, vz… See the full description on the dataset page: https://huggingface.co/datasets/CSCSXX/Franka_place_container_plate.pick_place_cube_1.17This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 10716,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CSCSXX/pick_place_cube_1.17.Franka_beat_block_hammerThis dataset was created using LeRobot.
Dataset Description
Franka Dual-Arm Beat Block Hammer Dataset
Pick the hammer and strike the block.
Hardware
Robot: 2× Franka Emika Panda (7-DOF dual-arm setup)
Cameras: 3× Intel RealSense D435 (256×256 RGB, front + wrist views)
State Space (38 dimensions)
Left Arm (19 dimensions):
tcp_pose (6): End-effector pose [x, y, z, roll, pitch, yaw] in meters and radians
tcp_vel (6): End-effector velocity [vx, vy, vz… See the full description on the dataset page: https://huggingface.co/datasets/CSCSXX/Franka_beat_block_hammer.csc_eval_public
csc_eval_public
一、测评数据说明
1.1 测评数据来源
1.gen_de3.json(5545): '的地得'纠错, 由人民日报/学习强国/chinese-poetry等高质量数据人工生成;
2.lemon_v2.tet.json(1053): relm论文提出的数据, 多领域拼写纠错数据集(7个领域), ; 包括game(GAM), encyclopedia (ENC), contract (COT), medical care(MEC), car (CAR), novel (NOV), and news (NEW)等领域;
3.acc_rmrb.tet.json(4636): 来自NER-199801(人民日报高质量语料);
4.acc_xxqg.tet.json(5000): 来自学习强国网站的高质量语料;
5.gen_passage.tet.json(10000): 源数据为qwen生成的好词好句, 由几乎所有的开源数据汇总的混淆词典生成;… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_eval_public.csc_dataCSC数据:W271K:279,816 条,Medical:39,303 条,Lemon:22,259 条,ECSpell:6,688 条,CSCD:35,001 条。完整项目代码:https://github.com/TW-NLP/ChineseErrorCorrector
Franka_adjust_bottleThis dataset was created using LeRobot.
Dataset Description
Franka Dual-Arm Adjust Bottle Dataset
Lift the bottle head-up from the table.
Hardware
Robot: 2× Franka Emika Panda (7-DOF dual-arm setup)
Cameras: 3× Intel RealSense D435 (256×256 RGB, front + wrist views)
State Space (66 dimensions)
Left Arm (33 dimensions):
tcp_pose (6): End-effector pose [x, y, z, roll, pitch, yaw] in meters and radians
tcp_vel (6): End-effector velocity [vx, vy, vz… See the full description on the dataset page: https://huggingface.co/datasets/CSCSXX/Franka_adjust_bottle.Franka_place_container_plate_30_trajs_260321This dataset was created using LeRobot.
Dataset Description
Franka Single-Arm Place Container Plate Dataset
Move the container to the plate.
Hardware
Robot: 1× Franka Emika Panda (7-DOF each)
Cameras: Intel RealSense D435 RGB views (front + wrist)
State Space (33 dimensions)
Left Arm (33 dimensions):
tcp_force (3): [tcp_force_x, tcp_force_y, tcp_force_z]
gripper_pose (1): [gripper_pose]
joint_pos (7): [q0, q1, q2, q3, q4, q5, q6]
joint_vel (7):… See the full description on the dataset page: https://huggingface.co/datasets/CSCSXX/Franka_place_container_plate_30_trajs_260321.test-repopara_crawl_cscscsc_clean_wang271k
csc_eval_public
一、测评数据说明
1.1 数据清洗
余-馀: 替换为馀-余
other - 馀: 替换为余
覆-复: 替换为复-覆
other-覆: # 答疆/回覆/反覆
# 覆审
他-她:不纠
她-他:不纠
人名不纠: 识别人名并丢弃
的得地: 建议丢弃(标注得不准)
# # 的 - 地
# # 的 - 得
# # 它 - 他
# # 哪 - 那
# # 改-大小改: 余-馀 覆-复 借-藉 功-工 琅-瑯 震-振 百-白 也-叶 经-禁(经不起-禁不起)
# # 部分不变(人名): 小-晓 一-逸 佳-家 得-地(马哈得) 红-虹 民-明
# # 匹配上但是不改的: 惟-唯 象-像 查-察 立-利 止-只 建-健 他-它 地-的 定-订 带-戴 力-利 成-城 点-店
# # 匹配上但是不改的: 作-做 得-的 场-厂 身-生 有-由 种-重 理-里
# # 空白没匹配上: 今-在 年-今 前-目 当-在 目-在 者-是
# # 外国人名等:其-齐 课-科 博-波… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_clean_wang271k.kate-cd
Welcome to the KATE-CD Dataset
Welcome to the home page of Kahramanmaraş Türkiye Earthquake-Change Detection Dataset (KATE-CD). If you are reading this README, you are probably visiting one of the following places to learn more about KATE-CD Dataset and the associated study "Earthquake Damage Assessment with SAMCD: A Change Detection Approach for VHR Images", to be appear in Journal of Applied Remote Sensing.
Code Ocean Capsule in Open Science Library
GitHub Repository… See the full description on the dataset page: https://huggingface.co/datasets/CSCRS/kate-cd.csc-wireless-latency-synthetic-100k
CSC Wireless Latency Synthetic Dataset (100k)
This synthetic dataset provides 100,000 prompt-completion pairs designed for training and evaluating PHY/MAC cross-layer optimization models in hybrid Li-Fi/RF wireless networks.
Official Core Implementation & Runtime
To parse, simulate, or process this dataset according to the official protocol specifications, please utilize the official runtime library:
Core Protocol Library (npm):… See the full description on the dataset page: https://huggingface.co/datasets/csc-architecture/csc-wireless-latency-synthetic-100k.DianJin-CSC-Data
Qwen DianJin Platform |
Github |
ModelScope |
Paper
📢 Introduction
Effective customer support requires not only accurate problem-solving but also structured and empathetic communication aligned with professional standards. However, existing dialogue datasets often lack strategic guidance, and realworld service data is difficult to access and annotate. To address this, we introduce the task of Customer Support Conversation (CSC)… See the full description on the dataset page: https://huggingface.co/datasets/navilable/DianJin-CSC-Data.pick_place_cube_1.18This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 16491,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CSCSXX/pick_place_cube_1.18.CSC-gpt4
Dataset Card for Chinese Spelling Correction(gpt4 fixed version)
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC-gpt4.csc245-project1-audio
