datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
regx-benchmark
RegX
Cross-Domain Multi-View Point Cloud Registration Benchmark
RegX evaluates multi-view point cloud registration across scales spanning nine orders
of magnitude — nanometre-scale microscopy to kilometre-scale airborne maps — and sensors
never designed to be compared: clinical colonoscopes, RGB-D cameras, spinning and
solid-state LiDAR, terrestrial and airborne laser scanners.
Most registration benchmarks fix one sensor and one scale. RegX asks a narrower question
instead: does… See the full description on the dataset page: https://huggingface.co/datasets/YuePanEdward/regx-benchmark.SlimPajama-6B_km-ip-d512so101_cube_tape_task_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 29,
"total_frames": 10345,
"total_tasks": 1,
"total_videos": 58,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:29"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yue1006/so101_cube_tape_task_merged.so101_cube_tape_task_50_newThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 18960,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yue1006/so101_cube_tape_task_50_new.yue_datast_shake_handsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/2yuef/yue_datast_shake_hands.SlimPajama-6B_km_4_8_cos-d512voxbox_cosyvoice2my_dataset
my_dataset
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
so101_cube_tape_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 3633,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yue1006/so101_cube_tape_task.speech2speech_vocalnetyue-math-preference
Cantonese Math Preference
This dataset is a Cantonese and Simplified Chinese translation of argilla/distilabel-math-preference-dpo. For more detailed information about the original dataset, please refer to the provided link.
This dataset is translated by Gemini Pro and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
License
This dataset is provided under the same license as the… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue-math-preference.emilia_lhotse_manifesttech-debt-ai-coding
Debt Behind the AI Boom — Replication Data
Data for the paper:
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo
📄 arXiv:2603.28592 · 💻 Code: github.com/yueyueL/tech-debt-ai-coding
We mined 302.6K AI-authored commits from 6,299 GitHub repositories across five
AI coding assistants (GitHub Copilot, Claude, Cursor, Gemini, Devin), ran static
analysis… See the full description on the dataset page: https://huggingface.co/datasets/yueyuel/tech-debt-ai-coding.douban_movie_info该数据集为豆瓣电影信息维表。
更多信息请参考文章《数据获取:豆瓣电影信息爬取》。
so101_cube_tape_task_30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 29,
"total_frames": 10345,
"total_tasks": 1,
"total_videos": 58,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:29"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yue1006/so101_cube_tape_task_30.FlipGuardData
FlipGuardData
This dataset contains the attack samples presented in the paper FlipAttack: Jailbreak LLMs via Flipping.
FlipAttack is a simple yet effective jailbreak attack against black-box LLMs that exploits their autoregressive nature by disguising harmful prompts using flipping transformations. FlipGuardData contains 45,000 attack samples generated against 8 different LLMs, including GPT-4o, Claude 3.5 Sonnet, and Llama 3.1.
Paper: https://huggingface.co/papers/2410.02832… See the full description on the dataset page: https://huggingface.co/datasets/yueliu1999/FlipGuardData.EmoSupportBench
EmoSupportBench
EmoSupportBench is a comprehensive dataset and benchmark for evaluating emotional support capabilities of large language models (LLMs). It provides a systematic framework to assess how well AI systems can provide empathetic, helpful, and psychologically-grounded support to users seeking emotional assistance.
🎯 Key Features
200-question bilingual evaluation set (English & Chinese) covering 8 major emotional support scenarios
Hierarchical scenario… See the full description on the dataset page: https://huggingface.co/datasets/YueyangWang/EmoSupportBench.nutribench-subset-cs-ja-yue
NutriBench Subset — Code-Switching JA/YUE Extension
Dataset Description
This dataset is a 1,000-sample subset of NutriBench v2,
extended with four additional meal-description columns covering Japanese and Cantonese
monolingual and code-switched variants.
It was constructed to study Research Question 2 (RQ2) of a UCL COMP0087 coursework project.
A sibling dataset covering RQ1 (six-language monolingual evaluation) is available at… See the full description on the dataset page: https://huggingface.co/datasets/chubao/nutribench-subset-cs-ja-yue.prosocial-dialog-yue_Hantso101_cube_tape_task_50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 6722,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yue1006/so101_cube_tape_task_50.cube_tape_task_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 29,
"total_frames": 10345,
"total_tasks": 1,
"total_videos": 58,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:29"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yue1006/cube_tape_task_100.lhotse_issue_1478NOAA_narratives_NERpittsburgh_floods_street_levelThe dataset provides fine-grained spatiotemporal information on urban floods occurring inside the city of Pittsburgh, PA, USA, from 2015 to 2024 by integrating publicly available data sources. The data sources include NOAA storm events database and Pittsburgh 311 flooding requests. Each row corresponds to one segment flood event, characterized by the street segment defined by a distinct combination of "u_node", "v_node", "length_m", and time. Each flood event was mapped to the street segments… See the full description on the dataset page: https://huggingface.co/datasets/yueq92/pittsburgh_floods_street_level.yue
yue cleaned speech corpus
This refresh provides a converged full train manifest, a frozen eval manifest,
and a token-coverage subset of approximately 5000 hours. All three manifests
reuse the existing validated FLAC-in-Parquet audio pool; no audio was copied,
repacked, or transcoded for this refresh.
Corpus overview
Variant
Rows
Hours
train
4,615,779
24933.792
coverage_5000h
927,949
5000.000
eval
1,187
2.000
All audio locators are portable and… See the full description on the dataset page: https://huggingface.co/datasets/WTForbes/yue.cable_insertion_absTCP
cable_insertion_absTCP
Flexiv flexiv_assembly converted to LeRobot v3.0.
State
observation.state is 8D:
TCP X
TCP Y
TCP Z
TCP RX
TCP RY
TCP RZ
gripper
padding = 0
The first six dimensions are absolute TCP pose, the seventh is gripper width-derived state, and the eighth is always zero padding.
Cameras
The dataset contains:
cam_0
cam_1
Image crop
IMAGE_CROP = {
"cam0": lambda img: img[img.shape[0] // 3 :, img.shape[1] // 4 :]… See the full description on the dataset page: https://huggingface.co/datasets/YueJ001/cable_insertion_absTCP.MPRA_VarCREnyc_floods_street_levelThe dataset provides fine-grained spatiotemporal information on urban floods occurring inside New York City, USA, from 2020 to 2024 by integrating publicly available data sources. The data sources include NOAA storm events database, NYC 311 flood-related requests, FloodNet sensor data, and MyCoast NY flood watch records.
Each row corresponds to one spatiotemporal segment flood event, characterized by the street segment defined by a distinct combination of "u_node", "v_node", "length_m", and… See the full description on the dataset page: https://huggingface.co/datasets/yueq92/nyc_floods_street_level.cable_insertion_relTCP
cable_insertion_relTCP
Flexiv flexiv_assembly converted to LeRobot v3.0.
State
observation.state is the raw ConRFT latest state, taken directly from latest_state without converting it to absTCP.
It is a 19D vector:
gripper pose
tcp force x
tcp force y
tcp force z
relative tcp x
relative tcp y
relative tcp z
relative tcp rx
relative tcp ry
relative tcp rz
tcp torque x
tcp torque y
tcp torque z
tcp vel x
tcp vel y
tcp vel z
tcp vel rx
tcp vel ry
tcp vel rz
This is… See the full description on the dataset page: https://huggingface.co/datasets/YueJ001/cable_insertion_relTCP.emilia_cosyvoice_v2_token
