datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rdpRDD2022
RDD2022: Multi-National Road Damage Detection Dataset (4-Class YOLO Export)
Unofficial redistribution of the RDD2022 multi-national road-damage dataset, reduced to the 4-class CRDDC2022 taxonomy and reformatted into a standardized YOLO-compatible directory layout, under the original CC BY-SA 4.0 license.
Disclaimer
This repository is not an official release of the RDD2022 dataset.
RDD2022 was created by Deeksha Arya, Hiroya Maeda, Sanjay Kumar Ghosh… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/RDD2022.rdmap_traj_0rdt-ft-data
Dataset Card
This is the fine-tuning dataset used in the paper RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.
Source
Project Page: https://rdt-robotics.github.io/rdt-robotics/
Paper: https://arxiv.org/pdf/2410.07864
Code: https://github.com/thu-ml/RoboticsDiffusionTransformer
Model: https://huggingface.co/robotics-diffusion-transformer/rdt-1b
Uses
Download all archive files and use the following command to extract:
cat rdt_data.tar.gz.* |… See the full description on the dataset page: https://huggingface.co/datasets/robotics-diffusion-transformer/rdt-ft-data.rdpdegentic_rd0
Dataset Card for Degentic Games
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
swe-marathon
SWE Marathon: Ultra Long-Horizon Software Engineering Tasks
20 ultra long-horizon software-engineering tasks designed to challenge frontier coding agents. Each task ships with a containerized environment, a precise instruction, comprehensive tests, and a reference oracle solution. All tasks pass NOP-baseline / Oracle-fix validation.
Homepage: https://github.com/abundant-ai/swe-marathon
License: Apache 2.0
Format: Harbor task format (task.toml + instruction.md + environment/ +… See the full description on the dataset page: https://huggingface.co/datasets/rdesai2/swe-marathon.tintedglass-rdt-synth
TintedGlass RDT — 贴膜玻璃合成反射数据 + 实拍配对测试集
面向贴膜玻璃(车窗/建筑膜)场景的单图反射去除合成数据。核心构造:
I = α(x)·T + α(x)^p·R + n,其中 α(x) ∈ (0,1] 是空间变化的膜透过率场
(覆盖突变 + 弧状缓变),p=0 —— 反射强度不随膜衰减,这是贴膜场景与通用
SIRR 合成集(一般 p=1)的关键差别。
仓库结构
train/rdt_0915/ 合成训练三元组(4000 场景 × 5 档 = 20000 样本, 448²)
T/{sid}.png 透射层(各档共享)
R/{sid}.png 反射层(各档共享)
I/{sid}_l{030,045,060,075,100}.png 合成输入(5 档)
alpha/{sid}_l{...}.npy 透过率场(float16, 448×448)
m/{sid}.npy… See the full description on the dataset page: https://huggingface.co/datasets/zmy1234567890/tintedglass-rdt-synth.CIC-IDS2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily.
NIDS Datasets
The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/CIC-IDS2017.UNSW-NB15We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily.
NIDS Datasets
The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/UNSW-NB15.rdc-manipulation-datasetsprocedural-engine-sounds
Procedural Engine Sounds Dataset (Official)
⚠️ Canonical Source
This is the original and official release of the Procedural Engine Sounds Dataset, created by Robin Doerfler:
Project page (start here): https://rdoerfler.github.io/procedural-engine-sounds-page/
Hugging Face repository: https://huggingface.co/datasets/rdoerfler/procedural-engine-sounds
Zenodo DOI (primary citation): https://doi.org/10.5281/zenodo.16883336
Other versions of this dataset on… See the full description on the dataset page: https://huggingface.co/datasets/rdoerfler/procedural-engine-sounds.Rdiffusion-audio
Audio dump for Rdiffusion dataset
libero_90_lerobot_pathmask_rdpThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jesbu1/libero_90_lerobot_pathmask_rdp.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.inatFGVC51m2r-real-tapeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 56,
"total_frames": 23054,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:56"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-real-tape.sim-handover-corrective-1000-0820This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda-panda",
"total_episodes": 2000,
"total_frames": 303710,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/sim-handover-corrective-1000-0820.ds003604-session-rdms
ds003604 session RDMs
Brain representational dissimilarity matrices for the ds003604 auditory language fMRI
dataset, one per (task × session), so the fMRI preprocessing never has to be repeated.
12 of 12 cells present: Sem, Phon, Gram, Plaus × ses-5, ses-7, ses-9.
Each <Task>/session_rdm_<ses>.npz contains:
key
contents
rdm
the RDM, (n_stim, n_stim), correlation distance. 72×72 for Sem/Phon, 60×60 for Gram/Plaus (perceptual-control trials excluded)
stimuli
stimulus… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/ds003604-session-rdms.rdkit_featuresRDD2022mb-atmospheric_dust_cls_rdr
mb-atmospheric_dust_cls_rdr_upd
A Mars image classification dataset for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-22
Cite As: TBD
Classes
This dataset contains the following classes:
0: dusty
1: not_dusty
Statistics
train: 9817 images
test: 5214 images
val: 4969 images
few_shot_train_2_shot: 4 images
few_shot_train_1_shot: 2 images… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-atmospheric_dust_cls_rdr.bhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/rdmpage/bhl-impact-gt.multi-robot-basketThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 41,
"total_frames": 25531,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:41"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/multi-robot-basket.joint-handover3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 45,
"total_frames": 44370,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/joint-handover3.1m2r-tape2-bvThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 56,
"total_frames": 49290,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:56"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-tape2-bv.1m2r-handover2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova-arx",
"total_episodes": 44,
"total_frames": 63492,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:44"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-handover2.1m2r-handover2-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "kinova-arx",
"total_episodes": 44,
"total_frames": 63492,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:44"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-handover2-v30.rag-logskinova-pick-placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova",
"total_episodes": 50,
"total_frames": 8396,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/kinova-pick-place.
