datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
copyrightgpt-v1
CopyrightGPT
A fast first-draft video dataset + fingerprint store, built to answer one question:
has this video (or something very similar to it) been seen before?
Each ingested video gets a YouTube-style random id and a folder ("bin") of binary
data files. A perceptual hash (pHash) is computed for sampled frames, so a new
video can be checked against everything already stored to flag likely duplicates /
re-uploads.
This is an early, intentionally simple draft — perceptual-hash… See the full description on the dataset page: https://huggingface.co/datasets/hdcli/copyrightgpt-v1.hd_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 37801,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nanyong/hd_cube.hdc_dataset_h1Zarr format datasets for Humanoid Diffusion Controller training , each containing 200 k trajectories collected on the H1 humanoid robot.:
h1_3_200k.zarr – 3-point motion-tracking trajectories
h1_22_200k.zarr – 22-point motion-tracking trajectories
hd_cube_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 177,
"total_frames": 82401,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:177"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nanyong/hd_cube_data.hd_cube_data2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 261,
"total_frames": 164949,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:261"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nanyong/hd_cube_data2.hd_cube_data3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 86,
"total_frames": 34163,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:86"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nanyong/hd_cube_data3.HD-car-img-captionhd_cube2hd_cylinderHdchdc-context-sentenceshd_cu_3_captwitter-Hdclip08-2025.12.27-2004829967679701410-TNJRJOogCJb-uer7-part1Hdc2HDCBHDCB Dataset
HDCBHDCB Dataset
HDC-Checkpointshdcz1Hdc3
