datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
American-Sign-Language-Dataset
American Sign Language (ASL) Dataset
Description:This dataset contains 108,618 videos representing 2,208 ASL words, with each word having a minimum of 30 videos. The videos were scraped, collected from multiple sources, and preprocessed to ensure consistency, quality, and usability for machine learning and gesture recognition tasks. Each video is ≤10 MB, optimized for storage and model training.The dataset can be used for ASL gesture recognition, video-based ML tasks, and model… See the full description on the dataset page: https://huggingface.co/datasets/shpouladi/American-Sign-Language-Dataset.American-Sign-Language-Dataset
American Sign Language (ASL) Dataset
Description:This dataset contains 108,618 videos representing 2,208 ASL words, with each word having a minimum of 30 videos. The videos were scraped, collected from multiple sources, and preprocessed to ensure consistency, quality, and usability for machine learning and gesture recognition tasks. Each video is ≤10 MB, optimized for storage and model training.The dataset can be used for ASL gesture recognition, video-based ML tasks, and model… See the full description on the dataset page: https://huggingface.co/datasets/akasheroor/American-Sign-Language-Dataset.Indian_Sign_Language_Data.gov_Rencoded
Indian_Sign_Language_Data.gov_Rencoded
Dataset Overview
Dataset name: Indian_Sign_Language_Data.gov_RencodedHugging Face repository: silentone0725/Indian_Sign_Language_Data.gov_RencodedModality: Video (H.265 / HEVC)Total size: ~75 GBOriginal size: ~200 GBLanguage: Indian Sign Language (ISL)License: MIT
This dataset is a re-encoded and curated version of the Indian Sign Language Dictionary originally published on the Government of India Open Data Portal (data.gov.in).… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/Indian_Sign_Language_Data.gov_Rencoded.language_table_train_115000_120000_augmented
language_table_train_115000_120000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,828
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_115000_120000_augmented.language_table_lerobot_v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "xarm",
"total_episodes": 442226,
"total_frames": 7045476,
"total_tasks": 127605,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:442226"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/tailong-wu/language_table_lerobot_v30.language_table_train_75000_80000_augmented
language_table_train_75000_80000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,557
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_75000_80000_augmented.language_table_train_55000_60000_augmented
language_table_train_55000_60000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,295
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_55000_60000_augmented.language_table_train_5000_10000_augmented
language_table_train_5000_10000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,151
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_5000_10000_augmented.sign-language-avatar-gloss-dgs
Dataset for German Sign Language Avatar Training
Dataset Summary
This dataset provides curated resources for training data-driven avatars
to perform isolated signs in German Sign Language (Deutsche Gebärdensprache, DGS). It includes videos of individual signs as well as corresponding pose estimation results in a structured and reusable format.
The data is at this moment just sourced from SignDict.org and organized into three primary folders:
videos-raw: Original… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/sign-language-avatar-gloss-dgs.language_table_train_95000_100000_augmented
language_table_train_95000_100000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,860
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_95000_100000_augmented.language_table_train_345000_350000_augmented
language_table_train_345000_350000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,459
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_345000_350000_augmented.language_table_train_65000_70000_augmented
language_table_train_65000_70000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,376
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_65000_70000_augmented.language_table_train_140000_145000_augmented
language_table_train_140000_145000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,121
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_140000_145000_augmented.American-Sign-Language-Dataset
American Sign Language (ASL) Dataset
Description:This dataset contains 108,618 videos representing 2,208 ASL words, with each word having a minimum of 30 videos. The videos were scraped, collected from multiple sources, and preprocessed to ensure consistency, quality, and usability for machine learning and gesture recognition tasks. Each video is ≤10 MB, optimized for storage and model training.The dataset can be used for ASL gesture recognition, video-based ML tasks, and model… See the full description on the dataset page: https://huggingface.co/datasets/LaluDhsjsklsls/American-Sign-Language-Dataset.language_table_train_15000_20000_augmented
language_table_train_15000_20000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,558
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_15000_20000_augmented.language_table_train_130000_135000_augmented
language_table_train_130000_135000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,357
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_130000_135000_augmented.language_table_train_125000_130000_augmented
language_table_train_125000_130000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 78,601
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_125000_130000_augmented.language_table_train_110000_115000_augmented
language_table_train_110000_115000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,139
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_110000_115000_augmented.language_table_train_70000_75000_augmented
language_table_train_70000_75000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,047
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_70000_75000_augmented.language_table_train_390000_395000_augmented
language_table_train_390000_395000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,774
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_390000_395000_augmented.cube_stack_language_pilotultimate-life-visual-language
Ultimate Life Visual Language — Public-Domain Visual Grammar Studies
Public, U.S.-first research/reference dataset. It contains silent, source-derived visual-grammar studies for controlled MiniMax H3 reference-video evaluation. It is not an Ultimate Life identity model, a character dataset, or a permission grant over third-party restorations, music, titles, or branding.
Contents
clips/ — 20 short H.264 MP4 visual-grammar studies with paired captions.… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/ultimate-life-visual-language.language_table_train_215000_220000_augmented
language_table_train_215000_220000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,565
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_215000_220000_augmented.language_table_train_170000_175000_augmented
language_table_train_170000_175000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,318
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_170000_175000_augmented.language_table_train_420000_425000_augmented
language_table_train_420000_425000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,111
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_420000_425000_augmented.s3-language-following-v7big
S3 Language-Following v7big
Dual-arm tabletop pick corpus for the S3 language-following benchmark.
5,531 episodes over 2,520 scenes (textured rendering, 4 cameras: head / front / wrist×2, 256×256 @10fps, 98 frames/episode)
Task family: "pick up the apple 〈relation〉 the 〈landmark〉" — 4 spatial relations × 6 landmark objects (18 populated cells), 3 identical apples per scene, 1 distractor landmark, 8 caption phrasings per cell
Balanced ~62 demos per (cell × distractor-config);… See the full description on the dataset page: https://huggingface.co/datasets/szang18/s3-language-following-v7big.language_table_train_120000_125000_augmented
language_table_train_120000_125000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,140
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_120000_125000_augmented.language_table_train_30000_35000_augmented
language_table_train_30000_35000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,860
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_30000_35000_augmented.language_table_train_10000_15000_augmented
language_table_train_10000_15000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,689
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_10000_15000_augmented.language_table_train_135000_140000_augmented
language_table_train_135000_140000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,556
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_135000_140000_augmented.
