datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.hdr-demo-clips
HDR Demo Clips (Lightricks SDR→HDR)
Paired SDR (input) / HDR (output) frame sequences from the Lightricks SDR-to-HDR pipeline (IC-LoRA on LTX-2).
Each clip contains:
hdr_exr/frame_XXXXX.exr — HDR output (f16, linear Rec.709/sRGB primaries, scene-referred)
sdr_png/frame_XXXXX.png — SDR input (8-bit sRGB, display-referred)
thumbnail.jpg — 280px preview from the middle frame
Dimensions: HDR is symmetrically cropped from SDR to match model-friendly dimensions (typically 28–56px… See the full description on the dataset page: https://huggingface.co/datasets/oumoumad/hdr-demo-clips.cobotmagic_Sim_click_bellCSL-Clinic
Dataset Card for CSL-Clinic
CSL-Clinic is a gloss-annotated Chinese Sign Language dataset for sign language understanding in the clinical and healthcare domain. It was introduced in Variational Sign Language Translation, published in the International Journal of Computer Vision.
Public Release and Full-Dataset Access
This Hugging Face repository publicly distributes only the 500-example test split. The complete dataset contains 5,972 examples; its train and dev… See the full description on the dataset page: https://huggingface.co/datasets/rzhao/CSL-Clinic.nextturn-study-clipsmy-clipspragya-h3-clipsCLIP-CC
📚 CLIP-CC Dataset (Movie Clips Edition)
Paper | arXiv | Project Page | Benchmark Code (CLIP-CC-Bench) | Dataset Repo (CLIP-CC)
CLIP-CC is a curated dataset for long-form video description: 200 movie clips sourced from YouTube, each about 90 seconds long (~5 hours in total) and drawn from more than 140 films spanning 1959–2024, each paired with one human-written English reference description averaging 402 ± 208 words. The references were written by four graduate-student… See the full description on the dataset page: https://huggingface.co/datasets/MINT-SDSU/CLIP-CC.nvidia-av-trajectory-clips-256ClickTargetPreprocessThreeCamerasSetUpOneThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 79,
"total_frames": 28063,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:79"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessThreeCamerasSetUpOne.ClickTargetPreprocessThreeCamerasSetUpOneMultiplePiecesAIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_remote_control_clip_box_water_bottle.ClickTargetPreprocessCleanThreeCamerasThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 80,
"total_frames": 80714,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessCleanThreeCameras.how2sign-asl-clips
how2sign-asl-clips
Sentence-level clips from the How2Sign ASL
dataset, cut to the realigned timestamps in how2sign_realigned_train.csv.
Built for the EPFL CS-503 hand2string project.
This is a work-in-progress mirror of a single shard (351 source
videos, 4991 clips). More splits and shards will be added as we
download them.
Schema (metadata.parquet)
column
type
notes
sentence_id
string
primary key, e.g. --7E2sU6zP4_10
sentence_name
string
full How2Sign… See the full description on the dataset page: https://huggingface.co/datasets/martinctl/how2sign-asl-clips.cobotmagic_Real_click_bell
CobotMagic Real Dataset: click_bell
Generated by convert_piper_hdf5_to_lerobot.py.
RoboSynChallenge Alignment
Task dir: click_bell
Environment ID: ClickBell
Real dataset: RoboSynChallenge/cobotmagic_Real_click_bell
Aligned simulation dataset: click_bell/cobotmagic_Sim_click_the_bell_000
Prompt: Click the bell
Summary
Robot type: aloha
FPS: 10
Sampling mode: full
Leading group: slave_cam_high
Crop strategy: LATEST_START
Default leading tolerance:… See the full description on the dataset page: https://huggingface.co/datasets/RoboSynChallenge/cobotmagic_Real_click_bell.how2sign-front-clips
How2Sign RGB Front Clips (VideoFolder)
This directory packages the frontal-view, sentence-level How2Sign RGB clips in
Hugging Face VideoFolder format. It contains the official train, validation,
and test splits with English sentence annotations.
Directory layout
how2sign_videofolder/
├── README.md
├── train/
│ ├── metadata.parquet
│ └── shard_001_032/ ... shard_032_032/
├── validation/
│ ├── metadata.parquet
│ └── shard_001_018/ ... shard_018_018/
└──… See the full description on the dataset page: https://huggingface.co/datasets/WayenVan/how2sign-front-clips.ClickTargetPreprocessThreeCamerasSetUpOneRedTriangleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 32888,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessThreeCamerasSetUpOneRedTriangle.nvidia-av-trajectory-clips-384VidPair-Halluc-ClipClickTargetPreprocessTestThreeCameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 120,
"total_frames": 37577,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:120"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessTestThreeCamera.ClickTargetPreprocessTestTwoCameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 80,
"total_frames": 19585,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/ClickTargetPreprocessTestTwoCamera.trossen_place_bead_on_string_10_gr00t_clip_01eval_act_ClickTargetPreprocessThreeTwoCameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 19,
"total_frames": 7271,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:19"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/eval_act_ClickTargetPreprocessThreeTwoCamera.openp2p-action-clips-media-3000-20260911agilex_pick_same_color_clipThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 20,
"total_frames": 16189,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:20"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_pick_same_color_clip.eval_act_ClickTargetPreprocessThreeCamerasSetUpOneThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 15,
"total_frames": 5165,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/eval_act_ClickTargetPreprocessThreeCamerasSetUpOne.eval_act_ClickTargetPreprocessThreeTwoCamera_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 15,
"total_frames": 5546,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Greynar/eval_act_ClickTargetPreprocessThreeTwoCamera_2.so101_teleop_vials_clean_completion_clips_v6
SO101 vials — completion clips V6
Derived from ltchang/so101_teleop_vials_clean. Source files 001, 002, and 020 are omitted. The remaining 56 full episodes are retained, and 53 synchronized completion clips are appended.
Per-episode video layout
Each episode has one MP4 per camera under videos/<camera>/chunk-000/file-XXX.mp4, matching the source layout. Full episodes retain their original source IDs (004–060, with omissions); appended clips are files 061–113.… See the full description on the dataset page: https://huggingface.co/datasets/ltchang/so101_teleop_vials_clean_completion_clips_v6.groot_cable_clip_v2_h264pragyadex-clips
