CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RekaAI /RekaDaily-10k-raw RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.imagevideo-classification100K<n<1M22 likes226k downloads9d agoHugging Face02mvp-lab /LLaVA-OneVision-2-Data LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. At a Glance The dataset is split across two Hugging Face repositories because of its size: Repository What it contains Part 1 (this repository) ~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.imagevideo-text-to-textn<1K39 likes131k downloads23d agoHugging Face03simple-world-lab /HiFi-UMI-2K HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data 2,000 hours released · 6 synchronized camera views · 480+ scenes · 3 mm pose accuracy · <40 µs synchronization 🌐 Project Website | 📦 Dataset | 📄 Paper: arXiv:2607.25895 Examples from the HiFi-UMI corpus. Click the image to play the video. 📚 Introduction HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.… See the full description on the dataset page: https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K.tabularrobotics100M<n<1B55 likes123k downloads2mo agoHugging Face04yaak-ai /L2DTL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot goes to driving school 90+ TeraBytes of multimodal data (5000+ hours of driving) from 30 cities in Germany 6x surrounding HD cameras and complete vehicle state: Speed/Heading/GPS/IMU Continuous: Gas/Brake/Steering and discrete actions: Gear/Turn Signals Environment state: Lane count, Road type (highway|residential), Road surface (asphalt, cobbled, sett), Max speed limit.… See the full description on the dataset page: https://huggingface.co/datasets/yaak-ai/L2D.tabularrobotics10M<n<100M52 likes72k downloads4mo agoHugging Face05cadene /agibot_alpha_v30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "AgiBot_A2D", "total_episodes": 28122, "total_frames": 47613574, "total_tasks": 30, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:28122"}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/agibot_alpha_v30.tabularrobotics10M<n<100M1 likes72k downloads1y agoHugging Face06lmms-eval /Video-MMEtext1K<n<10K96 likes40k downloads2y agoHugging Face07HRDexDB /HRDexDB HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments Authors Jongbin Lim¹⋆, Taeyun Ha¹⋆, Seongho Cha, Kanghyun Cho, Mingi Choi¹, Subin Jeon¹, Jisoo Kim¹, Byungjun Kim¹, Hanbyul Joo¹²† ¹ Seoul National University² RLWRLD ⋆ Equal contribution† Corresponding author News (2026.09.20) The full set of Robotiq 2F-85 data has been uploaded! (2026.07.27) We are improving the quality of the object mesh and the tracking results.… See the full description on the dataset page: https://huggingface.co/datasets/HRDexDB/HRDexDB.3d10K<n<100K16 likes39k downloads2d agoHugging Face08zekaiwang /trex_dataset T-Rex Dataset A large-scale, tactile-reactive bimanual manipulation dataset, collected via teleoperation on a Dexmate Vega-1 robot with two Sharpa Wave dexterous hands. Stored as a LeRobotDataset v3.0. 🌐 Project Page · ✍️ Paper (arXiv) · 💻 Code (T-Rex) · 🚀 Dataset Quickstart · 📓 Colab notebook One episode from each of 20 motor primitives (head-camera view, cropped to the workspace), each with a different object. Teleoperation setup: Manus gloves + VIVE… See the full description on the dataset page: https://huggingface.co/datasets/zekaiwang/trex_dataset.tabularrobotics1M<n<10M30 likes30k downloads3mo agoHugging Face09configinc /HABIT HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation TL;DR: HABIT is a large-scale robot demonstration dataset for human-present environments, designed to teach robot policies human-aware behaviors. Keywords: Robot Manipulation Dataset, Human-Robot Interaction, Vision-Language-Action Model Overview Data-driven approaches have emerged as a promising direction for training robotic manipulation policies. Recent robot datasets have… See the full description on the dataset page: https://huggingface.co/datasets/configinc/HABIT.tabularrobotics1M<n<10M6 likes30k downloads3mo agoHugging Face10brandonyang /artem-fold-towelThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "shape": [ 14 ], "names": [ "umi1_x", "umi1_y", "umi1_z", "umi1_rx", "umi1_ry", "umi1_rz", "umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.tabularrobotics1M<n<10M0 likes23k downloads23d agoHugging Face11allenai /MolmoAct2-BimanualYAM-DatasetThis dataset was created using LeRobot. MolmoAct2-BimanualYAM Dataset This repository is the merged ckpt / merged LeRobot dataset artifact for the MolmoAct2-BimanualYAM Dataset, a large-scale collection of bimanual robot manipulation demonstrations collected for MolmoAct2. Across the full collection, MolmoAct2-BimanualYAM contains more than 720 hours of training demonstrations spanning diverse tabletop manipulation tasks. Language Annotations This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-BimanualYAM-Dataset.tabularrobotics10M<n<100M10 likes22k downloads3mo agoHugging Face12laion /BVD-I-300M-URLs LAION-BVD - 300M Video Frame URLs This repository contains the URLs for ~300 million keyframes extracted from publicly available web videos. No image data is included, only the source video URL and the frame timestamp needed to reproduce each frame. Frames were extracted from BVD-RAW and cover YouTube, Dailymotion, and Vimeo content. Dataset structure Column Type Description webpage_url string URL of the source video frame_pts_time float Presentation… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-I-300M-URLs.textimage-text-to-text100M<n<1B2 likes22k downloads26d agoHugging Face13HuggingFaceFV /finevideogated FineVideo FineVideo Description Dataset Explorer Revisions Dataset Distribution How to download and use FineVideo Using datasets Using huggingface_hub Load a subset of the dataset Dataset StructureData Instances Data Fields Dataset Creation License CC-By Considerations for Using the Data Social Impact of Dataset Discussion of Biases Additional Information Credits Future Work Opting out of FineVideo Citation Information Terms of use for FineVideo… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFV/finevideo.textvisual-question-answering10K<n<100K382 likes19k downloads5mo agoHugging Face14lerobot /droid_1.0.1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/droid_1.0.1.tabularrobotics10M<n<100M24 likes17k downloads3mo agoHugging Face15lerobot /aloha_sim_transfer_cube_humanThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 20000, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_human.tabularrobotics10K<n<100K16 likes16k downloads4mo agoHugging Face16laion /BVD-V-55M-URLs LAION-BVD - 55M Video Clips (URL Release) This repository contains the metadata and captions for ~55 million scene-level video clips sourced from 2.4M randomly sampled videos from BVD-RAW. The 2.4M original videos are filtered to only include videos between 10s and 30min duration and are then split into the ~55M scene clips using PySceneDetect. No video or audio files are included; only URLs, timestamps, and text annotations are provided. Repository structure… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-V-55M-URLs.imagevideo-text-to-text10M<n<100M5 likes16k downloads26d agoHugging Face17lerobot /pushtThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 206, "total_frames": 25650, "total_tasks":1, "total_videos": 206, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:206" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/pusht.tabularrobotics10K<n<100K59 likes16k downloads1y agoHugging Face18liang12121 /dreamzero-egoverse-360-pretraintabular10M<n<100M1 likes15k downloads5mo agoHugging Face19BitRobot /HIW-500-LeRobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.images.head": { "dtype": "video", "shape": [ 480, 1280, 3 ], "names": [ "height", "width", "channels" ], "info": { "video.height":… See the full description on the dataset page: https://huggingface.co/datasets/BitRobot/HIW-500-LeRobot.tabularrobotics10M<n<100M24 likes14k downloads2mo agoHugging Face20ember-lab-berkeley /robocasa365-pretrain-mg Pretraining (MimicGen) — atomic MimicGen-generated rollouts across 60 atomic tasks (~10,000 demos/task). 1,615 hours total, generated by scripted augmentation from human demonstrations. Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable. Stats Episodes: 536,030 Frames: 116,246,439 (20 fps → 1615 h) Tasks: 720 (natural-language phrasings; underlying RoboCasa task classes: 60) Cameras: 3 × 256×256 h264 video… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-pretrain-mg.tabularrobotics100M<n<1B2 likes13k downloads4mo agoHugging Face21lerobot /liberoThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "panda", "total_episodes": 1693, "total_frames": 273465, "total_tasks": 40, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 10.0, "splits": { "train": "0:1693" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/libero.tabularrobotics100K<n<1M21 likes12k downloads4mo agoHugging Face22ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads9h agoHugging Face23andlyu /Public-YAM-runs Public-YAM-runs Physical bimanual YAM episodes recorded by the BluPe operator station. Each run adds an episode to this repository. Failed, interrupted, stopped and timed-out runs are retained and labeled; these are not all successful demonstrations. A model saying done is not independently verified task success. Loading from datasets import load_dataset runs = load_dataset("andlyu/Public-YAM-runs", split="train") usable = runs.filter(lambda row:… See the full description on the dataset page: https://huggingface.co/datasets/andlyu/Public-YAM-runs.image100K<n<1M2 likes10k downloads1d agoHugging Face24MME-Benchmarks /Video-MME-v2 🔥 News 2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch. 2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups. 🤗 About This Repo This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.textvideo-text-to-text1K<n<10K48 likes9.3k downloads1mo agoHugging Face25lerobot /aloha_sim_insertion_humanThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 25000, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_human.tabularrobotics10K<n<100K13 likes8.6k downloads4mo agoHugging Face26RekaAI /RekaDaily-10k-processed RekaDaily-10k (processed) Short first-person clips cut from the RekaDaily-10k recordings — unscripted daily-life video collected through Claru, Reka's data collection marketplace, recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Every clip carries one dense caption and a multi-question Q&A exchange written in the second person ("What am I doing in this video?"), so the corpus drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.imagevideo-text-to-text1M<n<10M2 likes8.5k downloads9d agoHugging Face27nyu-visionx /VSI-Bench Dataset arXiv Website Code VSI-Bench VSI-Bench-Debiased v1 [!IMPORTANT] [Aug. 9, 2026] PROVENANCE UPDATE: The existing "Debiased" subset is VSI-Bench-Debiased v1, a designer-in-the-loop manual pilot created with bespoke per-question-type filtering heuristics. It predates and was not generated by the automated Iterative Bias Pruning (IBP) algorithm. We retain v1 for reproducibility and will version any future automated subset separately.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/VSI-Bench.textvisual-question-answering10K<n<100K71 likes8.3k downloads1mo agoHugging Face28lerobot /abc_130k_v3_trainThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "shape": [ 14 ], "names": [ "left_arm_joint_1", "left_arm_joint_2", "left_arm_joint_3", "left_arm_joint_4", "left_arm_joint_5"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/abc_130k_v3_train.tabularrobotics100M<n<1B1 likes8.3k downloads3mo agoHugging Face29laion /BVD-A-10M-URLs LAION-BVD — 10M Audio Clip URLs This repository contains the metadata and captions for ~10 million audio clips randomly sampled from BVD-V-55M for large-scale audio pre-training. The audio itself is not included in this repository — every clip is described by the URL of its source video plus the start_time/end_time offsets needed to reproduce it. The corresponding clip files are available in the gated laion/BVD-A-10M repository. Dataset structure One row per audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-10M-URLs.tabulartext-to-audio10M<n<100M1 likes7.5k downloads26d agoHugging Face30elonelonelon /2025-challenge-demosThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "R1Pro", "total_episodes": 10000, "total_frames": 119094660, "total_tasks": 50, "total_videos": 90000, "chunks_size": 10000, "fps": 30, "splits": { "train": "0:10000" }, "data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/elonelonelon/2025-challenge-demos.tabularrobotics100M<n<1B1 likes7.4k downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.