datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.GoPro-Raw-Videos
Raw GoPro Videos for Four Robotic Manipulation Tasks
[Project Page]
[Paper]
[Code]
[Models]
[Processed Dataset]
This repository contains raw GoPro videos of robotic manipulation tasks collected in-the-wild using UMI, as described in the paper "Data Scaling Laws in Imitation Learning for Robotic Manipulation". The dataset covers four tasks:
Pour Water
Arrange Mouse
Fold Towel
Unplug Charger
Dataset Folders:
arrange_mouse and pour_water: Each folder contains data… See the full description on the dataset page: https://huggingface.co/datasets/Fanqi-Lin/GoPro-Raw-Videos.GOKU-2M
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
GOKU-2M is a large-scale, unified instruction-based video-editing dataset covering 10 editing tasks. Each sample provides a source video, an edited target video, and one or more natural-language instructions describing the edit.
📦 Repositories
⚠️ Because a single Hugging Face account has a free storage quota of about 8.7 TB, the dataset is split across two… See the full description on the dataset page: https://huggingface.co/datasets/bigfacing/GOKU-2M.GOKU-2M
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
GOKU-2M is a large-scale, unified instruction-based video-editing dataset covering 10 editing tasks. Each sample provides a source video, an edited target video, and one or more natural-language instructions describing the edit.
📦 Repositories
⚠️ Because a single Hugging Face account has a free storage quota of about 8.7 TB, the dataset is split across two… See the full description on the dataset page: https://huggingface.co/datasets/Goku-2M/GOKU-2M.GOAI-2026libero_goal_no_noops_1.0.0_lerobotb1k-224x224-gop8-fixed
BEHAVIOR-1K 2026 — 224x224, GOP=8, upstream-matched encoding
A 224x224 re-encode of the 2026 challenge demos (100 tasks, 3 RGB cameras) whose image
statistics match the dataset the widely-used 50-task checkpoint was pretrained on
(IliaLarchenko/behavior_224_rgb), while keeping GOP=8 for fast random-frame access during
training.
Why "fixed"
Earlier 224 re-encodes of this data used bicubic + libx264 CRF 23, which lands 13% softer
(high-frequency content) than the… See the full description on the dataset page: https://huggingface.co/datasets/JackLiu0406/b1k-224x224-gop8-fixed.S10-Citadel-Core
Run and deploy your AI Studio app
This contains everything you need to run your app locally.
Run Locally
Prerequisites: Node.js
Install dependencies:
npm install
Set the GEMINI_API_KEY in .env.local to your Gemini API key
Run the app:
npm run dev
godseye-violence-detection-dataset
Keypoints-RWF-total Dataset
This dataset is a fusion of three distinct datasets:
RWF-2000: A dataset that includes videos of real-world fights and non-fight scenarios.
Hockey Violence Dataset: A dataset focused on violent interactions in hockey games.
Airtlab Violence Dataset: A dataset containing videos of violent incidents in various settings.
Important Note: We do not own these datasets. By using this dataset, you are agreeing to follow the rules and licensing agreements of… See the full description on the dataset page: https://huggingface.co/datasets/valiantlynxz/godseye-violence-detection-dataset.rlwrld_ICLRSynthSite
SynthSite
SynthSite is a curated benchmark of 227 synthetic construction site safety videos (115 unsafe, 112 safe) generated using four text-to-video models: Sora 2 Pro, Veo 3.1, Wan 2.2-14B, and Wan 2.6. Each video was independently labeled by 2–3 human reviewers for the presence of a Worker Under Suspended Load hazard, producing a binary classification: unsafe (True_Positive — worker remains in the suspended-load fall zone) or safe (False_Positive — no worker in fall zone or… See the full description on the dataset page: https://huggingface.co/datasets/govtech/SynthSite.DFDC-extracted-fullAgiBotWorld-Beta_G1_task_566_Place_the_goods_in_the_material_box_on_the_shelf
agibot_task_566
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 将物料箱中的货物放到货架上part_1
total_episodes: 281
total_tasks: 1
size: 31G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_566_Place_the_goods_in_the_material_box_on_the_shelf.pose6daug
pose6daug
Real-world Franka manipulation episodes with object-swap and action augmentation
artifacts. 120 training episodes over 4 objects (blue_cup, green_pear, kanu,
white_spray), dual ZED cameras (exo static + ego wrist-mounted).
Layout
Per-frame PNGs are packed into uncompressed tars per episode — the dataset has
~427k mask/plate frames and loose files hit Hugging Face's per-repo file
recommendation and API rate limits hard.
data/<object>/<NNNN>/
masks.tar… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/pose6daug.libero_plus_goalIndian_Sign_Language_Data.gov_Rencoded
Indian_Sign_Language_Data.gov_Rencoded
Dataset Overview
Dataset name: Indian_Sign_Language_Data.gov_RencodedHugging Face repository: silentone0725/Indian_Sign_Language_Data.gov_RencodedModality: Video (H.265 / HEVC)Total size: ~75 GBOriginal size: ~200 GBLanguage: Indian Sign Language (ISL)License: MIT
This dataset is a re-encoded and curated version of the Indian Sign Language Dictionary originally published on the Government of India Open Data Portal (data.gov.in).… See the full description on the dataset page: https://huggingface.co/datasets/silentone0725/Indian_Sign_Language_Data.gov_Rencoded.behavior_3egocentric-gopro-rgb-imu
Hub Egocentric: GoPro RGB+IMU
19 egocentric human-manipulation clips captured on GoPro HERO13, with high-rate IMU (~200 Hz GPMF) delivered as CSV/Parquet/JSON sidecars plus a Foxglove MCAP recording.
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on GoPro HERO13. Egocentric, human-demonstration data (passive; no robot action stream). July 2026.
Dataset structure
Each clip is a top-level folder named #NN_... holding its media, a… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-gopro-rgb-imu.AIRBOT_MMK2_place_the_shark_toys_and_gold_bars
AIRBOT_MMK2_place_the_shark_toys_and_gold_bars
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_shark_toys_and_gold_bars.war-gov-uap-release
War.gov UAP Corpus
This dataset is a provenance-first mirror and source index for public files
listed in the U.S. Department of War UAP/UFO release portal.
The goal is to preserve the official War.gov release listing, public source
URLs, verified file hashes, byte sizes, and raw source assets before
summarization or interpretation is added.
Contents
Current contents:
sources: the complete public source index, with one row per public source
asset listed or linked… See the full description on the dataset page: https://huggingface.co/datasets/proxima/war-gov-uap-release.go2_object_approach_v1
Go2 Object Approach v1
Keyboard-teleoperated Unitree Go2 EDU trajectories for the task:
"Approach the target object and stop in a manipulation-ready pose."
The collection contains 202 episodes and 48,889 frames at a nominal 20 Hz
(2,444.45 seconds, approximately 40.74 minutes). Episode lengths range from
109 to 502 frames (5.45–25.10 seconds). Data is stored in LeRobot Dataset v3
with Parquet telemetry and H.264 front-camera video. No audio or depth is included.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/dancher00/go2_object_approach_v1.AIRBOT_MMK2_place_the_glasses_case_and_gold_bars
AIRBOT_MMK2_place_the_glasses_case_and_gold_bars
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_glasses_case_and_gold_bars.deepfake_video_dataowm-iss-numerical-v1-nonoise-goal-dt50ms-500k
owm-iss-numerical-v1-nonoise-goal-dt50ms-500k
Docking approaches to the International Space Station: a 12-tonne Dragon-class chaser
manoeuvring from starts between 100 m and 225 m out to a station docking port,
under two-vehicle ECI propagation with configurable J2-J6 zonal gravity, third-body and drag perturbations, and inertial attitude dynamics and against the station's 313-box collision hull. Generated with
owm-envs for world-model training,
at 20 Hz (dt = 0.05 s).
The… See the full description on the dataset page: https://huggingface.co/datasets/sislaboratory/owm-iss-numerical-v1-nonoise-goal-dt50ms-500k.GoogleDeepMind-NEPTUNEdk1_2026-03-01This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_dk1_follower",
"total_episodes": 141,
"total_frames": 306425,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:141"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Gongsta/dk1_2026-03-01.go2_rl_gym_videos
go2 rl gym videos
This is all original videos for project:
Code:
Train: go2_rl_gym
Evaluation: RoboGauge
Deploy: unitree_cpp_deploy
Page: robogauge.github.io
Paper: Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion
aopoli-lv-libero_combined_no_noops_lerobot_v21This dataset was created using LeRobot.
Dataset Description
Combined version of following lerobot datasets
aopolin-lv/libero_spatial_no_noops_lerobot_v21
aopolin-lv/libero_object_no_noops_lerobot_v21
aopolin-lv/libero_goal_no_noops_lerobot_v21
aopolin-lv/libero_10_no_noops_lerobot_v21
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1"… See the full description on the dataset page: https://huggingface.co/datasets/godnpeter/aopoli-lv-libero_combined_no_noops_lerobot_v21.010_pickplace_googleball_3Cam_MQ
010_pickplace_googleball_3Cam_MQ
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
trlcdk1_pickplace
