datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
extended_video_activities_drop_4
Dataset Card for meva_mevid
This is a FiftyOne dataset with 201 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/extended_video_activities_drop_4")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/extended_video_activities_drop_4.DFDC-extracted-fullscene-extrapolation-benchmark
Scene-Extrapolation Benchmark Suite
A unified, honest evaluation of generative novel-view synthesis (NVS) in the regimes where
per-scene 3D-Gaussian-Splatting reconstruction fails:
Extrapolative — held-out, unobserved views (not interpolation between dense captures).
Dynamic — moving foreground content.
Long-horizon — chained / loop-closing camera trajectories where drift accumulates.
Memory / revisit — recall when the camera returns to a previously-observed pose (do methods… See the full description on the dataset page: https://huggingface.co/datasets/luuuulinnnn/scene-extrapolation-benchmark.extreme_randomization_6_brick_03This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 2108,
"total_frames": 419545,
"total_tasks": 1,
"total_videos": 4216,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2108"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/extreme_randomization_6_brick_03.GigaHands_extrastrawberry_picking_dataset_scara_extra
Strawberry Picking Dataset — SCARA Extra
Relationship to the paper: These are extra SCARA demonstrations collected
for other tasks and experimental setups. None of the data in this repository
were used in the training, evaluation, or other experiments reported in
Learning to Pick: A Visuomotor Policy for Clustered Strawberry Picking.
The paper is cited as related context for the robot platform and research area.
LeRobot v3.0 dataset converted from ACT/ALOHA HDF5 demos on… See the full description on the dataset page: https://huggingface.co/datasets/zfff/strawberry_picking_dataset_scara_extra.metal_part_sort_v10_plus_extra_20260702
metal_part_sort_v10_plus_extra_20260702
Merged LeRobot v2.1 dataset for Unitree G1 + Inspire DFX metal part sorting.
Base dataset: PID0930/metal_part_sort_v10 @ 17fe43673ca30e840b1baa6e05c1f34b7dd43b91
Extra dataset: PID0930/metal_part_sort_v10_extra_20260701 @ 09c5a9113830443f99a8d91c11fd5754b238fc6a
Episodes: 329
Frames: 110966
Cameras: external, left wrist, right wrist, plus left_high/head slot retained for compatibility
GR00T training uses external as the logical head… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/metal_part_sort_v10_plus_extra_20260702.robotwin_extrinsicsThis dataset was created using LeRobot.
Dataset Description
RoboTwin 2.0 (aloha-agilex) with end-effector poses and extrinsics
Generated from lerobot/robotwin_unified at commit 1287871839fae2296bc27b88a5457c3e1eba8e1f by
benchmarks/robotwin/augment_robotwin_dataset.py. Original state and action columns are
unchanged: 14 joint drive targets [left arm(6), left gripper, right arm(6), right gripper],
grippers in [0, 1] with 1 = open, at 30 Hz; the action is the… See the full description on the dataset page: https://huggingface.co/datasets/dgrachev/robotwin_extrinsics.b1k_2026_extras
BEHAVIOR-1K 2026 challenge extras
Precomputed tables and preview clips derived from
behavior-1k/2026-challenge-demos.
The source release is 3.0 TB across 17,093 video shards; this is 111 MB, and it is enough
to read the whole annotation layer and watch one demonstration per task without touching
the release at all.
Contents
path
what it is
b1k_skill_segments.parquet
one row per skill span, all 20,000 episodes: 407,054 rows of episode_index, task_index… See the full description on the dataset page: https://huggingface.co/datasets/mahgoobi/b1k_2026_extras.I3D_Extract_Videosphoton47-extended-film-highlights
Photon47 extended-film highlights
Seven at-least-60-second, silent H.264/MP4 real-gameplay reels derived from the public creator film
Photon47: Multiplayer Cute Monster Battle Game.
Each website mode card plays its matching reel first, followed by the existing
automated mouse/keyboard gameplay capture.
Contents
moba.addendum.mp4 — team clashes, hero swarms, beams, ice, and vortex effects
arena.addendum.mp4 — card deployment, counterpush, and finale… See the full description on the dataset page: https://huggingface.co/datasets/ryan-superman/photon47-extended-film-highlights.so101_strawberry_extDeepFake_Extracted_Face_Imagestooth_extraction_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 200,
"total_frames": 76053,
"total_tasks": 1,
"total_videos": 400,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.tooth_extraction_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 32879,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_3.so101_red_screwdriver_to_yellow_container_extra_03_clean60This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/poi69420/so101_red_screwdriver_to_yellow_container_extra_03_clean60.DS1_1B1O_FF_removeidle_rand80_extendlfpick-and-place-foam-blockso101_pickup_20260516_182617_task3_extra_camera_pose_fixed
Task3 Camera Pose Fixed
Camera Pose Fix
Videos visually unchanged but re-encoded to H.264: 6 (file-000.mp4 through file-005.mp4).
Videos transformed: 84 (file-006.mp4 onward).
Transform: rotate counter-clockwise by 2.1 degrees, translate content left by 30 pixels, then apply overlay_task3.png as an RGBA alpha overlay.
Data parquet files and metadata are unchanged from the source dataset.
extreme-when-bench
ExtremeWhenBench
Hour-scale natural-language temporal grounding benchmark.
2,273 open-form natural-language questions over 194 hour-long
videos (mean 75.7 min, max 9 hr) sourced from LVBench, MLVU, and
VideoMME. The median GT event is 9 s — matched to Charades-STA's 7.1 s
— so the same event grain now sits inside a search space ~153× larger.
Companion to Natural-Language Temporal Grounding in Hour-Long Videos is a
Search Problem: A Benchmark and Empirical Decomposition —… See the full description on the dataset page: https://huggingface.co/datasets/min1321/extreme-when-bench.SO100_evaluate_generalize_pick_pos_extendThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 30,
"total_frames": 11184,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos_extend.metal_part_sort_v10_extra_20260701
metal_part_sort_v10_extra_20260701
Extra LeRobot v2.1 dataset recorded on Unitree G1 + Inspire DFX for metal part sorting.
Source raw episodes: episode_0311 through episode_0330 from local metal_part_sort_v10
Episodes: 20
Frames: 9372
Cameras: external, left wrist, right wrist, plus left_high/head slot retained for compatibility
GR00T training uses external as the logical head camera input
State/action: 26 dims (left arm 7, right arm 7, left hand 6, right hand 6)
dexterous-hand-pick-and-placeablation2_extract_cube_B2_10fps
Ablation2 Extract Cube B2 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Extract the cube from the pocket and place it on the target marker."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 100 (31,509 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/ablation2_extract_cube_B2_10fps.astra_grab_floor_toys_extendedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "astra_joint",
"total_episodes": 80,
"total_frames": 113547,
"total_tasks": 1,
"total_videos": 240,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lookas/astra_grab_floor_toys_extended.single_so101_cubes_extendedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/alex-luci/single_so101_cubes_extended.UCF_crime_extract_eventomx_ExtendedTest3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "aiworker",
"total_episodes": 50,
"total_frames": 17413,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ROBOTIS-junha/omx_ExtendedTest3.extract_cube_ours_10fps
Extract Cube Ours 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Extract the cube from the pocket and place it on the target marker."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 100 (31,555 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/extract_cube_ours_10fps.fastwam-robotwin2-extracted
