datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
e94fjt-v654-datacmp-v6-base108-renderkbcpjv-v654-datal36l5h-v654-datal36l5h-v654-rawe94fjt-v654-rawkbcpjv-v654-rawVR-egoverse-annotation-curated-v6.0
VR-egoverse-annotation-full (egoverse_v06)
Egocentric VR hand-tracking corpus: per-frame hand pose, wrist/camera trajectories, and
language narratives paired with head-mounted video clips. Schema 0.6.0, generated
2026-07-23, tooling commit e885040.
Private, in-progress upload. This release is being synced from local storage in
the background, incrementally and in chunks (see Layout below); clip counts on
the Hub will grow until the sync catches up to the full local release.… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egoverse-annotation-curated-v6.0.midjourney-v6-recap
Midjourney v6 Recaptioned
~1.2M Midjourney v6 images with captions from three VLMs:
llava: Original LLaVA captions from the source dataset
gemini: Gemini Flash 1.5 captions
qwen3: Qwen3 VL 8B captions
Caption coverage
llava: available for all 1,235,432 images (from original dataset)
gemini and qwen3: available for 1,017,105 images (82.3%)
Source
Based on brivangl/midjourney-v6-llava.
train_v6_deepseekmidjourney-v6-llavaThis dataset based on https://huggingface.co/datasets/CortexLM/midjourney-v6 dataset, captioned with LLava-1.6 model.
This dataset was released as is. By accessing and using this dataset, you acknowledge and agree that Cortex Foundation and the author of this repo are not responsible for any copyright violations or legal consequences that may arise from the use of these images.
train_v6_qwen32bgenomes-v4-genome_set-animals-intervals-v6_256_128HELM-Easiness-Data-10B-Labeled-v6synthetic-gauges-v6train_v6_qwenmegatron-prof-data-v6VR-egodex-annotation-converted-v6.0
VR-egodex-annotation-converted-v6.0
EgoDex converted from LeRobot v2.1 into the Layer-1 v0.6.0 annotation schema, with
per-clip narration included as language sidecars.
314,839 clips · 78,282,306 frames · 724.8 hours @ 30 fps · 129 tasks
100% narration coverage (1 sidecar per clip)
71 GB annotations + 2.3 GB narratives
Videos are NOT included. This release contains annotations and narration only. Source
video lives in griffinlabs/EgoDex-LeRobot-v3.0;
orig_id in the manifest… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egodex-annotation-converted-v6.0.ngld-grape-leaf-vlm-w-img-without-diff-ref-v6
cleaned
Dataset Source and Credits
This dataset is derived from the Niphad Grape Leaf Disease Dataset (NGLD) published on Mendeley Data.
Original dataset:
Title: Niphad Grape Leaf Disease Dataset (NGLD)
Authors: Madhuri Dharrao, Deepak Dharrao, Rakesh Sonawane
Institution: Symbiosis Institute of Technology, Symbiosis International University
DOI: https://doi.org/10.17632/8nnd2ypcv3.1
License: CC BY 4.0
The original dataset contains high-quality images of table grape… See the full description on the dataset page: https://huggingface.co/datasets/qingwuuu/ngld-grape-leaf-vlm-w-img-without-diff-ref-v6.sre-2d-harbor-tasks-v6aime-solution-hint-v6-deepscaler-respgenso101_pick_plug_insert_v5_v6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 72,
"total_frames": 53488,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:72"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/777im/so101_pick_plug_insert_v5_v6.openarm_pick_v6
OpenArm Pick v6 - LeRobot Dataset
A LeRobot v2.1 dataset for bimanual robot manipulation, recorded in Isaac Sim with teleoperation.
Dataset Description
This dataset contains teleoperated demonstrations of a bimanual OpenArm robot performing pick-and-place tasks. It was created for training the Pi0.5 model for robotic manipulation.
Quick Stats
Property
Value
Robot
openarm_bimanual
Episodes
100
Total Frames
65,190
Duration
~36.2… See the full description on the dataset page: https://huggingface.co/datasets/qualiadev/openarm_pick_v6.rook_to_d4_v6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 3099,
"total_tasks":1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/7jep7/rook_to_d4_v6.grab_cup_2cam_v6
grab_cup_2cam_v6 — SO-101, two cameras
LeRobot v3.0 dataset for the SO-101 arm. Task prompt (exact string):
Grab the white cup. Cameras fixed and handeye, 30 fps.
grab_cup_2cam_v4 with the dead episode, the frozen-action lead-in, and the
long standing-still stretches removed.
v4
v6
episodes
146
145
frames
104,819
77,881
dead chunks
24.5%
0.0%
fabricated action steps
2.4%
2.8%
Why
A policy trained on v4 drives the arm to a correct pose and… See the full description on the dataset page: https://huggingface.co/datasets/endoard/grab_cup_2cam_v6.tblock-all-piper-clean-v6-bi_piper_followerThis dataset was created using LeRobot.
Dataset Description
Joint-level bimanual piper dataset derived from local/tblock-all-piper-clean-v6. The state is the LeRobot bi_piper_follower follower vector (left_shoulder_pan.pos, left_shoulder_lift.pos, left_elbow_flex.pos, ...) in that plugin's own units, so it trains and deploys through its stack unchanged. observation.state[t] contains the command at t and action[t] contains the command at t+1.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/murobotics/tblock-all-piper-clean-v6-bi_piper_follower.so100_test_v6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 40,
"total_frames": 23918,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/baptiste-04/so100_test_v6.mmu_jwst_primer_uds_grizli_v6.0_all_96
mmu_jwst_primer_uds_grizli_v6.0_all_96 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_jwst_primer_uds_grizli_v6.0_all_96.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_jwst_primer_uds_grizli_v6.0_all_96.W_hausa_v6v1v1d_docmatix_1k_normal_v6
