datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MVLUMLVU: Multi-task Long Video Understanding Benchmark
This repo contains the annotation data and evaluation code for the paper "MLVU: A Comprehensive Benchmark for Multi-Task Long Video Understanding".
🔔 News:
🆕 7/28/2024: The data for the MLVU-Test set has been released (🤗 Link)! The test set includes 11 different tasks, featuring our newly added Sports Question Answering (SQA, single-detail LVU) and Tutorial… See the full description on the dataset page: https://huggingface.co/datasets/MLVU/MVLU.MLVU_devSuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.MLVUg1-inspire-pipette-tip
G1 + Inspire — "eject the pipette tip into the rack" (GR00T N1.7 / G1_INSPIRE)
Teleoperated Unitree G1 (29-DoF) + Inspire RH56DFTP hands manipulation data,
converted to the GR00T-flavored LeRobot v2.1 format for fine-tuning GR00T N1.7
as a custom NEW_EMBODIMENT (here called G1_INSPIRE).
Same schema as
MLeggiero/g1-gr00t-inspire-pick_and_place,
with two additions: observation.effort (per-joint torque) and native
1280×720 ego-view video and depth instead of 424×240. Read
Known… See the full description on the dataset page: https://huggingface.co/datasets/MLeggiero/g1-inspire-pipette-tip.MLVU_TestOmni-VFX
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
🔥 Updates
[2025/08] We release the CogVideoX-1.5 finetuned on our Omni-VFX dataset !
[2025/08] We release the controllable single-VFX/Multi-VFX version of Omni-Effects!
📣 Overview
Visual effects (VFX) are essential visual enhancements fundamental to modern cinematic production. Although video generation models offer cost-efficient solutions for VFX production, current… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/Omni-VFX.dpgs_sim_faucet_5kMLVUml-lectures
ML/Math Lecture Archive
Archived lecture videos (1080p MP4) with English subtitles (.vtt) from publicly
available university course recordings on YouTube.
268 videos, ~72 GB.
Contents
Folder
Course
Videos
18.065_Strang/
MIT 18.065 — Matrix Methods in Data Analysis, Signal Processing, and Machine Learning (Gilbert Strang)
36
18.06SC_LinearAlgebra/
MIT 18.06SC — Linear Algebra, Fall 2011 (Gilbert Strang)
74
CS109_Piech/
Stanford CS109 — Introduction… See the full description on the dataset page: https://huggingface.co/datasets/olive5/ml-lectures.MLVUbehavior_3MLV-Bench
MedHorizon / MLV-Bench
MedHorizon, also released as MLV-Bench, is a long-context medical video benchmark for evaluating multimodal models on full-procedure clinical videos. The benchmark emphasizes two properties that are not captured by short-clip medical video datasets: extremely sparse evidence retrieval and multi-hop reasoning over observations distributed across a full procedure.
Dataset Contents
Videos: 340 full-procedure videos.
Questions: 1,253 multiple-choice QA… See the full description on the dataset page: https://huggingface.co/datasets/DBD123/MLV-Bench.synwts
SynWTS: Synthetic Woven Traffic Safety Dataset
SynWTS is a high-fidelity synthetic dataset built as a Digital Twin of the Woven Traffic Safety (WTS) dataset. It is developed for the 2026 AI City Challenge (Track 2) to advance Sim2Real research in transportation safety understanding.
Dataset Summary
Participants in the Sim2Real challenge must train models exclusively on this synthetic data and evaluate performance on real-world video. SynWTS provides a geometric… See the full description on the dataset page: https://huggingface.co/datasets/mlcglab/synwts.MLVUTestvace-aug-224-gr00t
VACE-augmented RoboCasa dataset — 224 episodes (GR00T / LeRobot v2.1)
The exact 224-episode training set used to fine-tune GR00T-1.5 in the VACE-augmentation
baseline. Built by swapping a target object into RoboCasa pick-and-place episodes with the
VACE video-diffusion model, then converting to the GR00T "gr00t_views" (LeRobot v2.1) format.
224 episodes, 62,445 frames, 3 camera views (left_view / right_view / wrist_view
= RoboCasa robot0_agentview_left / robot0_agentview_right… See the full description on the dataset page: https://huggingface.co/datasets/mlnha/vace-aug-224-gr00t.g1-inspire-turn-page-twist2
G1 + Inspire — "turn the page of the notebook" (TWIST2 high-level)
Teleoperated Unitree G1 (29-DoF) + Inspire RH56DFTP hands manipulation data,
converted to LeRobot v3.0 with the TWIST2 converter
(deploy_real/convert_twist2_to_lerobot.py, --action_mode high_level).
A left-handed, thin-deformable manipulation task: the robot slides a single
notebook page off the stack and flips it over. Same recorder, same schema and the
same 48/49-dim vector layout as… See the full description on the dataset page: https://huggingface.co/datasets/MLeggiero/g1-inspire-turn-page-twist2.pi0_conversion_no_pad_videoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1417,
"total_frames": 166855,
"total_tasks": 33,
"total_videos": 2834,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:1417"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mlfu7/pi0_conversion_no_pad_video.telearms_yam_clean_30hz_recovereddpgs_sim_led_5kg1-inspire-nut-in-box-twist2
G1 + Inspire — "pick up the tweezers and put the nut in the box" (TWIST2 high-level)
Teleoperated Unitree G1 (29-DoF) + Inspire RH56DFTP hands manipulation data,
converted to LeRobot v3.0 with the TWIST2 converter
(deploy_real/convert_twist2_to_lerobot.py, --action_mode high_level).
This capture adds two wrist cameras on top of the onboard ego-view — so there
are three synchronized RGB streams (one head + two wrist), not the single head
view of… See the full description on the dataset page: https://huggingface.co/datasets/MLeggiero/g1-inspire-nut-in-box-twist2.so101_3cam_mlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 6,
"total_frames": 1621,
"total_tasks": 1,
"total_videos": 18,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ganondorofu/so101_3cam_ml.so100_brickThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 297,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_brick.so100_test_10_brick
so100_test_10_brick
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
lerobot-hackathon-v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 5,
"total_frames": 1462,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ml6team/lerobot-hackathon-v3.libero_spatial_no_noops_1.0.0_leroboticrt_otter_conversionso100_test_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 596,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mldev19/so100_test_4.rollout_eval_tac_mlp
rollout_eval_tac_mlp
Rollout evaluation of Jingyi-Z/act-sotac-tac-mlp (ACT + raw tactile frame
token, trained on sotac ep 0-70) on the SO-101 bench, 2026-09-16.
Protocol: bowl fixed at point 1, ball at points 6, 5, 4, 3, 2, 7; five
consecutive 60 s episodes per scene, 30 s reset. Stock lerobot 0.6.1 rollout
with the so101_paxini robot type; observation.state is 318-dim (6 joints +
312 tactile forces at 30 Hz), so tactile is recorded in every episode.
Cameras: top idx 0, wrist… See the full description on the dataset page: https://huggingface.co/datasets/Jingyi-Z/rollout_eval_tac_mlp.lerobot-hackathon-dummy-dataset-v2-red-framesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 5,
"total_frames": 1462,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ml6team/lerobot-hackathon-dummy-dataset-v2-red-frames.
