datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multitask_vqa_benchmarkThis dataset is a part of .
🍈 MMT-47: Multimodal Multi-Task Benchmark
47 Tasks · 7 Categories · 3 Modalities (Image, Video, Text)
Cite our ICML-2026 paper for this dataset
@article{kowsher2026lime,
title={LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning},
author={Kowsher, Md and Mansoor, Haris and Prottasha, Nusrat Jahan and Garibay, Ozlem and Zhu, Victor and Ji, Zhengping and Chen, Chen},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/Kowsher/multitask_vqa_benchmark.VQA-MultiTaskReal_World_Multi_Taskur5e_multitask_000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 270,
"total_frames": 106900,
"total_tasks": 11,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:270"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Perseus101/ur5e_multitask_000.franka_multi_task_v1_imgThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 75843,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/franka_multi_task_v1_img.ur10e-multitask
UR10e Multitask Teleop
335 teleoperated manipulation episodes on a Universal Robots UR10e with a Robotiq 2F gripper, covering 7 tabletop tasks. Packaged in LeRobot v2.1 format with precomputed normalization statistics, so it can be dropped into a VLA finetuning run the same way LIBERO is.
Episodes
335
Frames
60,416
Tasks
7
Control rate
15 fps
Robot
UR10e, 6-DoF, Robotiq 2F gripper
Cameras
2 exterior (Azure Kinect) + 1 wrist (RealSense)
Image size
180 x… See the full description on the dataset page: https://huggingface.co/datasets/bag100/ur10e-multitask.xlerobot_multitask_part10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 5,
"total_frames": 3694,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 0.001,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ArthurWangSawau/xlerobot_multitask_part10.xlerobot_multitask_part11This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 6,
"total_frames": 5290,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 0.001,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ArthurWangSawau/xlerobot_multitask_part11.xlerobot_multitask_part13This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 21,
"total_frames": 13987,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 0.001,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ArthurWangSawau/xlerobot_multitask_part13.MultiTasks-v2
MultiTasks-v2
MultiTasks-v2 is a collection of twelve multimodal benchmark subsets normalized
into a shared image-question-answer format. Each subset provides a train split
and a test split.
Dataset Structure
Each example contains:
id: a unique sample identifier in the form {Dataset}_{split}_{index}.
images: a list containing one image. Images are stored as JPEG bytes.
problem: the prompt shown to the model.
answer: the target answer.
RefAdv is the only subset… See the full description on the dataset page: https://huggingface.co/datasets/LoserLi/MultiTasks-v2.satquery-multitask-datasetfranka_multi_task_v1_imgThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 75843,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/franka_multi_task_v1_img.ur5_multitaskbatch_indexing_machine_multitask
Dataset Card for "batch_indexing_machine_multitask"
More Information needed
satellite-multitask-omni
🛰️ Satellite Multi-Task Omni Dataset
A unified, multi-task satellite/aerial imaging dataset designed for training omni-models that work with image+text as both input and output modalities. All data is converted to a consistent ChatML conversational format.
📊 Dataset Overview
Metric
Value
Total Samples
34,894
Train / Val / Test
31,404 / 1,744 / 1,746
Tasks
9 distinct task types
Sources
10 source datasets
Format
ChatML conversations + images… See the full description on the dataset page: https://huggingface.co/datasets/rahuldshetty/satellite-multitask-omni.so100_multitask_80-trajectoryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 80,
"total_frames": 38620,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/so100_multitask_80-trajectory.multi_task_v1MultiTasks
MultiTasks
This dataset collection contains six multimodal benchmark subsets. Each subset
provides a train split and a test split with the columns images,
problem, and answer.
RefAdv uses a list-valued answer for bounding boxes. The other subsets use
string-valued answers.
multitask_vtab9kweb-agent-multitask-flat-failed-stepsmultitask-diagnostic-suite-vlmaloha_real_agilex_pull_drawer_dummy_multitask-20250730multitask_mllmweb-agent-multitask-runs-failedweb-agent-multitask-flat-successful-shortest-steps
