datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-task-tcc-robosuite
XIRL
Overview
Setup
Datasets
Code Navigation
Experiments: Reproducing Paper Results
Extending XIRL
Acknowledgments
Overview
Code release for our CoRL 2021 conference paper:
XIRL: Cross-embodiment Inverse Reinforcement Learning
Kevin Zakka1,3, Andy Zeng1, Pete Florence1, Jonathan Tompson1, Jeannette Bohg2, and Debidatta Dwibedi1
Conference on Robot Learning (CoRL) 2021
1Robotics at Google,
2Stanford… See the full description on the dataset page: https://huggingface.co/datasets/Renton-Ren/multi-task-tcc-robosuite.multitask_vqa_benchmarkThis dataset is a part of .
🍈 MMT-47: Multimodal Multi-Task Benchmark
47 Tasks · 7 Categories · 3 Modalities (Image, Video, Text)
Cite our ICML-2026 paper for this dataset
@article{kowsher2026lime,
title={LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning},
author={Kowsher, Md and Mansoor, Haris and Prottasha, Nusrat Jahan and Garibay, Ozlem and Zhu, Victor and Ji, Zhengping and Chen, Chen},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/Kowsher/multitask_vqa_benchmark.VQA-MultiTaskReal_World_Multi_Taskur5e_multitask_000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 270,
"total_frames": 106900,
"total_tasks": 11,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:270"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Perseus101/ur5e_multitask_000.franka_multi_task_v1_imgThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 75843,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/franka_multi_task_v1_img.ur10e-multitask
UR10e Multitask Teleop
335 teleoperated manipulation episodes on a Universal Robots UR10e with a Robotiq 2F gripper, covering 7 tabletop tasks. Packaged in LeRobot v2.1 format with precomputed normalization statistics, so it can be dropped into a VLA finetuning run the same way LIBERO is.
Episodes
335
Frames
60,416
Tasks
7
Control rate
15 fps
Robot
UR10e, 6-DoF, Robotiq 2F gripper
Cameras
2 exterior (Azure Kinect) + 1 wrist (RealSense)
Image size
180 x… See the full description on the dataset page: https://huggingface.co/datasets/bag100/ur10e-multitask.xlerobot_multitask_part13This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 21,
"total_frames": 13987,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 0.001,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ArthurWangSawau/xlerobot_multitask_part13.coastal-multitask-380
Coastal & Rural Bangladesh — Multi-Task Visual Dataset
379 field photographs (JPEG, native resolution as shot — see classification/metadata.csv for per-image width/height) collected on foot along the Bakkhali river embankment and surrounding villages/farmland near Cox's Bazar, Bangladesh, structured into three ML-task "levels": classification, semantic segmentation, and change detection.
Source: huggingface data 06 (Golam Rob / Tawhid Enterprise photo collection).… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/coastal-multitask-380.GPRadar-Defect-MultiTask
GPRadar-Defect-MultiTask 数据集
本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。
数据集结构
数据集组织如下:
dataset/
├── annotations/ - 包含JSON和JSONL格式的标注文件
│ ├── _annotations.train.jsonl - 训练集标注
│ ├── _annotations.valid.jsonl - 验证集标注
│ ├── _annotations.test.jsonl - 测试集标注
│ ├── p-1.v1i.paligemma/ - 主数据集元数据
│ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据
├── images/ - 包含所有图像文件
特点
包含874张带注释的地质雷达扫描图像
图像预处理为640x640像素大小
支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/GPRadar-Defect-MultiTask.paligemma-multitask-dataset
PaliGemma Multitask Dataset
This dataset is designed for training and evaluating the PaliGemma multitask model for defect detection and analysis. It combines a base set of annotated samples with an extended collection of 874 real-world structural inspection images.
Dataset Description
Overview
The dataset contains images of structural defects along with their corresponding annotations for:
Object detection (bounding boxes)
Defect classification… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/paligemma-multitask-dataset.xlerobot_multitask_part10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 5,
"total_frames": 3694,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 0.001,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ArthurWangSawau/xlerobot_multitask_part10.GPRadar-Defect-MultiTask
GPRadar-Defect-MultiTask 数据集
本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。
数据集结构
数据集组织如下:
dataset/
├── annotations/ - 包含JSON和JSONL格式的标注文件
│ ├── _annotations.train.jsonl - 训练集标注
│ ├── _annotations.valid.jsonl - 验证集标注
│ ├── _annotations.test.jsonl - 测试集标注
│ ├── p-1.v1i.paligemma/ - 主数据集元数据
│ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据
├── images/ - 包含所有图像文件
特点
包含874张带注释的地质雷达扫描图像
图像预处理为640x640像素大小
支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/LiZHENGzai/GPRadar-Defect-MultiTask.xlerobot_multitask_part11This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 6,
"total_frames": 5290,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 0.001,
"fps": 30,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ArthurWangSawau/xlerobot_multitask_part11.Multitask-RoboTwinMultiTasks-v2
MultiTasks-v2
MultiTasks-v2 is a collection of twelve multimodal benchmark subsets normalized
into a shared image-question-answer format. Each subset provides a train split
and a test split.
Dataset Structure
Each example contains:
id: a unique sample identifier in the form {Dataset}_{split}_{index}.
images: a list containing one image. Images are stored as JPEG bytes.
problem: the prompt shown to the model.
answer: the target answer.
RefAdv is the only subset… See the full description on the dataset page: https://huggingface.co/datasets/LoserLi/MultiTasks-v2.satquery-multitask-datasetfranka_multi_task_v1_imgThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 75843,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lirislab/franka_multi_task_v1_img.ur5_multitaskMultiTaskNLP-TestDataset
MultiTaskNLP-Dataset
1. Introduction
The MultiTaskNLP-Dataset has undergone significant quality improvements through iterative data curation. In the latest version, we have substantially enhanced the data completeness and label accuracy by implementing rigorous annotation protocols and multi-stage quality assurance mechanisms. The dataset demonstrates outstanding quality metrics across various dimensions, including completeness, accuracy… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/MultiTaskNLP-TestDataset.batch_indexing_machine_multitask
Dataset Card for "batch_indexing_machine_multitask"
More Information needed
satellite-multitask-omni
🛰️ Satellite Multi-Task Omni Dataset
A unified, multi-task satellite/aerial imaging dataset designed for training omni-models that work with image+text as both input and output modalities. All data is converted to a consistent ChatML conversational format.
📊 Dataset Overview
Metric
Value
Total Samples
34,894
Train / Val / Test
31,404 / 1,744 / 1,746
Tasks
9 distinct task types
Sources
10 source datasets
Format
ChatML conversations + images… See the full description on the dataset page: https://huggingface.co/datasets/rahuldshetty/satellite-multitask-omni.so100_multitask_80-trajectoryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 80,
"total_frames": 38620,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/edmos7/so100_multitask_80-trajectory.multi_task_v1multitask_vtab9kmultitask-diagnostic-suite-vlmaloha_real_agilex_pull_drawer_dummy_multitask-20250730multitask_mllmweb-agent-multitask-flat-failed-stepsMultiTasks
MultiTasks
This dataset collection contains six multimodal benchmark subsets. Each subset
provides a train split and a test split with the columns images,
problem, and answer.
RefAdv uses a list-valued answer for bounding boxes. The other subsets use
string-valued answers.
