datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
calvin-task-ABCD-D-lerobotproduct-photography-v1-tiny-prompts-tasks-collage-filteredTaskMeAnything-v1-imageqa-2024
Dataset Card for TaskMeAnything-v1-imageqa-2024
TaskMeAnything-v1-imageqa-2024 benchmark dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-2024
TaskMeAnything-v1-imageqa-2024 is a benchmark for reflecting the current progress of MLMs by automatically finding tasks that SOTA MLMs struggle with using the TaskMeAnything Top-K queries.
This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-imageqa-2024.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.PENGWIN_Task2
PENGWIN Task 2: Pelvic Fragment Segmentation on Synthetic X-ray Images
Mirror of the training split of Task 2 of the MICCAI 2024 PENGWIN challenge
(https://pengwin.grand-challenge.org/), from the official Zenodo record
10913196 (train.zip, md5 9c90215dae54d8f494a85cfc7b19bc96).
These are SYNTHETIC X-rays, not real radiographs: DeepDRR renders of the 100 PENGWIN
Task 1 training CTs simulating intraoperative C-arm fluoroscopy, 500 random poses per CT
= 50,000 image/mask pairs.… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/PENGWIN_Task2.TaskMeAnything-v1-imageqa-random
Dataset Card for TaskMeAnything-v1-imageqa-random
TaskMeAnything-v1-imageqa-random dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-Random
TaskMeAnything-v1-imageqa-random is a dataset which using
randomly sampled questions from TaskMeAnything-v1, including 5,700 ImageQA questions. The dataset contains 19 splits, while each splits contains 300… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-imageqa-random.calvin-task-ABC-D-lerobot-v30calvin_task_D_D_scale_100_lerobo_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 5124,
"total_frames": 303794,
"total_tasks": 389,
"total_videos": 0,
"total_chunks": 6,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:5124"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/calvin_task_D_D_scale_100_lerobo_format.PENGWIN_Task1
PENGWIN Task 1 — Pelvic Fracture Segmentation on CT
The CT task of the PENGWIN 2024 challenge (PElvic bone fraGment (WIN)dow,
MICCAI 2024): segment the sacrum, left hipbone and right hipbone, and the
individual fracture fragments of each, in preoperative pelvic trauma CT.
This is an instance segmentation task, not a 3-class semantic one — the
label value identifies which fragment of which bone, and the fragment count
varies per case.
What this mirror contains — read… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/PENGWIN_Task1.calvin-task-D-D-lerobottask_all_origtask_book_leavesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 38,
"total_frames": 8664,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:38"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_book_leaves.UVT-Explanatory-based-Vision-Tasks
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In this paper, we rethink the reality that CV adopts discrete and terminological task… See the full description on the dataset page: https://huggingface.co/datasets/axxkaya/UVT-Explanatory-based-Vision-Tasks.task_pour_leavesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 15685,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_pour_leaves.task_book_redThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 46,
"total_frames": 12634,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:46"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_book_red.object_task_suiteThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "mobile_aloha",
"total_episodes": 301,
"total_frames": 77500,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/object_task_suite.task_pull_leavesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 9649,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_pull_leaves.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.GUIrilla-Task
GUIrilla-Task
Ground-truth Click & Type actions for macOS screenshots
Dataset Summary
GUIrilla-Task pairs real macOS screenshots with free-form natural-language instructions and precise GUI actions.
Every sample asks an agent either to:
Click a specific on-screen element, or
Type a given text into an input field.
Targets are labelled with bounding-box geometry, enabling exact evaluation of visual-language grounding models.
Data were gathered automatically by… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/GUIrilla-Task.UVT-Terminological-based-Vision-Tasks
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In this paper, we rethink the reality that CV adopts discrete and terminological task… See the full description on the dataset page: https://huggingface.co/datasets/axxkaya/UVT-Terminological-based-Vision-Tasks.task_pull_redThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 21,
"total_frames": 4548,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_pull_red.task_pour_redThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 17345,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_pour_red.realworld_task820mbeir_webqa_task2libero_10_image_task_7This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 13476,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/libero_10_image_task_7.libero_10_image_task_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 48,
"total_frames": 12346,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:48"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/libero_10_image_task_4.robotwin_aloha_50_tasks_clean_mergedtask3-12This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 95,
"total_frames": 78004,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:95"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jasontchan/task3-12.libero_10_image_task_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 13298,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/libero_10_image_task_2.libero_40_image_task_0_30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 155,
"total_frames": 27179,
"total_tasks": 31,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:155"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/libero_40_image_task_0_30.
