datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tasksThis dataset is for storing assets for https://huggingface.co/tasks and https://github.com/huggingface/huggingface.js/tree/main/packages/tasks
calvin-task-ABC-D-lerobotcalvin-task-ABCD-D-lerobotOCR_Taskvision-flan_191-task_1k
🚀 Vision-Flan Dataset
vision-flan_191-task-1k is a human-labeled visual instruction tuning dataset consisting of 191 diverse tasks and 1,000 examples for each task.
It is constructed for visual instruction tuning and for building large-scale vision-language models.
Paper or blog for more information:
https://github.com/VT-NLP/MultiInstruct/
https://vision-flan.github.io/
Paper coming soon 😊
Citation
Paper coming soon 😊. If you use Vision-Flan, please use the… See the full description on the dataset page: https://huggingface.co/datasets/Vision-Flan/vision-flan_191-task_1k.product-photography-v1-tiny-prompts-tasks-collage-filteredTaskMeAnything-v1-imageqa-2024
Dataset Card for TaskMeAnything-v1-imageqa-2024
TaskMeAnything-v1-imageqa-2024 benchmark dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-2024
TaskMeAnything-v1-imageqa-2024 is a benchmark for reflecting the current progress of MLMs by automatically finding tasks that SOTA MLMs struggle with using the TaskMeAnything Top-K queries.
This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-imageqa-2024.mind2web_multimodal_test_task
Dataset Card for Multimodal Mind2Web "Cross-Task" Test Split
Note: This dataset is the test split of the Cross-Task dataset introduced in the paper.
This is a FiftyOne dataset with 1338 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_task.TaskMeAnything-v1-imageqa-random
Dataset Card for TaskMeAnything-v1-imageqa-random
TaskMeAnything-v1-imageqa-random dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-Random
TaskMeAnything-v1-imageqa-random is a dataset which using
randomly sampled questions from TaskMeAnything-v1, including 5,700 ImageQA questions. The dataset contains 19 splits, while each splits contains 300… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-imageqa-random.TaskMeAnything-v1-sourcePENGWIN_Task2
PENGWIN Task 2: Pelvic Fragment Segmentation on Synthetic X-ray Images
Mirror of the training split of Task 2 of the MICCAI 2024 PENGWIN challenge
(https://pengwin.grand-challenge.org/), from the official Zenodo record
10913196 (train.zip, md5 9c90215dae54d8f494a85cfc7b19bc96).
These are SYNTHETIC X-rays, not real radiographs: DeepDRR renders of the 100 PENGWIN
Task 1 training CTs simulating intraoperative C-arm fluoroscopy, 500 random poses per CT
= 50,000 image/mask pairs.… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/PENGWIN_Task2.calvin-task-ABC-D-lerobot-v30Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.multi-task-tcc-robosuite
XIRL
Overview
Setup
Datasets
Code Navigation
Experiments: Reproducing Paper Results
Extending XIRL
Acknowledgments
Overview
Code release for our CoRL 2021 conference paper:
XIRL: Cross-embodiment Inverse Reinforcement Learning
Kevin Zakka1,3, Andy Zeng1, Pete Florence1, Jonathan Tompson1, Jeannette Bohg2, and Debidatta Dwibedi1
Conference on Robot Learning (CoRL) 2021
1Robotics at Google,
2Stanford… See the full description on the dataset page: https://huggingface.co/datasets/Renton-Ren/multi-task-tcc-robosuite.calvin-task-D-D-lerobottask_pour_leavesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 15685,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_pour_leaves.TaskGrasp-Pro
TaskGrasp-Pro dataset
The TaskGrasp-Pro dataset extends the original TaskGrasp dataset by providing fine-grained part decompositions and part-level physical property annotations for 190 household objects. For each object instance, we design three types of tasks: category-related tasks, part-related tasks, and part-irrelevant tasks, resulting in a total of 2,850 tasks.
Files
scans/ contains the point clouds of objects, multi-view RGB-D images, different types of task… See the full description on the dataset page: https://huggingface.co/datasets/WCL-Robotics/TaskGrasp-Pro.task_book_redThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 46,
"total_frames": 12634,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:46"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_book_red.task_book_leavesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 38,
"total_frames": 8664,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:38"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_book_leaves.object_task_suiteThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "mobile_aloha",
"total_episodes": 301,
"total_frames": 77500,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/object_task_suite.task_all_origtask_pull_leavesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 9649,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_pull_leaves.autocad-bench-tasks
AutoCAD-Bench Tasks
This dataset contains the 50 AutoCAD tasks used for model evaluation in
AutoCAD-Bench.
Contents
50 total evaluation tasks
21 dimensioned 2D drafting tasks
29 3D modeling and presentation tasks
millimetres as the task unit
one reference PNG and one scorer-side gold DWG per task
the exact 2D and 3D task prompts used by the benchmark
Each task is stored under data/task-NNN/:
data/task-001/
├── reference.png
└── gold.dwg
tasks.jsonl records the… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/autocad-bench-tasks.PENGWIN_Task1
PENGWIN Task 1 — Pelvic Fracture Segmentation on CT
The CT task of the PENGWIN 2024 challenge (PElvic bone fraGment (WIN)dow,
MICCAI 2024): segment the sacrum, left hipbone and right hipbone, and the
individual fracture fragments of each, in preoperative pelvic trauma CT.
This is an instance segmentation task, not a 3-class semantic one — the
label value identifies which fragment of which bone, and the fragment count
varies per case.
What this mirror contains — read… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/PENGWIN_Task1.task_pull_redThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 21,
"total_frames": 4548,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/task_pull_red.realworld_task820UVT-Explanatory-based-Vision-Tasks
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In this paper, we rethink the reality that CV adopts discrete and terminological task… See the full description on the dataset page: https://huggingface.co/datasets/axxkaya/UVT-Explanatory-based-Vision-Tasks.GUIrilla-Task
GUIrilla-Task
Ground-truth Click & Type actions for macOS screenshots
Dataset Summary
GUIrilla-Task pairs real macOS screenshots with free-form natural-language instructions and precise GUI actions.
Every sample asks an agent either to:
Click a specific on-screen element, or
Type a given text into an input field.
Targets are labelled with bounding-box geometry, enabling exact evaluation of visual-language grounding models.
Data were gathered automatically by… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/GUIrilla-Task.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.sciclaimeval-shared-task
SciClaimEval Shared Task: All information is available at sciclaimeval.github.io
Evaluation scripts & examples: github.com/SciClaimEval/sciclaimeval-shared-task
More Information: paper
Version Info
Please use the latest version, v1.1.
Changes from v1.0 to v1.1
Compared with v1.0, v1.1 includes the following changes.
Removed Samples
The following 20 samples have been removed:
val_tab_1594
val_tab_0067… See the full description on the dataset page: https://huggingface.co/datasets/alabnii/sciclaimeval-shared-task.
