datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLS-Bench-Tasks
MLS-Bench Tasks
MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.mind2web_multimodal_test_task
Dataset Card for Multimodal Mind2Web "Cross-Task" Test Split
Note: This dataset is the test split of the Cross-Task dataset introduced in the paper.
This is a FiftyOne dataset with 1338 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_task.osworld_tasks_filestt-task-tensor-batches
TikTok TechJam Task 5 Tensor Batches
This private dataset contains local tensor batches prepared for TikTok TechJam
2026 Track 5: robust AI-generated image detection under real-world
transformations.
Each .npz file stores one original image and its transformed variants:
images: [20, 3, H, W]
label: bool scalar, False=real, True=AI-generated
transform_names: clean + 19 transform names
image_paths: relative paths used when the tensor batch was built
layout: images are [B, C, H… See the full description on the dataset page: https://huggingface.co/datasets/Miyano-siho/tt-task-tensor-batches.geoguesser-tasks
GeoGuesser Task Splits
Task indexes for the GeoGuesser OpenEnv environment.
Each line is one episode: an ordered list of panorama frames with coordinates,
headings and capture dates, plus the sequence and contributor it came from.
Split
Tasks
Countries
Frames
Fully mirrored
eval
200
73
4673
200/200
train
3452
130
80179
3448/3452
What a task is
These files carry metadata only, not imagery. Every frame's coordinates,
heading and capture date are… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/geoguesser-tasks.Recruitment-Task-3
DeepWeeds - AI-MED AGH convenience mirror
This is a convenience mirror of the official DeepWeeds image archive and the
upstream annotations pinned to a specific commit. original/images.zip is
preserved unchanged; images are not extracted or duplicated here. models.zip
from the source authors is deliberately not mirrored.
Dataset facts
17,509 in-situ images from Queensland, Australia.
Nine classes: eight weed species plus Negative.
The authors publish five folds… See the full description on the dataset page: https://huggingface.co/datasets/AI-MED-AGH/Recruitment-Task-3.ArGuard-Task1
ArGuard – Track A: Arabic Hateful Memes
This repository hosts the official dataset for Track A of the
ArGuard shared task: multimodal hateful-meme detection in Arabic.
Each instance is an Arabic meme (image + OCR-extracted overlaid text)
manually annotated for hatefulness and fine-grained sub-types.
Content warning. The dataset contains text and imagery that is
offensive, discriminatory, or otherwise harmful by design. Handle
with care.
Track A subtasks
Given a… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ArGuard-Task1.Recruitment-Task-2A
BloodMNIST - AI-MED AGH convenience mirror
This repository is an AI-MED AGH convenience mirror of the unchanged official
BloodMNIST NPZ distribution from the MedMNIST project.
The original bloodmnist.npz is preserved byte-for-byte in original/.
Dataset facts
28x28 RGB images with eight blood-cell classes.
Official splits: 11,959 training images, 1,712 validation images, and 3,421 test images.
Classes:
basophil
eosinophil
erythroblast
immature granulocytes… See the full description on the dataset page: https://huggingface.co/datasets/AI-MED-AGH/Recruitment-Task-2A.ProbeScout-tasks
ProbeScout Main17 task packages
Use with ProbeScout source and setup instructions.
Download only the dataset(s) you need, and extract each ZIP into the code repository root.
The ZIP paths already include dataset/ and visual_analytics/.
File
Tasks
Compressed size
cars.zip
7
93.74 MB
hico.zip
8
320.29 MB
celeba.zip
2
439.43 MB
Each package includes task definitions, attributes, query image IDs, ordered
records, original VQA source/fit/Validation labels, fixed… See the full description on the dataset page: https://huggingface.co/datasets/Ian100/ProbeScout-tasks.video-dataset-task-02
Video Dataset - task-02
Dataset Description
This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips.
Dataset Structure
frames/ — extracted frames (first frame from each segment)
segments/ — video clips for each annotation interval
annotations/ — original JSON annotation
transcriptions/ — transcription files (full_transcription.txt + per segment)
dataset.csv — mapping… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-task-02.video-dataset-task02
Video Dataset - task02
Dataset Description
This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips.
Dataset Structure
frames/ — extracted frames (first frame from each segment)
segments/ — video clips for each annotation interval
annotations/ — original JSON annotation
transcriptions/ — transcription files (full_transcription.txt + per segment)
dataset.csv — mapping… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-task02.FLARE-Task4-CT-FM
MICCAI FLARE25
Task 4: Foundation Models for 3D CT and MRI Scans (Homepage)
This is the official dataset for CT image foundation model development. We provide 10,000+ CT scans for model pretraining.
Downstream tasks include:
Abdominal disease classification
Abdominal lesion segmentation
Abdominal organ segmentation
Lung lesion segmentation
Dataset
Dataset Name
Task
Metric
Source
License
Abdominal Disease Classification
multi-label… See the full description on the dataset page: https://huggingface.co/datasets/FLARE-MedFM/FLARE-Task4-CT-FM.from-desktop-task01
Video Dataset - task01
Dataset Description
This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips.
Dataset Structure
frames/ — extracted frames (first frame from each segment)
segments/ — video clips for each annotation interval
annotations/ — original JSON annotation
transcriptions/ — transcription files (full_transcription.txt + per segment)
dataset.csv — mapping… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/from-desktop-task01.FLARE-Task4-MRI-FM
MICCAI FLARE25
Task 4: Foundation Models for 3D CT and MRI Scans (Homepage)
This is the official dataset for MRI image foundation model development. We provide 10,000+ MRI scans for model pretraining.
Downstream tasks include:
Liver tumor segmentation: ATLAS23_liver_tumor_seg
Cardiac tissue and pathology segmentation: EMIDEC_heart_seg_and_class
Heart myocardial status classification (binary label: pathological vs normal): EMIDEC_heart_seg_and_class
Autism Diagnosis:… See the full description on the dataset page: https://huggingface.co/datasets/FLARE-MedFM/FLARE-Task4-MRI-FM.
