datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multiple-sclerosis-dataset
Multiple Sclerosis Dataset, Brain MRI Object Detection & Segmentation Dataset
The dataset consists of .dcm files containing MRI scans of the brain of the person with a multiple sclerosis. The images are labeled by the doctors and accompanied by report in PDF-format.
The dataset includes 13 studies, made from the different angles which provide a comprehensive understanding of a multiple sclerosis as a condition.
MRI study angles in the dataset
💴 For… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/multiple-sclerosis-dataset.mind2web_multimodal_test_domain
Dataset Card for "Cross-Domain" Test Split in Multimodal Mind2Web
Note: This dataset is the test split of the Cross-Domain dataset introduced in the paper.
This is a FiftyOne dataset with 4050 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_domain.mind2web_multimodal_test_task
Dataset Card for Multimodal Mind2Web "Cross-Task" Test Split
Note: This dataset is the test split of the Cross-Task dataset introduced in the paper.
This is a FiftyOne dataset with 1338 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_task.mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.hard-intersection-multimodal-sample
Hard Intersection Multimodal Samples
Release Notes
Release
Description
v1.0.0
Initial public release.
v1.1.0
Added Unreal Engine assets.Fixed issues in the OpenDRIVE map data.Updated the README to improve documentation and usability.
Dataset Summary
Hard Intersection Multimodal Samples is a curated multimodal dataset of accident-prone urban intersection in Japan for autonomous driving research and development.It provides… See the full description on the dataset page: https://huggingface.co/datasets/dynamic-maps/hard-intersection-multimodal-sample.Multimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans
Multimodal Chest X ray dataset for Normal and Bacterial Pneumonia in Africans | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: imagefolder - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Multimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.pd-voice-full-multimodal-dataset
Parkinson Voice — Full Multimodal Dataset
Complete Parkinson’s vs healthy voice package for classification and explainable Mel reasoning research (EDGE).
Not Mel-only: raw audio, 10 visual modalities, feature CSVs, plus Gemma reasoning traces for Mel.
Contents
Path
Description
audio/
Waveform clips (healthy / parkinsons), 1134 files
images/mel/
Mel spectrograms
images/spectrogram/
Linear spectrograms
images/mfcc/
MFCC maps
images/delta_mfcc/… See the full description on the dataset page: https://huggingface.co/datasets/mdimamhosen/pd-voice-full-multimodal-dataset.color-multi-fractal-db-1k
Dataset Card for Color Multi Fractal DB 1k
This is a pre-generated 1k classes, 1M images colored-multi-fractal-images dataset based on Improving Fractal Pre-training by Connor Anderson et al. and Multi-Fractal-Dataset by FYSignate1009.
We have changed some fractal parameters so that our ViT pretraining can converge. Modified parameters can be found on this repo.
You can pretrain vision transformers without worrying about dataset licensing for commercial use.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/color-multi-fractal-db-1k.multicare-subset
MultiCaRe (subset)
A subset of the MultiCaRe open-source
clinical case dataset, re-packaged for the Hugging Face Hub with images
embedded alongside captions, clinical case text, and article metadata.
Source: https://zenodo.org/records/20416562 (DOI: 10.5281/zenodo.20416562)
Original paper: https://doi.org/10.3390/data10080123
License: the dataset as a whole is CC BY-NC-SA 4.0. Individual rows may
carry a less restrictive license (CC BY, CC BY-NC, CC0) recorded in the
license… See the full description on the dataset page: https://huggingface.co/datasets/MohamedAhmedAE/multicare-subset.Uddessho-Bangla-Multimodal-Intent-Classification
📊 Uddessho Dataset — Multimodal Author Intent Classification
Uddessho (meaning "Intent" in English) is a multimodal dataset created for author intent classification in the low-resource Bangla language.It contains 3,048 social media posts (text + images) labeled into six distinct intent types.
🏷️ Intent Categories & Label Mapping
Label ID
Class Name
0
Advocative
1
Controversial
2
Exhibitionist
3
Expressive
4
Informative
5
Promotive
📂… See the full description on the dataset page: https://huggingface.co/datasets/Mukaffi28/Uddessho-Bangla-Multimodal-Intent-Classification.multi-label-food-recognition
Multi-Label Food Recognition Dataset
This is a multi-label food recognition dataset generated from single-class food images.
Each image contains 2-5 different food items composited together using natural composition methods.
Dataset Details
Total Images: 13,000
Training Images: 10,400 (80%)
Validation Images: 2,600 (20%)
Number of Classes: 90
Labels per Image: 2-5 labels
Image Format: RGB, 512x512 pixels
File Format: Parquet
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.REFUGE-MultiRater
REFUGE
REFUGE Challenge provides a data set of 1200 fundus images with ground truth segmentations and clinical glaucoma labels, currently the largest existing one.
This dataset supplied multi-rater annotations of REFUGE Challenge Dataset. The challenge dataset releases majority vote (with some modifications) results of seven independent
annotations. We release the scource seven annotations here.
Cite
@article{fang2022refuge2,
title={REFUGE2 Challenge: Treasure for… See the full description on the dataset page: https://huggingface.co/datasets/realslimman/REFUGE-MultiRater.xraydar-multimodal
X-Raydar Multimodal Chest X-Ray Dataset
A multimodal dataset of 979 chest X-ray examinations, each with:
Chest X-ray image (full-resolution PNG, anonymised)
Consensus image-level labels (37 radiological findings, agreed by two expert radiologists with adjudication)
Bounding box annotations on the image from each annotator independently (localising each finding)
Original radiology report text
Report span annotations (token-level labels across 45 categories)
This dataset combines… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/xraydar-multimodal.multicare-images
MultiCaRe: Open-Source Clinical Case Dataset
MultiCaRe is an open-source, multimodal clinical case dataset built from the PubMed Central Open Access (OA) Case Report articles. It aggregates de-identified, open-access case narratives, figure images, captions, and rich article metadata across diverse specialties (radiology, pathology, surgery, ophthalmology, etc.). The data is normalized so images, cases, and articles can be joined via stable IDs.
Source and process: OA case reports… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/multicare-images.AID_MultiLabel
Dataset Card for "AID_MultiLabel"
Licensing Information
CC0: Public Domain
Citation Information
Imagery:
AID: A benchmark data set for performance evaluation of aerial scene classification
Multilabels:
Relation Network for Multi-label Aerial Image Classification
@article{xia2017aid,
title = {AID: A benchmark data set for performance evaluation of aerial scene classification},
author = {Xia, Gui-Song and Hu, Jingwen and Hu, Fan and Shi, Baoguang… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/AID_MultiLabel.NSFW-MultiDomain-Classification
NSFW_MultiDomain
The NSFW_MultiDomain dataset is a curated image classification dataset focused on multi-domain adult content recognition. It consists of 5 distinct categories aimed at facilitating the development of robust NSFW (Not Safe For Work) image classification models. This dataset enables training and benchmarking of models that can distinguish between subtle variations in explicit and non-explicit content across artistic, animated, and real-world imagery.… See the full description on the dataset page: https://huggingface.co/datasets/strangerguardhf/NSFW-MultiDomain-Classification.cocoa_agroforestry_multispectral
Cocoa Agroforestry Multispectral
An unlabeled image dataset of Cocoa Agroforestry Multispectral. The dataset contains 1,272 images with no classification, segmentation, or bounding-box annotations.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{lammoglia2024high,
title={High-resolution multispectral and RGB dataset from UAV surveys of ten cocoa agroforestry typologies in Côte d'Ivoire}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/cocoa_agroforestry_multispectral.coastal-multitask-380
Coastal & Rural Bangladesh — Multi-Task Visual Dataset
379 field photographs (JPEG, native resolution as shot — see classification/metadata.csv for per-image width/height) collected on foot along the Bakkhali river embankment and surrounding villages/farmland near Cox's Bazar, Bangladesh, structured into three ML-task "levels": classification, semantic segmentation, and change detection.
Source: huggingface data 06 (Golam Rob / Tawhid Enterprise photo collection).… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/coastal-multitask-380.multi-material-fingerprint-spoofing
Fingerprint Spoofing
The dataset contains over 4,000+ photos from 100 people, consisting of fingerprints images and spoofing attacks created using various spoofing materials such as alginate, plasticine, and silicone. It serves as essential training data for biometric systems focused on fingerprint recognition and spoof detection.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in biometric security and spoof detection… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/multi-material-fingerprint-spoofing.ChinaHeritaQA
Images
This folder contains visual data for the ChinaHeritaQA benchmark: https://arxiv.org/abs/2606.08959
Contents
Folder
Description
Image_data/
Chinese UNESCO World Heritage Site images (2,279 images from 51 sites)
worlds_data/
Non-Chinese World Heritage Site images (133 images from 23 sites)
Overview
The image dataset includes a comprehensive collection of photographs from both Chinese and international UNESCO World Heritage… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/ChinaHeritaQA.tekno21-brain-stroke-dataset-multi
Dataset Card for BTX24/tekno21-brain-stroke-dataset-multi
🔗 Dataset Sources
Dataset Source: TEKNOFEST-2021 Stroke Dataset
Kaggle: İnme Veri Seti (Stroke Dataset)
Sağlık Bakanlığı Açık Veri Portalı: İnme Veri Seti
Dataset Structure
Format: PNG
Total Images: 7,369
Categories:
hemorajik/ (Hemorrhagic stroke images)
iskemik/ (Ischemic stroke images)
normal/ (Non-stroke images)
The dataset is structured in a folder-based format where images are grouped into… See the full description on the dataset page: https://huggingface.co/datasets/BTX24/tekno21-brain-stroke-dataset-multi.multimodal-shapes-subset
Dataset Card for multimodal_shapes_subset
This is a grouped multimodal FiftyOne dataset with 2000 samples, each consisting of an rgb image and LiDAR data of objects.
The samples are labeled as cube or sphere. The classes are perfectly balanced (1000 cubes, 1000 spheres).
This dataset is used in this project on github: https://github.com/MatthiasCr/Computer-Vision-Assignment-2/tree/main.
There it is also explained how to use this dataset, visualize it in fiftyone, and convert… See the full description on the dataset page: https://huggingface.co/datasets/MatthiasCr/multimodal-shapes-subset.wheat_mosaic_classification_multispectral
Wheat Mosaic Classification Multispectral
This dataset provides real multispectral images of wheat plants in field environments across South Africa, collected during August and October 2023. Captured using a specially adapted Canon EOS 800D DSLR, the images focus on early detection of Wheat Stripe Mosaic Virus, depicting plants at various stages of disease progression. The dataset contains 424 images across 2 classes: diseased, early.Images per class:
diseased: 252
early: 172… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/wheat_mosaic_classification_multispectral.2026-24679-HW1-Multimodal-Original
Straight-member torque: image and structured statics data
eandujar/2026-24679-HW1-Multimodal-Original
100 synthetic planar-statics cases containing a rendered diagram and
structured numerical/categorical features describing the member geometry,
supports, and applied loads.
The prediction task has two outputs:
Torque direction — clockwise or counterclockwise.
Torque magnitude — absolute moment about the pin in N m.
This therefore supports both classification and regression… See the full description on the dataset page: https://huggingface.co/datasets/eandujar/2026-24679-HW1-Multimodal-Original.Multilabel-Portrait-18K
Multilabel-Portrait-18K
Multilabel-Portrait-18K is a multi-label portrait classification dataset designed to analyze and categorize different styles of portrait images. It supports classification into the following four portrait types:
0 — Anime Portrait
1 — Cartoon Portrait
2 — Real Portrait
3 — Sketch Portrait
This dataset is ideal for training and evaluating machine learning models in the domain of portrait-style classification. The goal is to enable accurate recognition… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Multilabel-Portrait-18K.hw1-paper-triage-multimodal
HW1 Multimodal Paper Triage
Purpose
This dataset supports a classroom exercise in assembling and augmenting multimodal data for a personalized research-paper triage system.
Composition and splits
The dataset begins with 100 original multimodal samples. The original samples were split before augmentation using a fixed random seed and stratification by the binary target.
train: 10,070 samples consisting of 70 training originals and 10,000 augmented… See the full description on the dataset page: https://huggingface.co/datasets/ishaanamahajan/hw1-paper-triage-multimodal.2026-24679-HW1-Multimodal
Straight-member torque: image and structured statics data
eandujar/2026-24679-HW1-Multimodal
100 synthetic planar-statics cases containing a rendered diagram and
structured numerical/categorical features describing the member geometry,
supports, and applied loads.
The prediction task has two outputs:
Torque direction — clockwise or counterclockwise.
Torque magnitude — absolute moment about the pin in N m.
This therefore supports both classification and regression experiments.… See the full description on the dataset page: https://huggingface.co/datasets/eandujar/2026-24679-HW1-Multimodal.indonesian-gambling-multimodal-dataset
Indonesian Online Gambling Promotion Multimodal Dataset
This dataset is intended for machine learning research on detecting and filtering online gambling promotions on Indonesian social media. It is a research artifact for content moderation and harmful content detection, not a gambling website or promotional resource.
Research Context
This dataset is created for academic research in harmful content detection and multimodal machine learning. It is intended to… See the full description on the dataset page: https://huggingface.co/datasets/0xRafie/indonesian-gambling-multimodal-dataset.
