datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.Heliconius-Collection_Cambridge-Butterfly
Dataset Card for Heliconius Collection (Cambridge Butterfly)
Dataset Description
Dataset Summary
Subset of the collection records from Chris Jiggins' research group at the University of Cambridge, collection covers nearly 20 years of field studies.
This subset contains approximately 36,189 RGB images of 11,962 specimens (29,134 images of 10,086 specimens across all Heliconius). Many records have both images and locality data.
Most images were… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Heliconius-Collection_Cambridge-Butterfly.CrossViewUrbanTrafficDataset
Cross-View Urban Traffic Dataset
Dataset Summary
The Cross-View Urban Traffic Dataset (CVUTD) is a benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections in Regensburg, Germany.
The dataset is designed to support two linked tasks:
Cross-view identity matching between street-view and drone-view object tracks
Ego-to-BEV prediction using aerial supervision… See the full description on the dataset page: https://huggingface.co/datasets/prakharbh/CrossViewUrbanTrafficDataset.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.conflux-chest-ct
CONFLUX Chest-CT
200,000 synthetic 3D chest CT volumes with structured abnormality and demographic labels, generated by CONFLUX.
Released with the paper CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training.
Paper (arXiv) •
Model •
Code — coming soon
About
CONFLUX is a conditional 3D latent generative model for chest CT: a VAE tokenizer
compresses each volume into a compact 16-channel latent, a… See the full description on the dataset page: https://huggingface.co/datasets/gevaertlab/conflux-chest-ct.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.bigearthnet
BigEarthNet - HDF5 version
This repository contains an export of the existing BigEarthNet dataset in HDF5 format. All Sentinel-2 acquisitions are exported according to TorchGeo's dataset (120x120 pixels resolution).
Sentinel-1 is not contained in this repository for the moment.
CSV files contain for each satellite acquisition the corresponding HDF5 file and the index.
A PyTorch dataset class which can be used to iterate over this dataset can be found here, as well as the script used… See the full description on the dataset page: https://huggingface.co/datasets/lc-col/bigearthnet.Ransomware_PE_Header_Feature_Dataset
Dataset Card for Ransomware PE Header Feature Dataset
Dataset Description
Dataset Summary
This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header.
Supported Tasks
Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.CleanPatrick
CleanPatrick: A Benchmark for Data Cleaning
Welcome to CleanPatrick, the first large-scale benchmark designed for data cleaning in the image domain.
Built on the Fitzpatrick17k dermatology dataset, CleanPatrick is a dataset for measuring the performance in detecting three major data quality issues:
off-topic samples, near-duplicates, and label errors.
Overview
CleanPatrick consists of dermatological images annotated with over 500,000 binary labels across three data… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Dermatology/CleanPatrick.phenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.cnn-based-drowsiness-detection-data
CNN-Based Drowsiness Detection - Dataset
Preprocessed, auto-labeled face-crop images used to train the model in
notgoodkeeper/cnn-based-drowsiness-detection.
Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection
Collection
Frames were captured from a webcam, then run through:
Haar Cascade face detection -> crop + pad + resize to 412x412
MediaPipe Selfie Segmentation -> background replaced with white
CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.typhoon-intensity-classification
Typhoon - Image Classification Dataset
This dataset comes from PTIT AI Challenge and is organized for a multi-class image classification task focusing on tropical cyclone (typhoon) intensity estimation.
Dataset Structure
The directory structure is organized as follows:
train/
├── images/
│ ├── image1.jpg
│ └── ...
└── annotations.csv (only present in the train folder)
The public_test and private_test sets are used to evaluate and score the… See the full description on the dataset page: https://huggingface.co/datasets/star092304/typhoon-intensity-classification.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/Rosalia1212/cbis-ddsm-r.CulturalBiases-2025Preprint : [https://arxiv.org/pdf/2505.14729?]
Crop-Disease-Image-Eval-Synthetic
Crop, Category, Disease and Pest Test Set
11,057 smallholder-farmer photographs sent to FarmerChat from Ethiopia, India, Kenya and Nigeria, each
labelled with the crop, whether the problem is a disease or a pest, and which one. This is the held-out
test split of a four-head classification benchmark, restricted to the rows whose labels came from an
independent model council rather than from the production vendor.
Why 11,057 and not 16,275
The full held-out split is… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Crop-Disease-Image-Eval-Synthetic.CrossViewUrbanTrafficSubsetDataset
Cross-View Urban Traffic Subset Dataset
This repository provides a representative sample of the Cross-View Urban Traffic Dataset (CVUTD) for reviewer inspection and qualitative verification.
The subset is intended to let reviewers:
inspect the raw and annotated data format,
verify annotation quality,
understand the cross-view correspondence structure,
and assess the benchmark design without downloading the full dataset.
The full dataset is hosted separately and is available at… See the full description on the dataset page: https://huggingface.co/datasets/prakharbh/CrossViewUrbanTrafficSubsetDataset.testing-goldstandard-cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/jrw2989/testing-goldstandard-cuthill.japanese-image-classification-evaluation-dataset
recruit-jp/japanese-image-classification-evaluation-dataset
Overview
Developed by: Recruit Co., Ltd.
Dataset type: Image Classification
Language(s): Japanese
LICENSE: CC-BY-4.0
More details are described in our tech blog post.
日本語CLIP学習済みモデルとその評価用データセットの公開
Dataset Details
This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks.
jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.OsteosarcomaTumorAssessment
Osteosarcoma data from UT Southwestern/UT Dallas for Viable and Necrotic Tumor Assessment (Osteosarcoma-Tumor-Assessment)
Unofficial fork.
Folder Structure
ML_Features_1144.csv # Contains 1144 rows for all the image tiles and 69 columns for filename, classification, and 65 machine learning features.
OsteosarcomaTumorAssessment.tar.zst
|-- Osteosarcoma-UT.sums
|-- Training-Set-1 # 11 folders with 547 images. Each folder contains 48~50 image tiles and 1 csv for… See the full description on the dataset page: https://huggingface.co/datasets/CAIR-M3LLM/OsteosarcomaTumorAssessment.ad-creative-quality-human-vs-llm
Human Expert vs LLM Judge: Facebook Ad Creative Quality
500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.
The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.mlcd-mteb-cifar-eval
MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation
Evaluation results accompanying the MTEB integration of two MLCD image encoders
(PR #5406, resolving
issue #2571).
Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the
reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100
image-classification tasks alongside size-matched OpenAI CLIP baselines.
What was measured
Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/helloerikaaa/cbis-ddsm-r.ImageNet-CJ
JPEG Re-encoding Confound Control Dataset
A controlled-experiment dataset that isolates one acknowledged-but-unmeasured confound in
ImageNet-C. Hendrycks & Dietterich (Benchmarking Neural Network Robustness to Common
Corruptions and Perturbations, ICLR 2019, arXiv:1903.12261)
save every corrupted image as a lightly compressed JPEG. The benchmark therefore never measures a
corruption c applied to an image x in isolation — it measures JPEG(c(x)). This dataset lets
you quantify how… See the full description on the dataset page: https://huggingface.co/datasets/atharvadagaonkar/ImageNet-CJ.humanoid-basic-actions-dataset-v1
Humanoid Basic Actions Dataset v1
Synthetic dataset for humanoid robot training simulation.
Description
This dataset contains labeled humanoid robot action images for basic movement recognition tasks.
Classes
walk
run
sit
stand
wave
pick_object
turn_left
turn_right
Structure
dataset/
├── train/
├── validation/
Each folder contains subfolders named after action labels.
Format
Image Classification (RGB Images 224x224)
Total… See the full description on the dataset page: https://huggingface.co/datasets/Caplin43/humanoid-basic-actions-dataset-v1.CleanPatrick
CleanPatrick: A Benchmark for Data Cleaning
Welcome to CleanPatrick, the first large-scale benchmark designed for data cleaning in the image domain.
Built on the Fitzpatrick17k dermatology dataset, CleanPatrick is a dataset for measuring the performance in detecting three major data quality issues:
off-topic samples, near-duplicates, and label errors.
Overview
CleanPatrick consists of dermatological images annotated with over 500,000 binary labels across three… See the full description on the dataset page: https://huggingface.co/datasets/akshanshmittal12/CleanPatrick.Fashion-MNIST-CSVThis dataset is a direct copy of Fashion-MNIST, originally published by Zalando Research on Kaggle https://www.kaggle.com/datasets/zalando-research/fashionmnist.
Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. Zalando intends Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for… See the full description on the dataset page: https://huggingface.co/datasets/vincent-espitalier/Fashion-MNIST-CSV.chinesechatdeaf-isl-gesturesChatDEAF ISL Dataset-Initial README and metadata update
license: cc-by-4.0
pretty_name: ChatDEAF ISL Gesture Dataset
tags:
sign-language
accessibility
isl
gesture-recognition
multimodal
chatdeaf
task_categories:
image-classification
size_categories:
10<n<100
language:
- en
tags:
sign-language
isl
international-sign
gestures
accessibility
chatdeaf
image-classification
visual-language
{
"belt": "Gesture for 'belt' in ISL.",
"glasses": "Gesture for 'glasses' in… See the full description on the dataset page: https://huggingface.co/datasets/yasodeafs/chatdeaf-isl-gestures.dataset-openmoji
Dataset OpenMoji
Creator: https://www.kaggle.com/krayc81This is base on https://openmoji.org/License https://creativecommons.org/licenses/by-sa/4.0
Files:
README.md this :)
data.csv containing all data see bellow description
openmoji folder containing the image files
The data.csv contains:
idx the character as int
character text representation
bytes representation
hex representation (replace Ox with U+ for unicode)
description of the emoji
path_black path to the bw image… See the full description on the dataset page: https://huggingface.co/datasets/Kray-C/dataset-openmoji.K-MNIST-CSV
Kuzushiji-MNIST
This dataset is a direct CSV conversion of Kuzushiji-MNIST, originally sourced from the GitHub repository https://github.com/rois-codh/kmnist.
Kuzushiji-MNIST is a drop-in replacement for the MNIST dataset (28x28 grayscale, 70,000 images).
