datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gastric-X
Gastric-X
Multi-phase abdominal CT cohort paired with structured laboratory panels
and free-text radiology reports, in proficient medical English with
the original Simplified Chinese preserved alongside.
Changelog
2026-06-26
Added per-phase organ masks (<phase>_organ_mask.nii.gz) — CADS
multi-organ segmentation on each phase's CT grid (e.g. label 6 = stomach);
all 4897 phases.
Added per-phase gastric tumor masks (<phase>_tumor_mask.nii.gz,
binary) — a patient's… See the full description on the dataset page: https://huggingface.co/datasets/HaoChen2/Gastric-X.clearwrist-pediatric-wrist-xrayClearWrist: Pediatric Wrist Fracture X-Ray Dataset
20,327 labeled pediatric wrist radiographs, rebuilt from GRAZPEDWRI-DX with clean patient-level splits, verified fracture ground truth, and YOLO-style bounding box annotations.
Overview
This dataset packages the full GRAZPEDWRI-DX corpus, 20,327 pediatric wrist radiographs from 6,091 patients treated at the Department for Pediatric Surgery of the University Hospital Graz between… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/clearwrist-pediatric-wrist-xray.MCD-2.6m
MCD-2.6m
MCD-2.6m is a collection of 2,604,450 agricultural and plant images distributed in 49 Parquet shards. It combines images of multiple crops collected across several institutions and field-imaging projects.
The release contains one train split. Images are embedded in the Parquet files and can be decoded directly with the Hugging Face datasets library.
Dataset Structure
Each example contains exactly three columns:
Column
Type
Description
row_id… See the full description on the dataset page: https://huggingface.co/datasets/XIANG-Shuai/MCD-2.6m.chest-xray-classification
Dataset Labels
['NORMAL', 'PNEUMONIA']
Number of Images
{'train': 4077, 'test': 582, 'valid': 1165}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("keremberke/chest-xray-classification", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/mohamed-traore-2ekkp/chest-x-rays-qjmia/dataset/2
Citation… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/chest-xray-classification.xAI_Aurora_t2i_human_preferences
Rapidata Aurora Preference
This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Aurora across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/xAI_Aurora_t2i_human_preferences.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-1k-wds.NIH-Chest-XRay-Federated
NIH Chest X-ray Federated Learning Dataset
Federated learning splits designed for the [Cold Start:] Distributed AI Hack Berlin 2025.
The dataset is based on the NIH Chest X-ray14 dataset, which contains ~112,000 X-ray images from 30,805 unique patients, and models a federated learning scenario with non-IID characteristics across three hospitals, plus an out-of-distribution test set.
Dataset Description
The data was partitioned using a scoring algorithm that creates… See the full description on the dataset page: https://huggingface.co/datasets/exalsius/NIH-Chest-XRay-Federated.chest-xray-14-320
NIH Chest X-ray14 - 320x320 Processed for CheXVision
Project Resources
GitHub repository
Presentation deck
Live demo
Scratch model
DenseNet model
This dataset repackages the raw NIH Chest X-ray14 source dataset from
alkzar90/NIH-Chest-X-ray-dataset
into a data-only Parquet dataset for the CheXVision project.
Dataset Summary
Source format: 12 ZIP archives of original chest X-ray images plus CSV manifests
Output format: data-only Parquet shards under data/… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/chest-xray-14-320.chest-xray-14
NIH Chest X-ray14 — Processed for CheXVision
This dataset wraps the NIH Chest X-ray14 dataset, preprocessed for the CheXVision project.
Labels
Label
Count
Prevalence
Infiltration
19,894
17.7%
Effusion
13,317
11.9%
Atelectasis
11,559
10.3%
Nodule
6,331
5.6%
Mass
5,782
5.2%
Pneumothorax
5,302
4.7%
Consolidation
4,667
4.2%
Pleural_Thickening
3,385
3.0%
Cardiomegaly
2,776
2.5%
Emphysema
2,516
2.2%
Edema
2,303
2.1%
Fibrosis
1,686
1.5%… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/chest-xray-14.REOBench
Folder/File Descriptions
AID/AID_train.zip: Contains all AID images in the training set.
AID/AID_test.zip: Contains images in the test set under perturbation.
AID/AID_JSON/: Contains JSON files for zero-shot evaluation of LLM-based models.
Potsdam/Potsdam_Images_trian.zip: Contains all Potsdam images in the training set.
Potsdam/Potsdam_Anns_trian.zip: Contains annotations for images in the training set.
Potsdam/Potsdam_Images_test.zip: Contains Potsdam test images under… See the full description on the dataset page: https://huggingface.co/datasets/xiang709/REOBench.chest-xray-14
NIH Chest X-ray14 — Processed for CheXVision
This dataset wraps the NIH Chest X-ray14 dataset, preprocessed for the CheXVision project.
Labels
Label
Count
Prevalence
Infiltration
19,894
17.7%
Effusion
13,317
11.9%
Atelectasis
11,559
10.3%
Nodule
6,331
5.6%
Mass
5,782
5.2%
Pneumothorax
5,302
4.7%
Consolidation
4,667
4.2%
Pleural_Thickening
3,385
3.0%
Cardiomegaly
2,776
2.5%
Emphysema
2,516
2.2%
Edema
2,303
2.1%
Fibrosis
1,686
1.5%… See the full description on the dataset page: https://huggingface.co/datasets/Sharon2105/chest-xray-14.XJTU-perception
Roles
Roles: perception view of XJTU — annot is the source label (inner_race / normal / outer_race), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the repo ships a bearing's vibration in four image encodings as four equal-sized configs — reshaped (consecutive samples arranged as the rows of a grayscale square), scalogram (a continuous-wavelet time-scale view), spectrogram (a short-time Fourier transform) and waveform… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/XJTU-perception.XJTU
Roles
Roles: canon repo — annot is the source label, kept machine-parseable as the gold for verification and reward parsing; there is no filled reasoning column and this repo is not itself a training view. Derived repos each state their own regime on their own card.
XJTU-SY — fault classification from the envelope spectrum (reasoning track)
Third signal dataset in the AI4Manufacturing FORGE corpus (Category C, task T-C1), from 15 accelerated run-to-failure tests.… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/XJTU.DiagramSketch
DiagramSketch
This repository contains the full dataset archive and a smaller sample archive.
The sample is provided so users can inspect the directory structure, test data
loading code, and review the metadata format before downloading the full
dataset.
Files
dataset.tar.gz: full dataset archive.
dataset_sample.tar.gz: small sample archive hosted alongside the full
archive.
dataset_sample_manifest.jsonl: lightweight manifest used by the Hugging
Face Dataset… See the full description on the dataset page: https://huggingface.co/datasets/xinjiangyu/DiagramSketch.xai-attack-detection-cifar10
XAI Attack Detection — CIFAR-10 PGD
This private research dataset contains balanced, paired clean and adversarial images for
studying whether an attack can be detected from a classifier explanation map.
Dataset construction
The source is the CIFAR-10 test split. A fine-tuned OpenCLIP ViT-B/16 classifies each
clean image. Clean-correct examples are attacked with untargeted L-infinity PGD using
epsilon 8/255, step size 2/255, 10 steps, and deterministic random… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-cifar10.chest-xray-classification
Dataset Labels
['PNEUMONIA', 'NORMAL']
Number of Images
{'test': 582, 'valid': 1165, 'train': 12230}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("trpakov/chest-xray-classification", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/mohamed-traore-2ekkp/chest-x-rays-qjmia/dataset/3
Citation
License… See the full description on the dataset page: https://huggingface.co/datasets/trpakov/chest-xray-classification.SUN-R-D-T
📚 SUN-R-D-T
SUN-R-D-T is a multi-view/modal benchmark built on top of SUN RGB-D.Each scene is represented by:
a RGB image
a Depth map
a MLLM-generated caption (text view)
a 19-way scene label (train/test split follows SUN RGB-D)
The text descriptions are generated automatically by Qwen3-VL-32B-Instruct with a carefully designed prompt, aiming to capture salient scene content while avoiding label leakage and hallucinated details.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/XLearning-SCU/SUN-R-D-T.chest-xray-14
NIH Chest X-ray14 Dataset
Dataset Description
This dataset contains 112120 chest X-ray images with multiple disease labels per image.
Labels
Atelectasis, Cardiomegaly, Consolidation, Edema, Effusion, Emphysema, Fibrosis, Hernia, Infiltration, Mass, No Finding, Nodule, Pleural_Thickening, Pneumonia, Pneumothorax
Dataset Structure
Train split: 78484 images
Validation split: 16818 images
Test split: 16818 images
Data Format
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Manas2703/chest-xray-14.wbcbench2026
WBCBench 2026 - Robust White Blood Cell Classification Under Class Imbalance
Dataset for the WBCBench 2026 ISBI Challenge.
A 13-class white-blood-cell classification benchmark with mixed pristine and degraded images for
robustness evaluation under realistic imaging conditions.
Paper: arXiv:2604.10797
Challenge website: https://xudong-ma.github.io/WBCBench2026-Robust-White-Blood-Cell-Classification/
Kaggle competition: https://www.kaggle.com/competitions/wbc-bench-2026/overview… See the full description on the dataset page: https://huggingface.co/datasets/Xin-Tian/wbcbench2026.xenium-senescence-demo
Xenium Senescence Benchmark (Demo Preview)
Note: This is a demo preview with 234 sample images. The full dataset (62K labeled cells across 2 tissue samples and 3 magnification scales) will be released upon paper publication.
Overview
A benchmark dataset for predicting cellular senescence from spatial transcriptomics (Xenium) cell images. Cell images are DAPI fluorescence microscopy captures at multiple magnification scales, labeled with senescence scores derived from… See the full description on the dataset page: https://huggingface.co/datasets/Xiang-zx-zx/xenium-senescence-demo.XJTU-annotated
Roles
Roles: reasoning view of XJTU — annot is the source label (inner_race / normal / outer_race), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the image is an envelope spectrum with the theoretical fault frequencies marked. The reasoning column is filled on all 907 records and is the SFT imitation target for this repo; the query enumerates the closed set of labels the answer must come from, and annot remains the… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/XJTU-annotated.mfg010-sample
MFG-010 — Manufacturing Defects Dataset (Sample)
A schema-identical preview of MFG-010, the XpertSystems.ai synthetic
defect events with visual-inspection ML metadata dataset for AOI
(Automated Optical Inspection) ML training, FMEA RPN modeling,
Ishikawa root cause classification, CAPA workflow simulation, and
defect-cohort quality engineering research. The full product covers
10,000-100,000 records. This sample is HF-sized at 3,000 records.
Built by XpertSystems.ai — Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/xpertsystems/mfg010-sample.xai-attack-detection-imagenette
XAI Attack Detection: Imagenette targeted BIM/PGD on ViT-B/16
Private research dataset of paired clean and targeted adversarial Imagenette images. It is
built to study how adversarial attacks change a Vision Transformer's explanation maps and to
support later work on attack detection. Each row is one source image with its clean and its
attacked version.
Summary
Pairs
12,420 (train 8,690 · validation 1,860 · test 1,870)
Source images
Imagenette v2… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-imagenette.audioform_dataset
AAA UUUUUUUU UUUUUUUUDDDDDDDDDDDDD IIIIIIIIII OOOOOOOOO FFFFFFFFFFFFFFFFFFFFFF OOOOOOOOO RRRRRRRRRRRRRRRRR MMMMMMMM MMMMMMMM
A:::A U::::::U U::::::UD::::::::::::DDD I::::::::I OO:::::::::OO F::::::::::::::::::::F OO:::::::::OO R::::::::::::::::R M:::::::M M:::::::M
A:::::A U::::::U U::::::UD:::::::::::::::DD I::::::::I OO:::::::::::::OO… See the full description on the dataset page: https://huggingface.co/datasets/xuyuefan111/audioform_dataset.cifar10
Dataset Specifications
Contains the entire CIFAR10 dataset, downloaded via PyTorch, then split and saved as .png files representing 32x32 images.
There a three splits, perfectly balanced class-wise:
train: 49,000 out of the original 50,000 samples from the training set of CIFAR10;
calibration: 1,000 left-out samples from the training set;
test: 10,000 samples, the entire original test set.
File Structure
Files are archives <split>/<classname>.zip. Each… See the full description on the dataset page: https://huggingface.co/datasets/xjy0123/cifar10.subaps_s1sm_x96_x5
SRSD Sub-Aperture S1SM x96
This dataset is packaged as one tar shard per metadata prefix to keep the file count manageable for Hugging Face Hub uploads.
Layout
shards/<prefix>.tar: archive containing metadata/<prefix>.csv and the matching tensors/*.safetensors files.
shard_index.csv: one row per shard with file counts and byte totals.
Source
Generated from /lustre/scratch/1001/rdelprete/srsd_patches/dataset_sm_subaps_x5.
Notes
Each… See the full description on the dataset page: https://huggingface.co/datasets/juanfra54/subaps_s1sm_x96_x5.chest_x_ray
Dataset Card for NIH Chest X-ray dataset
Dataset Summary
ChestX-ray dataset comprises 112,120 frontal-view X-ray images of 30,805 unique patients with the text-mined fourteen disease image labels (where each image can have multi-labels), mined from the associated radiological reports using natural language processing. Fourteen common thoracic pathologies include Atelectasis, Consolidation, Infiltration, Pneumothorax, Edema, Emphysema, Fibrosis, Effusion, Pneumonia… See the full description on the dataset page: https://huggingface.co/datasets/Shee2001/chest_x_ray.subaps_s1sm_x96
SRSD Sub-Aperture S1SM x96
This dataset is packaged as one tar shard per metadata prefix to keep the file count manageable for Hugging Face Hub uploads.
Layout
shards/<prefix>.tar: archive containing metadata/<prefix>.csv and the matching tensors/*.safetensors files.
shard_index.csv: one row per shard with file counts and byte totals.
Notes
Each tar shard groups tensor patches by the same prefix used by its metadata CSV.
Safetensor filenames are preserved… See the full description on the dataset page: https://huggingface.co/datasets/juanfra54/subaps_s1sm_x96.
