datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.REPID
REPID: Rendering Evaluation of Photographic Image Dataset
REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA) in paper Beyond distortions: a benchmark for subjective evaluation of image rendering quality.
Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.retina-age-analysis
Retina Age Analysis Dataset
Dataset Description
This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks.
Dataset Summary
Task: Age prediction from retinal fundus images
Images: 9,857 high-quality retinal images
Patients: 5,393 unique patients
Age Range: 5-97 years
Image Format: JPEG
Average Image Size: ~1 MB
Supported Tasks
Regression: Predict continuous age (5-97 years)
Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.Ransomware_PE_Header_Feature_Dataset
Dataset Card for Ransomware PE Header Feature Dataset
Dataset Description
Dataset Summary
This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header.
Supported Tasks
Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/Rosalia1212/cbis-ddsm-r.humancentric-scenes-ai
HumanCentric-Scenes-AI
A multimodal benchmark of 296 AI-generated human-centric scenes across four domains:
CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise.
Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/rjmaftv33/humancentric-scenes-ai.honeybee-samples
HoneyBee Sample Files
Sample data and resource files for the HoneyBee framework — a scalable, modular toolkit for multimodal AI in oncology.
These files are used by the HoneyBee example notebooks (clinical, pathology, radiology) and by HoneyBee's molecular processing code at runtime (Hugo_symbols.tsv is fetched on first use of DNA mutation preprocessing).
Paper: HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models… See the full description on the dataset page: https://huggingface.co/datasets/Lab-Rasool/honeybee-samples.radr
Rad-R: A Real-World Raw-ADC Dataset and Benchmark for mmWave Radar Robustness
Webpage |
Code |
Paper |
PyPI
This is a demo release with 20 clips. The full dataset (~5,000 clips) will be available upon paper acceptance at NeurIPS 2026 Evaluations & Datasets Track.
Overview
Rad-R is the first mmWave radar dataset combining:
Raw ADC captures from a TI MMWCAS-RF-EVM 77 GHz cascaded radar (12 TX × 16 Rx = 192 virtual channels)
Controlled hardware fault… See the full description on the dataset page: https://huggingface.co/datasets/gtaxcenter/radr.Recruitment-Task-3
DeepWeeds - AI-MED AGH convenience mirror
This is a convenience mirror of the official DeepWeeds image archive and the
upstream annotations pinned to a specific commit. original/images.zip is
preserved unchanged; images are not extracted or duplicated here. models.zip
from the source authors is deliberately not mirrored.
Dataset facts
17,509 in-situ images from Queensland, Australia.
Nine classes: eight weed species plus Negative.
The authors publish five folds… See the full description on the dataset page: https://huggingface.co/datasets/AI-MED-AGH/Recruitment-Task-3.japanese-image-classification-evaluation-dataset
recruit-jp/japanese-image-classification-evaluation-dataset
Overview
Developed by: Recruit Co., Ltd.
Dataset type: Image Classification
Language(s): Japanese
LICENSE: CC-BY-4.0
More details are described in our tech blog post.
日本語CLIP学習済みモデルとその評価用データセットの公開
Dataset Details
This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks.
jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.hmdb51-pick-run-stand
HMDB51 — Pick / Run / Stand (frames procesados)
Subconjunto procesado del dataset HMDB51 para un caso de uso de
clasificación de productividad de empleados en almacén mediante visión por
computadora: distinguir entre trabajador activo (recogiendo / corriendo)
e inactivo/pausado (parado).
Clases
Clase HMDB51
Etiqueta de negocio
pick
Activo (Pick Up)
run
Activo (Running)
stand
Inactivo/Pausado (Standing)
Estadísticas
492 videos… See the full description on the dataset page: https://huggingface.co/datasets/treborDev/hmdb51-pick-run-stand.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/helloerikaaa/cbis-ddsm-r.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.DSGR
DSGR: Domain Shift across Geographic Regions
A large-scale Domain Generalisation (DG) benchmark for land-use classification in satellite imagery under spatial domain shift.
Official dataset of the paper: "Analysing Satellite Imagery Classification under Spatial Domain Shift across Geographic Regions", International Journal of Computer Vision (IJCV), 2025.
Sara A. Al-Emadi, Yin Yang, Ferda Ofli — Qatar Computing Research Institute (QCRI), Hamad Bin Khalifa University (HBKU)… See the full description on the dataset page: https://huggingface.co/datasets/RWGAI/DSGR.AmericanSignLanguageMNISTBased on Kaggle - Sign Language MNIST.Repackaged both CSV's into a single CSV with a field datasetType to assign each to its type.
The class mapping:
0: 'A', 1: 'B', 2: 'C', 3: 'D', 4: 'E', 5: 'F', 6: 'G', 7: 'H', 8: 'I',
10: 'K', 11: 'L', 12: 'M', 13: 'N', 14: 'O', 15: 'P', 16: 'Q', 17: 'R',
18: 'S', 19: 'T', 20: 'U', 21: 'V', 22: 'W', 23: 'X', 24: 'Y'
Labels 9 (J) and 25 (Z) are excluded as these letters require motion in ASL hence no such images are available.
TriALS-Report
TriALS-Report: A Multi-Center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT
Study workflow. Non-contrast CT volumes are paired with the triphasic contrast-enhanced report of the same patient; findings are extracted from the report to form the label space, and models are evaluated on disease diagnosis and report generation.
TriALS-Report is a multi-centre benchmark for abdominal disease diagnosis from non-contrast CT (NCCT), where the… See the full description on the dataset page: https://huggingface.co/datasets/marwankefah/TriALS-Report.reflection_boundary_challenge_v01ClarusC64/reflection_boundary_challenge_v01
Dataset summary
This dataset tests whether models handle mirrors and reflections without breaking container logic.Scenes contain real objects and mirrored views.Some reflections are consistent with the room.Others violate boundaries or basic geometry.
Main goals
detect when a reflection conflicts with the layout
keep track of entities visible only in mirrors
avoid placing entities across walls or through barriers
respect gravity and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reflection_boundary_challenge_v01.MNISTKyrgyzTest400The data is based on Kyrgyz MNIST.It is based on the Test Set.
Reproduce by:
numSamplesPerCls = 400
seedNum = 512
dfData = pd.read_csv(r'test.csv')
dfT = dfData.groupby('label', group_keys = False).sample(n = numSamplesPerCls, replace = False, random_state = seedNum)
dfT = dfT.reset_index(drop = False)
dfT = dfT.rename(columns = {'index': 'img_index'})
dfT.to_csv(r'MNISTKyrgyzTest400.csv', index = False)
MNISTThe dataset contains various MNIST like datasets in teh form of a csv files.
MNIST
Based on the MNIST Dataset in OpenML: OpenML mnist_784.
The way to reproduce:
from sklearn.datasets import fetch_openml
dfX, dsY = fetch_openml('mnist_784', version = 1, return_X_y = True, as_frame = True)
dfX.columns = [str(ii) for ii in range(dfX.shape[1])]
dfX['Label'] = dsY
dfX.to_csv('MNIST.csv')
Fashion MNIST
Based on Zalando Research - FashionMNIST.
Packaged into a CSV in a Row… See the full description on the dataset page: https://huggingface.co/datasets/Royi/MNIST.minicar-dataset
🏎️ MiniCar Autonomous Driving Dataset
自動運転ミニカー用のトレーニングデータセット
概要
このデータセットには以下が含まれます:
カメラ画像
センサーデータ(IMU等)
アノテーション(ステアリング角度、スロットル)
データ構造
minicar-dataset/
├── train/
│ ├── images/ # カメラ画像 (JPG/PNG)
│ ├── sensors/ # センサーデータ (CSV)
│ └── annotations.csv # ラベルデータ
├── test/
│ └── ...
└── README.md
使い方
from datasets import load_dataset
dataset = load_dataset("Romihi50/minicar-dataset")
# トレーニングデータ
for sample in… See the full description on the dataset page: https://huggingface.co/datasets/Romihi50/minicar-dataset.
