datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.REPID
REPID: Rendering Evaluation of Photographic Image Dataset
REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA) in paper Beyond distortions: a benchmark for subjective evaluation of image rendering quality.
Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.relaion2b-natural-embeddings
LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart)
LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7).
Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings
Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.retina-age-analysis
Retina Age Analysis Dataset
Dataset Description
This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks.
Dataset Summary
Task: Age prediction from retinal fundus images
Images: 9,857 high-quality retinal images
Patients: 5,393 unique patients
Age Range: 5-97 years
Image Format: JPEG
Average Image Size: ~1 MB
Supported Tasks
Regression: Predict continuous age (5-97 years)
Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.rlbenchfail_val_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.HUGO-Bench-Paper-Reproducibility
HUGO-Bench Paper Reproducibility
Supplementary data and reproducibility materials for the paper:
Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study - https://arxiv.org/abs/2602.03894
Hugo Markoff, Stefan Hein Bengtson, Michael Ørsted
Aalborg University, Denmark
Dataset Description
This repository contains complete experimental results, pre-computed embeddings, and execution logs from our comprehensive benchmarking study… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench-Paper-Reproducibility.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.Recap-DataComp-1B-FoodOrDrink
Recap-DataComp-1B: Food or Drink
A filtered subset of Recap-DataComp-1B containing 106,230,157 rows classified as food/drink content, enriched with structured food/drink extraction from FoodExtract-v2.
Overview
Count
Percentage
Total rows
106,230,157
100%
Food/drink (Stage 5 label)
96,618,895
91.0%
Not food/drink (Stage 5 label)
9,611,262
9.0%
FoodExtract (re_caption): food/drink
79,519,489
74.9%
FoodExtract (re_caption): not food/drink
26,710,156… See the full description on the dataset page: https://huggingface.co/datasets/mrdbourke/Recap-DataComp-1B-FoodOrDrink.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.Ransomware_PE_Header_Feature_Dataset
Dataset Card for Ransomware PE Header Feature Dataset
Dataset Description
Dataset Summary
This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header.
Supported Tasks
Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/Rosalia1212/cbis-ddsm-r.openbrush-religious-art
OpenBrush Religious Art
Religious paintings from OpenBrush-75K — saints, biblical scenes, devotional works.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,119 you actually want.
Why this subset
A coherent visual genre: religious narrative painting from medieval through early modern. Heavy on Renaissance and Baroque eras. Common… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-religious-art.radr
Rad-R: A Real-World Raw-ADC Dataset and Benchmark for mmWave Radar Robustness
Webpage |
Code |
Paper |
PyPI
This is a demo release with 20 clips. The full dataset (~5,000 clips) will be available upon paper acceptance at NeurIPS 2026 Evaluations & Datasets Track.
Overview
Rad-R is the first mmWave radar dataset combining:
Raw ADC captures from a TI MMWCAS-RF-EVM 77 GHz cascaded radar (12 TX × 16 Rx = 192 virtual channels)
Controlled hardware fault… See the full description on the dataset page: https://huggingface.co/datasets/gtaxcenter/radr.image-aesthetic-scores
Rule34.nexus · Licence: Rule34.nexus Derived Dataset Licence 1.0
Rule34.nexus Image Aesthetic Scores
1. Overview
This dataset contains per-image aesthetic predictions for images in the Rule34.nexus corpus.
Predictions were generated using
discus0434/aesthetic-predictor-v2-5. Source images are not
included in this dataset — only opaque post identifiers, the source image's SHA-256 hash,
the post's content type, and the predicted score.… See the full description on the dataset page: https://huggingface.co/datasets/rule34nexus/image-aesthetic-scores.openbrush-renaissance
OpenBrush Renaissance
Renaissance works from OpenBrush-75K, combining Northern, Early, High, and Mannerism Late Renaissance into one period subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,565 you actually want.
Why this subset
Combined Renaissance period subset spanning ~1300–1600. Heavy on religious painting, portraits… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-renaissance.Light-RAG-Marketing-Assets-Agent
🖼️ Light RAG Marketing Assets Agent — Pre-ingested Data
Pre-ingested LightRAG knowledge graph and vector data from 420 marketing images
analyzed with Gemini Vision API (gemini-3.5-flash) and processed through GPT-4o
for entity extraction and relationship mapping.
GitHub repo: 0xrphl/Light-RAG-Marketing-Assets-Agent
📊 Dataset Statistics
Metric
Value
Source images
420 (JPG/PNG/WebP)
Text chunks
2,095 (5 per image: core, visual, people/setting… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/Light-RAG-Marketing-Assets-Agent.radread-public-results
RadRead — public results
Rollout-level results for RadRead, a benchmark of frontier models reading 150
radiographs. Every row is one graded model read: 5 saved rollouts per study
per model, scored by a deterministic grader (no judge model).
A read passes only when every required checklist finding, lesion box (the grader's
IoU / centre / containment test), lexical diagnosis check and action-set membership
check match the reference rubric. No partial credit inside a study;… See the full description on the dataset page: https://huggingface.co/datasets/tirandazdylan/radread-public-results.war-gov-uap-release-1
Department of War UAP Release 1 — structured corpus
The first tranche of declassified U.S. government records on Unidentified
Anomalous Phenomena (UAP / UFOs), released by the Department of War on
8 May 2026 under the Presidential Unsealing and Reporting System for
UAP Encounters (PURSUE) directive.
This dataset is a structured, machine-readable companion to the source
material at https://www.war.gov/UFO/. It pairs every original document
with VLM-extracted page text, cropped… See the full description on the dataset page: https://huggingface.co/datasets/MTSlive/war-gov-uap-release-1.trace-rx-eval-predictions
TRACE-RX Evaluation Predictions
Per-image detector scores from an independent evaluation of the two TechJam 2026 TRACE-RX
detectors, run 30 Aug – 1 Sep 2026.
No images here. Every file contains scores, labels, asset ids and transform names only — this is
derived evaluation metadata, not a redistribution of any source imagery. The underlying corpora
(Joshyxwa/data_draft, Joshyxwa/techjam2026, techjam-aigc/wildfake-eval-subset) keep their own
terms, and data_draft's WildFake rows… See the full description on the dataset page: https://huggingface.co/datasets/joelleoqiyi/trace-rx-eval-predictions.rvl-cdip-filtered
RVL-CDIP Filtered Dataset
This dataset contains filtered images from the RVL-CDIP dataset, focusing on 4 specific document types.
Dataset Summary
A filtered subset of the RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset containing 100,000 images across 4 document categories. Each image is stored as base64-encoded data in Parquet format for efficient processing.
Classes
Label
Class Name
Description
0
letter
Personal and… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/rvl-cdip-filtered.openbrush-rembrandt
OpenBrush Rembrandt
All Rembrandt works from OpenBrush-75K — paintings, etchings, and sketches.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 776 you actually want.
Why this subset
The defining body of work for Baroque chiaroscuro and dramatic light. Includes religious scenes, portraits, self-portraits, and biblical narratives.… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-rembrandt.diabetic-retinopathy-multimodal-progression-africa
DR-Progression — Multimodal 5-Year Diabetic-Retinopathy Risk with Africa-Grounded Synthetic EHR
A multimodal dataset for predicting 5-year diabetic-retinopathy progression and
risk from a fundus image combined with a full systemic electronic health record
(EHR): glycaemic control, diabetes duration, blood pressure, renal function,
comorbidities, treatment, and access-to-care.
Version 1.0.0 · core dr_synth 1.0.0 · part of the DR-Africa dataset
family (see also dr-grading and… See the full description on the dataset page: https://huggingface.co/datasets/macular/diabetic-retinopathy-multimodal-progression-africa.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/helloerikaaa/cbis-ddsm-r.RA3D
RA3D collision and measurement research data
Data and intermediate results for Holographic Droplet Collision Measurement, by Dai Nakai. The RA3D Python repository provides the algorithms, selective downloader, training, inference, and offline visual review.
This release contains six synthetic scenes, each with 6,000 synchronized frame pairs and 2,000 collision events. Seeds 260740–260744 form the development/training split. Seed 260745 is the terminal validation scene and is… See the full description on the dataset page: https://huggingface.co/datasets/dnakaikit/RA3D.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.
