datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.Heliconius-Collection_Cambridge-Butterfly
Dataset Card for Heliconius Collection (Cambridge Butterfly)
Dataset Description
Dataset Summary
Subset of the collection records from Chris Jiggins' research group at the University of Cambridge, collection covers nearly 20 years of field studies.
This subset contains approximately 36,189 RGB images of 11,962 specimens (29,134 images of 10,086 specimens across all Heliconius). Many records have both images and locality data.
Most images were… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Heliconius-Collection_Cambridge-Butterfly.vecforge-paper-corpus
Note (rebuild in progress): figures are being re-extracted with a fixed extractor (cleaner crops). The image-preview config (Parquet with an inline column + difficulty/type/score labels) returns after re-classification. The config (paper metadata + links) is live now.
VecForge Paper Corpus
A pristine, deduplicated collection of 59,732 top-venue AI/ML/CV/NLP/Robotics papers (2020-2024) with
every captioned figure and full paper text, for research on figure understanding… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/vecforge-paper-corpus.CrossViewUrbanTrafficDataset
Cross-View Urban Traffic Dataset
Dataset Summary
The Cross-View Urban Traffic Dataset (CVUTD) is a benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections in Regensburg, Germany.
The dataset is designed to support two linked tasks:
Cross-view identity matching between street-view and drone-view object tracks
Ego-to-BEV prediction using aerial supervision… See the full description on the dataset page: https://huggingface.co/datasets/prakharbh/CrossViewUrbanTrafficDataset.coyo-labeled-300m
Dataset Card for COYO-Labeled-300M
Dataset Summary
COYO-Labeled-300M is a dataset of machine-labeled 300M images-multi-label pairs. We labeled subset of COYO-700M with a large model (efficientnetv2-xl) trained on imagenet-21k. We followed the same evaluation pipeline as in efficientnet-v2. The labels are top 50 most likely labels out of 21,841 classes from imagenet-21k. The label probabilies are provided rather than label so that the user can select threshold of their… See the full description on the dataset page: https://huggingface.co/datasets/kakaobrain/coyo-labeled-300m.ODELIA-Challenge-2025
ODELIA Challenge Dataset
This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning.
The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.central-florida-native-plants
DeepEarth Central Florida Native Plants Dataset v0.2.0
🌿 Dataset Summary
A comprehensive multimodal dataset featuring 33,665 observations of 232 native plant species from Central Florida. This dataset combines citizen science observations with state-of-the-art vision and language embeddings for advancing multimodal self-supervised ecological intelligence research.
Key Features
🌍 Spatiotemporal Coverage: Complete GPS coordinates and timestamps for all… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants.conflux-chest-ct
CONFLUX Chest-CT
200,000 synthetic 3D chest CT volumes with structured abnormality and demographic labels, generated by CONFLUX.
Released with the paper CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training.
Paper (arXiv) •
Model •
Code — coming soon
About
CONFLUX is a conditional 3D latent generative model for chest CT: a VAE tokenizer
compresses each volume into a compact 16-channel latent, a… See the full description on the dataset page: https://huggingface.co/datasets/gevaertlab/conflux-chest-ct.galaxy-chirality-catalog
DESI Legacy Galaxy Chirality Catalog
This dataset accompanies the current Paper IV manuscript, An Observed-Label Chirality-Dipole Null in 949,584 High-Confidence DESI Spirals and an 8.5-Million-Galaxy Catalog.
The primary high-confidence observed-label statistic is consistent with zero under fixed-occupancy label randomization (z=0.7053169638, one-sided empirical-rank p=0.2246775322). This is not a calibrated true-spin, physical-amplitude, or primordial-parity bound.… See the full description on the dataset page: https://huggingface.co/datasets/bamfai/galaxy-chirality-catalog.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting instances, which follows the COCO format… See the full description on the dataset page: https://huggingface.co/datasets/eelianafang/MOUNT-Cattle.Ransomware_PE_Header_Feature_Dataset
Dataset Card for Ransomware PE Header Feature Dataset
Dataset Description
Dataset Summary
This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header.
Supported Tasks
Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.similar-but-different
Similar But Different — Sentinel-2 patches with deceptive RGB
30,927 32×32 Sentinel-2 L2A multispectral patches (12 bands resampled to
10 m) across ten ESA WorldCover classes, selected so that the visible
bands are uninformative by construction: every patch sits in a region of
RGB-mean space dominated by patches of other classes, while its
NIR / red-edge / SWIR response stays class-informative.
The dataset is a controlled probe for one question: does a model actually
use the… See the full description on the dataset page: https://huggingface.co/datasets/calebrob6/similar-but-different.CleanPatrick
CleanPatrick: A Benchmark for Data Cleaning
Welcome to CleanPatrick, the first large-scale benchmark designed for data cleaning in the image domain.
Built on the Fitzpatrick17k dermatology dataset, CleanPatrick is a dataset for measuring the performance in detecting three major data quality issues:
off-topic samples, near-duplicates, and label errors.
Overview
CleanPatrick consists of dermatological images annotated with over 500,000 binary labels across three data… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Dermatology/CleanPatrick.phenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.cnn-based-drowsiness-detection-data
CNN-Based Drowsiness Detection - Dataset
Preprocessed, auto-labeled face-crop images used to train the model in
notgoodkeeper/cnn-based-drowsiness-detection.
Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection
Collection
Frames were captured from a webcam, then run through:
Haar Cascade face detection -> crop + pad + resize to 412x412
MediaPipe Selfie Segmentation -> background replaced with white
CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.PittImageVideoAdsDataset
Dataset Card for PittImageVideoAdsDataset
Dataset Summary
PittImageVideoAdsDataset is the image and video advertisement dataset released with Automatic Understanding of Image and Video Advertisements. The paper reports 64,832 image advertisements and 3,477 YouTube advertisement videos, with human annotations for topics, sentiments, slogans, persuasive strategies, symbolic references, and action/reason Q/A. This Hugging Face version exposes the public annotation… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PittImageVideoAdsDataset.CheXTemporal
CheXTemporal
A longitudinal chest-radiograph dataset of paired (current, prior) studies
with disease–progression labels, anatomy-aligned segmentation masks, and
sentence-level static/dynamic annotations, derived from CheXpert, MIMIC-CXR,
and ReXGradient. CheXTemporal pairs an expert-annotated gold evaluation
split with a much larger MedGemma-generated silver training corpus.
This release contains annotations only. Images must be downloaded
separately from each parent corpus under… See the full description on the dataset page: https://huggingface.co/datasets/anonaccount107240/CheXTemporal.crophelth
CropHelth — Unified Crop Disease & Watering Dataset (Round-1 Pre-training)
Author: Hansaka Rasanjana
Project Context: First-year IoT group mini-project at the Sri Lanka Institute of Information Technology (SLIIT)
~125,000 plant leaf images across 109 disease/healthy classes for 26 crops, complete with a per-image provenance manifest, treatment knowledge base, and per-crop watering recommendation rules.
Built as the core pre-training and round-1 dataset for our CropHelth system:… See the full description on the dataset page: https://huggingface.co/datasets/hansaka01/crophelth.vernier
vernier
Error bars on a dataset vendor's quality claim. Build AI publishes hand-visibility and
active-manipulation rates for Egocentric-10K / Egocentric-100K, judged once by
gemini-2.5-flash with no human gold, no interval, and no test that the judge scores a factory
floor and a home kitchen on the same scale. This release is the data behind an independent,
pre-registered measurement of that claim: human labels against a written rubric, a live
open-weights judge on the same… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/vernier.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/Rosalia1212/cbis-ddsm-r.drawvla-prompt-validation-clean
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-clean.lalafo-kg-cars
lalafo.kg — Kyrgyzstan Cars (used-car market)
Scraped from lalafo.kg, the largest informal classifieds
board in Kyrgyzstan — messier and larger than the curated boards, and closer to the
real street-level market. Field names are English; values are kept in the original
language (Russian).
Subsets
subset
rows
description
listings
54,518
one row per advertisement (every category, deal and region in scope)
users
49,685
sellers (ad authors), with… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/lalafo-kg-cars.CulturalBiases-2025Preprint : [https://arxiv.org/pdf/2505.14729?]
ClevelandMuseumArt
Cleveland Museum of Art Open Access Dataset
This dataset contains the complete Cleveland Museum of Art Open Access collection data, originally sourced from the ClevelandMuseumArt/openaccess GitHub repository and reuploaded for broader distribution, Parquet generation for high-performance analytics, and seamless integration with data science workflows.
Dataset Description
The Cleveland Museum of Art provides open access to information on more than 61,000 artworks in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/ClevelandMuseumArt.Full-Danbooru-Complement
Full Danbooru Complement
发布状态 / Release status(2026-08-06):基线版本、2026-06 月度归档和 2026-07 月度归档均已完成整理、验证并发布。
The baseline, 2026-06 monthly archive, and 2026-07 monthly archive have all been assembled, verified, and published.
概述 / Overview
Full Danbooru Complement 是 deepghs/danbooru2024-webp-4Mpixel 的后续补充数据集,用于延伸其 Danbooru 图像与 post metadata 覆盖范围。数据按不可变发布目录组织;每个 Parquet 行对应一个 Danbooru post,并在行内保存图像字节与相关 metadata。
Full Danbooru Complement is a follow-up complement to… See the full description on the dataset page: https://huggingface.co/datasets/Xuness/Full-Danbooru-Complement.womens-fashion-catalog
Livostyle Women's Fashion Catalog — Open Data
Open, machine-readable, weekly-updated catalog of 2,766+ curated women's fashion
products from Livostyle.com — a US DTC retailer
(Arcada LLC, Delaware). Free under MIT license for AI/LLM training,
recommender systems, fashion NLP research, and multimodal learning.
TL;DR
from datasets import load_dataset
ds = load_dataset("arturayupov/womens-fashion-catalog")
# ds["products"] → 2,766 products
# ds["images"] → 12,978… See the full description on the dataset page: https://huggingface.co/datasets/arturayupov/womens-fashion-catalog.NIH-CXR14-BiomedCLIP-Features
NIH-CXR14-BiomedCLIP-Features Dataset
This dataset is derived from the NIH Chest X-ray Dataset (NIH-CXR14) and processed using the BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 model from Microsoft. It contains image and text features extracted from chest X-ray images and their corresponding textual findings.
Dataset Description
The original NIH-CXR14 dataset comprises 112,120 chest X-ray images with disease labels from 30,805 unique patients. This processed dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yasintuncer/NIH-CXR14-BiomedCLIP-Features.
