datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.sawhill-dataset
Sawhill Numismatic Collection Dataset
Dataset Description
This dataset contains video recordings and extracted images of coins from the MacKenzie Art Gallery's Sawhill Numismatic Collection. The dataset is designed for research in automated coin identification, cultural heritage digitization, and computer vision applications in numismatics.
Dataset Summary
Source: MacKenzie Art Gallery Sawhill Numismatic Collection
Content: Handheld video recordings of coins… See the full description on the dataset page: https://huggingface.co/datasets/COIN-Research-Group/sawhill-dataset.MetaPKLot-Dataset
MetaPKLot
A Large-Scale Benchmark for Vision-Based Parking Lot Management
2,265,974 labeled samples · 1,366,185 new annotations · 3 research challenges · COCO-style annotations
MetaPKLot is a large-scale, harmonized dataset designed for research on vision-based parking lot management.
It extends and standardizes three existing parking datasets:
PKLot
CNRPark-EXT
PLds
MetaPKLot introduces new annotations, revises existing parking-space annotations, standardizes… See the full description on the dataset page: https://huggingface.co/datasets/DSBD-Research/MetaPKLot-Dataset.SCIN-Dermatology-Raw-Images
SCIN-Dermatology-Raw-Images
This dataset contains 6,517 patient-submitted photographs organized into 3,061 clinical cases of common skin diseases. The source images are curated from the public Google Skin Condition Image Network (SCIN) corpus, cleansed of quality and gradability conflicts, and paired with complete patient-reported demographics, clinical symptoms, and dermatologist gradings.
Dataset Structure
This repository follows the standard Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/SCIN-Dermatology-Raw-Images.cardio-mark
CardioMark Review Subset
This repository contains an anonymized review subset of the CardioMark benchmark introduced for automated vertebral heart score (VHS) estimation in canine thoracic radiographs.
The subset is provided to support reproducibility and data-quality inspection during peer review.
Dataset Overview
CardioMark is a large-scale benchmark for evaluating the complete VHS measurement pipeline, including:
cardiac landmark localization
geometric VHS estimation… See the full description on the dataset page: https://huggingface.co/datasets/gen-ai-researcher/cardio-mark.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.OpenJev-Vision-Research-v0.1
OpenJev Vision Research v0.1
12,832 image records, with public provenance, original synthetic scenes,
and programmatically derived decision questions.
This is an experimental research dataset for visual posterior learning and
compositional decisions, released with OpenJev.
It is not a reproduction of TypeSafe's proprietary Jev model or training method.
Three separate configurations
Config
Images
What the labels mean
License
synthetic
8,192
Exact… See the full description on the dataset page: https://huggingface.co/datasets/IamBusy/OpenJev-Vision-Research-v0.1.
