datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CommunityForensics-Small
Community Forensics: Using Thousands of Generators to Train Fake Image Detectors (CVPR 2025)
Paper / Project Page / Code (GitHub)
This is a small version of the Community Forensics dataset. It contains roughly 11% of the generated images of the base dataset and is paired with real data with redistributable license. This dataset is intended for easier prototyping as you do not have to download the corresponding real datasets separately.
We distribute this dataset with a… See the full description on the dataset page: https://huggingface.co/datasets/OwensLab/CommunityForensics-Small.smart-bin-detect
arudaev/smart-bin-detect
Training data for Smart Bin Recognition – a validator ("is there a bin?")
and an identifier ("which bin?"). The design lives in docs/04-ml-pipeline.md
in the project repo, which is private; the manifests here carry per-image
provenance and are the authoritative record of what this dataset contains.
Every image carries provenance: source, source URL, licence, region,
capture date, annotator where known, label origin (human / machine /
legacy /… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/smart-bin-detect.MOUSS_fish_imagery_dataset_grayscale_small
Dataset Card for Modular Optical Underwater Survey System (MOUSS) Imagery - Small Set
This dataset contains grayscale underwater imagery collected by NOAA's Modular Optical Underwater Survey System (MOUSS), specifically for object detection of fish. The dataset is intended for training and evaluating models like the YOLOv8n-based Fish Detector on grayscale underwater footage.
Dataset Details
Dataset Description
This dataset is composed of black-and-white… See the full description on the dataset page: https://huggingface.co/datasets/akridge/MOUSS_fish_imagery_dataset_grayscale_small.Smart-Projector-Image-Classification-Dataset
Smart Projector Image Classification Dataset
With the rapid development of smart devices, a variety of portable projectors have emerged in the market. In practical applications, accurately recognizing and classifying the appearance images of these projectors has become a technical challenge. Current image recognition technology often yields poor classification results due to insufficient datasets or inaccurate annotations when dealing with different brands and models of projectors.… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Smart-Projector-Image-Classification-Dataset.MV-VDB-photos-small
MV-VDB-photos-small
Media Vault - Vector Database Photos (Small)
A curated collection of 11,000 images from various computer vision datasets, designed for testing internal mechanisms in the Media Vault Vector Database system. This is the first small-scale dataset (targeting 10K samples, with NSFW split totaling 11K) for validation and testing purposes.
Dataset Structure
The dataset contains two splits:
sfw: All non-NSFW images (~10,000 images)
x_nsfw: Only NSFW images… See the full description on the dataset page: https://huggingface.co/datasets/SamoXXX/MV-VDB-photos-small.SAVANT-CODALM-small
SAVANT CODALM Small Dataset
This dataset is part of the SAVANT framework described in the SAVANT paper, currently under peer review.
This repository is provided for peer-review purposes only. After the review process, the dataset will be made publicly available through the authors' main account.
Dataset Description
CODALM small contains 100 real-world driving images (50 anomalous, 50 normal) derived from the CODA corner case dataset. Each image includes manual annotation… See the full description on the dataset page: https://huggingface.co/datasets/u94fmn391j/SAVANT-CODALM-small.Smart-Air-Conditioner-Image-Classification-Dataset
Smart Air Conditioner Image Classification Dataset
The current smart device industry faces challenges such as increased complexity in user interaction and insufficient accuracy in device recognition. Existing solutions often rely on traditional image recognition technology, resulting in low accuracy and poor adaptability. The construction of the Smart Air Conditioner Image Classification Dataset aims to address this issue by providing high-quality image data to help models better… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Smart-Air-Conditioner-Image-Classification-Dataset.index-card-blank-content
Index-card blank / content / divider classifier — dataset
Cropped single archival index cards labelled blank, content, or divider, for
training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata
extraction in card-catalogue digitisation pipelines.
Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library
of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection.
How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.Smart-Speaker-Image-Classification-Dataset
Smart Speaker Image Classification Dataset
The current smart speaker market is highly competitive with diverse brands and forms, leading to difficulties for consumers in identification when making purchases. Existing image datasets mostly focus on general products, lacking classification datasets specifically for smart speakers, resulting in obvious deficiencies in brand identification and appearance classification. This dataset aims to help develop more efficient smart speaker… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Smart-Speaker-Image-Classification-Dataset.text-2-image-dpo-human-preferences-small
Text-2-Image DPO Human Preferences (Small)
A quality-controlled human preference dataset for text-to-image generation. 40,000 trust-weighted pairwise judgments from calibrated annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the highest-annotator-quality subset. For the full 5,000-pair dataset, see datapointai/text-2-image-dpo-human-preferences.
Built on the Datapoint annotation platform — purpose-built… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-small.
