datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.FruitVision_quality_classification
FruitVision Quality Classification
A dataset for quality classification of apples, bananas, mangoes, grapes, and oranges. The dataset contains raw and augmented versions.The raw dataset contains 10,154 images.Images per class:
Formalin-mixed: 3,176
Fresh: 3,800
Rotten: 3,178
The augmented dataset contains 73,389 images.Images per class:
Formalin-mixed: 22,228
Fresh: 30,400
Rotten: 20,761
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/FruitVision_quality_classification.ad-creative-quality-human-vs-llm
Human Expert vs LLM Judge: Facebook Ad Creative Quality
500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.
The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.banana_guava_quality_classification
Banana Guava Quality Classification
A dataset for quality classification of bananas and guavas. The dataset contains 1,748 images across 3 classes: Class_A, Class_B, Defect.Images per class:
Class_A: 671
Class_B: 469
Defect: 608
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{kumari2024banana,
title={Banana and Guava dataset for machine learning and deep learning-based quality classification}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/banana_guava_quality_classification.jwst-quality-analysis-dataset
JWST Quality Analysis Dataset
Overview
This dataset contains comprehensive quality analysis for 2,709 JWST (James Webb Space Telescope) NIRCam images from the MAST archive. Each image has been automatically analyzed for quality metrics, artifact detection, and noise characteristics.
Dataset Information
Size: 2,709 images
Format: JSONL (JSON Lines)
Source: JWST NIRCam observations from MAST
Targets: M16, NGC 3132, NGC 3324, SMACS 0723, Stephan's Quintet… See the full description on the dataset page: https://huggingface.co/datasets/norbertm/jwst-quality-analysis-dataset.vintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
bangumibase-face-quality-cls
deepghs-cv/bangumibase-face-quality-cls
A 4-class anime face drawing-quality classification dataset (88,000 single-character close-up frames, 818 anime shows), derived from deepghs/bangumibase-webp-4Mpixel_x. Labels are {poor, ok, good, excellent} from face_geo_mean = sqrt(face_w · face_h) after a strict single-face + single-head + face∩head-overlap consistency filter.
⚠ Scope. Every row is a frame from an anime TV-series broadcast. The dataset is intended for training/evaluating… See the full description on the dataset page: https://huggingface.co/datasets/deepghs-cv/bangumibase-face-quality-cls.
