datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION_Aesthetics_1024LAION_Aesthetics_512Laion-Aesthetics-High-Resolution-GoT
Laion-Aesthetics-High-Resolution-GoT
Paper
Dataset Description
The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information.
Key Features
Size: 3.77 million samples
Modalities: Image, Text, and Grounding Annotations
Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.Laion_aesthetics_5plus_1024_33MAVA-aesthetics-10pct-min50-10bins
AVA Aesthetics 10% Subset (min50, 10 bins)
This dataset is a curated 10% subset of the AVA Aesthetics Dataset (or the original AVA dataset as described in Murray et al., 2012). It includes images that have at least 50 total votes and have been stratified into 10 bins based on their computed mean aesthetic scores.
Dataset Overview
Dataset Name: AVA Aesthetics 10% Subset (min50, 10 bins)
Subset Size: 10% of the original AVA dataset (after filtering for a minimum of 50… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-aesthetics-10pct-min50-10bins.TAD66K_for_Image_Aesthetics_Assessment
Introduction
We build a large-scale dataset called the Theme and Aesthetics Dataset with 66K images (TAD66K), which is specifically designed for IAA. Specifically, (1) it is a theme-oriented dataset containing 66K images covering 47 popular themes. All images were carefully selected by hand based on the theme. (2) In addition to common aesthetic criteria, we provide 47 criteria for the 47 themes. Images of each theme are annotated independently, and each image contains at least… See the full description on the dataset page: https://huggingface.co/datasets/Shuai1995/TAD66K_for_Image_Aesthetics_Assessment.improved_aesthetics_4.5plus-ultra-hrversion https://git-lfs.github.com/spec/v1
oid sha256:98b45ea81164d1e1a1dd82255207053b15cd6c69d922a1c5cf3387ce604d4b74
size 28
improved_aesthetics_6.5pluslaion-aesthetics-12m-umap
LAION-Aesthetics :: CLIP → UMAP
This dataset is a CLIP (text) → UMAP embedding of the LAION-Aesthetics dataset - specifically the improved_aesthetics_6plus version, which filters the full dataset to images with scores of > 6 under the "aesthetic" filtering model.
Thanks LAION for this amazing corpus!
The dataset here includes coordinates for 3x separate UMAP fits using different values for the n_neighbors parameter - 10, 30, and 60 - which are broken out as separate columns with… See the full description on the dataset page: https://huggingface.co/datasets/dclure/laion-aesthetics-12m-umap.aesthetics_v2_4.75laion_aesthetics_sketchlaion_aesthetics_v2_6.5plusLaion_aesthetics_5plus_1024_33M_csvaesthetics_v2_4.5laion-aesthetics-recap-qwen3p5-35b-a3b
LAION-Aesthetics recaptions with Qwen3.5-35B-A3B
Dataset laion-aesthetics: 24.290 Million caption rows.
This public caption-only repository contains 24,290,381 generated captions for 23,687,875 image assets and no image payload. It includes 602,506 additional distinct caption variants. Rows match BootsofLagrangian/laion-aesthetics-webp90-min256px-noresize through image_shard and image_member; URL and content hashes support independent reconciliation.
Captions were produced with… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/laion-aesthetics-recap-qwen3p5-35b-a3b.limingcv_LAION_Aesthetics_1024-sd-scripts-5000Each zip contains;
image.png/jpg/etc -> is the image
image.txt -> contains the caption
The captions are what I extracted from the json files at runtime. Hindsight says I should have kept the json but it is what it is for now.
I'll run a better one later. This one took quite a few hours as it was.
aesthetics_v2_4.75_filteredlaion_aesthetics_v2_6.0plusLAION_Aesthetics_1024_bucketed_512
LAION Aesthetics 1024 Bucketed 512 Captioned
This is a captioned bucketed-shards export of images from limingcv/LAION_Aesthetics_1024.
Images were filtered and resized/cropped into SDXL-style aspect-ratio buckets at a 512 base resolution, without upsampling. The export contains 382,144 images across 397 uncompressed WebDataset-style tar shards.
The .txt files now contain model-generated captions, not the original LAION web-scrape alt text or surrounding page text. Captions were… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_512.Aesthetics_X_Phone_720p_Images_Rec_Captioned_16_9
Aesthetics X Image Dataset
Overview
This dataset contains high-quality aesthetic images collected from Twitter user @aestheticsguyy. The collection features visually pleasing digital artwork, wallpapers, and photography with a focus on visual appeal and design inspiration.
Dataset Contents
• Image files in JPEG/PNG format• High-resolution wallpaper collections• Thematically organized visual content
Collection Methodology
Images were gathered from… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Aesthetics_X_Phone_720p_Images_Rec_Captioned_16_9.aestheticssemiautomatic-aestheticslaion_aesthetics_v2_6.25pluslabeled_aesthetics_simpsons
Dataset Card for "labeled_aesthetics_simpsons"
More Information needed
aesthetics_6_5pluspixiv-image-toplist-aesthetics
Introduction
I mannually choose 342 images from top-100 in weekly toplist of 2022.I have to mention that some of painters may not be consent to the use of their art works in AI training.
sample_font_aesthetics_dswater_glassbottle_aesthetics_rated
Dataset Card for Dataset Name
This dataset holds 121 images of glass bottles for drinking water. The aesthetics were rated by five participants from Germany across different demographics.
aesthetic_sim_under3_toekns_5_70patch-laion-aesthetics
