datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aesthetic-4K
Aesthetic-4K Dataset
We introduce Aesthetic-4K, a high-quality dataset for ultra-high-resolution image generation, featuring carefully selected images and captions generated by GPT-4o.
Additionally, we have meticulously filtered out low-quality images through manual inspection, excluding those with motion blur, focus issues, or mismatched text prompts.
For more details, please refer to our paper:
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models (CVPR… See the full description on the dataset page: https://huggingface.co/datasets/zhang0jhon/Aesthetic-4K.grayscale_image_aesthetic_3M
Dataset Card for "grayscale_image_aesthetic_3M"
More Information needed
highresolution-laioncoco-aesthetic-MEGThis dataset is filtered from laioncoco-aesthetic, which is used for academic research on mobile edge generation (MEG).
It includes high-resolution 1024-by-1024 text-to-image samples generated by a distilled SDXL with 4-12 denoising steps.
The dataset mainly involves the following fields:
caption: The text prompt of the image.
image: The target image corresponding to the prompt.
diffusion: The generative results of the distilled SDXL.
latents: The latent features of the distilled SDXL.
LAION_Aesthetics_1024LAION_Aesthetics_512laion_improved_aesthetics_6.5plus_with_imagesterminusresearch-photo-aesthetics
Photo Aesthetics — Webshart
This is the canonical, maintained location for the former terminusresearch/photo-aesthetics dataset. The migration completed on August 19, 2026, with 30,032 image samples repackaged into 371 indexed Webshart shards. Every sample now includes both the original CogVLM caption and a newer GLM-5.3-Flash caption.
The legacy repository now contains only a relocation notice. Its original tar archives, parquet file, and prior history were removed after this… See the full description on the dataset page: https://huggingface.co/datasets/webshart/terminusresearch-photo-aesthetics.Aesthetic-Train-V2
Aesthetic-Train-V2 Dataset
We introduce Aesthetic-Train-V2, a high-quality traing set for ultra-high-resolution image generation.
For more details, please refer to our paper:
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models (CVPR 2025)
Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
Source code is available at https://github.com/zhang0jhon/diffusion-4k.
Citation
If you find our paper or dataset is helpful in your… See the full description on the dataset page: https://huggingface.co/datasets/zhang0jhon/Aesthetic-Train-V2.OpenNiji-Dataset-Aesthetic-Finetune-0-15K
Dataset Card for "OpenNiji-Dataset-Aesthetic-Finetune-0-15K"
More Information needed
lunara-aesthetic-image-variations
Dataset Card for Moonworks Lunara Aesthetic II
This dataset introduces the second open-source release by Moonworks. This dataset contains original image and art created by Moonworks and their contextual variations generated by Moonworks Lunara, a sub-10B parameter model with a novel diffusion mixture architecture.
Paper: https://arxiv.org/pdf/2602.01666
While part 1 is intended for learning and evaluating regional as well as region-agnostic art styles, part 2 is intended for… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic-image-variations.Laion-Aesthetics-High-Resolution-GoT
Laion-Aesthetics-High-Resolution-GoT
Paper
Dataset Description
The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information.
Key Features
Size: 3.77 million samples
Modalities: Image, Text, and Grounding Annotations
Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.LAION_Aesthetics_1024_bucketed_1024
LAION Aesthetics 1024 Bucketed 1024 Captioned
Bucketed-shards export of images from limingcv/LAION_Aesthetics_1024, with a small high-resolution addon from limingcv/LAION_Aesthetics_512.
Images: 382,412
Shards: 398 uncompressed WebDataset-style TAR files
Format: bucketed_shards_v1
Base resolutions: [1024, 512]
Manifest: manifest.json
Images are organized under buckets/<bucket_id>/<shard>.tar. Each sample contains:
<key>.jpg
<key>.txt
<key>.json
The .txt files contain… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_1024.lunara-aesthetic
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B parameter… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic.Laion_aesthetics_5plus_1024_33Mlaplacian_image_aesthetic_3M
Dataset Card for "laplacian_image_aesthetic_3M"
More Information needed
Aesthetic3D_Sample_Part1AVA-aesthetics-10pct-min50-10bins
AVA Aesthetics 10% Subset (min50, 10 bins)
This dataset is a curated 10% subset of the AVA Aesthetics Dataset (or the original AVA dataset as described in Murray et al., 2012). It includes images that have at least 50 total votes and have been stratified into 10 bins based on their computed mean aesthetic scores.
Dataset Overview
Dataset Name: AVA Aesthetics 10% Subset (min50, 10 bins)
Subset Size: 10% of the original AVA dataset (after filtering for a minimum of 50… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-aesthetics-10pct-min50-10bins.laion2B-en-aestheticTAD66K_for_Image_Aesthetics_Assessment
Introduction
We build a large-scale dataset called the Theme and Aesthetics Dataset with 66K images (TAD66K), which is specifically designed for IAA. Specifically, (1) it is a theme-oriented dataset containing 66K images covering 47 popular themes. All images were carefully selected by hand based on the theme. (2) In addition to common aesthetic criteria, we provide 47 criteria for the 47 themes. Images of each theme are annotated independently, and each image contains at least… See the full description on the dataset page: https://huggingface.co/datasets/Shuai1995/TAD66K_for_Image_Aesthetics_Assessment.laion2B-en-aesthetic-seed
Dataset Card for "laion2B-en-aesthetic-seed"
More Information needed
photo-aesthetics
Photo Aesthetics Dataset
Pulled from Pexels in 2023.
Image filenames may be used as captions, or, the parquet table contains the same values.
This dataset contains the full images.
Captions were created with CogVLM.
improved_aesthetics_4.5plus-ultra-hrversion https://git-lfs.github.com/spec/v1
oid sha256:98b45ea81164d1e1a1dd82255207053b15cd6c69d922a1c5cf3387ce604d4b74
size 28
lunara-aesthetic
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B… See the full description on the dataset page: https://huggingface.co/datasets/rico2512/lunara-aesthetic.laion-aesthetics-12m-umap
LAION-Aesthetics :: CLIP → UMAP
This dataset is a CLIP (text) → UMAP embedding of the LAION-Aesthetics dataset - specifically the improved_aesthetics_6plus version, which filters the full dataset to images with scores of > 6 under the "aesthetic" filtering model.
Thanks LAION for this amazing corpus!
The dataset here includes coordinates for 3x separate UMAP fits using different values for the n_neighbors parameter - 10, 30, and 60 - which are broken out as separate columns with… See the full description on the dataset page: https://huggingface.co/datasets/dclure/laion-aesthetics-12m-umap.aesthetics_v2_4.75laion-coco-aesthetic
LAION COCO with aesthetic score and watermark score
This dataset contains 10% samples of the LAION-COCO dataset filtered by some text rules (remove url, special tokens, etc.), and image rules (image size > 384x384, aesthetic score>4.75 and watermark probability<0.5). There are total 8,563,753 data instances in this dataset. And the corresponding aesthetic score and watermark score are also included.
Noted: watermark score in the table means the probability of the existence of… See the full description on the dataset page: https://huggingface.co/datasets/guangyil/laion-coco-aesthetic.aesthetic-v2
Dataset Card for "aesthetic-v2"
More Information needed
lunara-aesthetic
Dataset Card for Moonworks Lunara Aesthetic Dataset
Sample Images
Dataset Summary
paper: https://arxiv.org/abs/2601.07941
The Lunara Aesthetic Dataset is a curated collection of 2,000 high-quality image–prompt pairs designed for controlled research on prompt grounding, style conditioning, and aesthetic alignment in text-to-image generation.
All images are generated using the Moonworks Lunara, a sub-10B… See the full description on the dataset page: https://huggingface.co/datasets/dowangoo/lunara-aesthetic.aesthetic_6klaion_aesthetics_sketch
