datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION_Aesthetics_1024LAION_Aesthetics_512laion_improved_aesthetics_6.5plus_with_imagesterminusresearch-photo-aesthetics
Photo Aesthetics — Webshart
This is the canonical, maintained location for the former terminusresearch/photo-aesthetics dataset. The migration completed on August 19, 2026, with 30,032 image samples repackaged into 371 indexed Webshart shards. Every sample now includes both the original CogVLM caption and a newer GLM-5.3-Flash caption.
The legacy repository now contains only a relocation notice. Its original tar archives, parquet file, and prior history were removed after this… See the full description on the dataset page: https://huggingface.co/datasets/webshart/terminusresearch-photo-aesthetics.Laion-Aesthetics-High-Resolution-GoT
Laion-Aesthetics-High-Resolution-GoT
Paper
Dataset Description
The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information.
Key Features
Size: 3.77 million samples
Modalities: Image, Text, and Grounding Annotations
Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.Laion_aesthetics_5plus_1024_33MLAION_Aesthetics_1024_bucketed_1024
LAION Aesthetics 1024 Bucketed 1024 Captioned
Bucketed-shards export of images from limingcv/LAION_Aesthetics_1024, with a small high-resolution addon from limingcv/LAION_Aesthetics_512.
Images: 382,412
Shards: 398 uncompressed WebDataset-style TAR files
Format: bucketed_shards_v1
Base resolutions: [1024, 512]
Manifest: manifest.json
Images are organized under buckets/<bucket_id>/<shard>.tar. Each sample contains:
<key>.jpg
<key>.txt
<key>.json
The .txt files contain… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_1024.AVA-aesthetics-10pct-min50-10bins
AVA Aesthetics 10% Subset (min50, 10 bins)
This dataset is a curated 10% subset of the AVA Aesthetics Dataset (or the original AVA dataset as described in Murray et al., 2012). It includes images that have at least 50 total votes and have been stratified into 10 bins based on their computed mean aesthetic scores.
Dataset Overview
Dataset Name: AVA Aesthetics 10% Subset (min50, 10 bins)
Subset Size: 10% of the original AVA dataset (after filtering for a minimum of 50… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-aesthetics-10pct-min50-10bins.TAD66K_for_Image_Aesthetics_Assessment
Introduction
We build a large-scale dataset called the Theme and Aesthetics Dataset with 66K images (TAD66K), which is specifically designed for IAA. Specifically, (1) it is a theme-oriented dataset containing 66K images covering 47 popular themes. All images were carefully selected by hand based on the theme. (2) In addition to common aesthetic criteria, we provide 47 criteria for the 47 themes. Images of each theme are annotated independently, and each image contains at least… See the full description on the dataset page: https://huggingface.co/datasets/Shuai1995/TAD66K_for_Image_Aesthetics_Assessment.photo-aesthetics
Photo Aesthetics Dataset
Pulled from Pexels in 2023.
Image filenames may be used as captions, or, the parquet table contains the same values.
This dataset contains the full images.
Captions were created with CogVLM.
improved_aesthetics_4.5plus-ultra-hrversion https://git-lfs.github.com/spec/v1
oid sha256:98b45ea81164d1e1a1dd82255207053b15cd6c69d922a1c5cf3387ce604d4b74
size 28
improved_aesthetics_6.5pluslaion-aesthetics-12m-umap
LAION-Aesthetics :: CLIP → UMAP
This dataset is a CLIP (text) → UMAP embedding of the LAION-Aesthetics dataset - specifically the improved_aesthetics_6plus version, which filters the full dataset to images with scores of > 6 under the "aesthetic" filtering model.
Thanks LAION for this amazing corpus!
The dataset here includes coordinates for 3x separate UMAP fits using different values for the n_neighbors parameter - 10, 30, and 60 - which are broken out as separate columns with… See the full description on the dataset page: https://huggingface.co/datasets/dclure/laion-aesthetics-12m-umap.aesthetics_v2_4.75laion_aesthetics_sketchphoto-aesthetics
Photo Aesthetics Dataset
Pulled from Pexels in 2023.
Image filenames may be used as captions, or, the parquet table contains the same values.
This dataset contains the full images.
Captions were created with CogVLM.
LAION_Aesthetics_512_bucketed_512
LAION Aesthetics 512 Bucketed 512 Captioned
This is a captioned bucketed-shards export of images from limingcv/LAION_Aesthetics_512.
Images were filtered and resized/cropped into SDXL-style aspect-ratio buckets at a 512 base resolution, without upsampling. The export contains 1,999,908 images across 1,976 uncompressed WebDataset-style tar shards.
The .txt files contain model-generated captions, not the original LAION web-scrape alt text or surrounding page text. Captions were… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_512_bucketed_512.laion_aesthetics_v2_6.5plusLaion_aesthetics_5plus_1024_33M_csvphoto-aesthetics
Photo Aesthetics has moved
The maintained dataset is now available at webshart/terminusresearch-photo-aesthetics.
The replacement is a fully repackaged Webshart dataset with 30,032 captioned image samples across 371 indexed shards. Its paired JSON indexes include byte offsets, image geometry, and embedded captions for efficient random HTTP range access.
The replacement dataset card contains ready-to-use SimpleTuner and Webshart Python examples.
This legacy repository's tar… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/photo-aesthetics.aesthetics-wiki
Introduction
This dataset is webscraped version of aesthetics-wiki. There are 1022 aesthetics captured.
Columns + dtype
title: str
description: str (raw representation, including \n because it could help in structuring data)
keywords_spacy: str (['NOUN', 'ADJ', 'VERB', 'NUM', 'PROPN'] keywords extracted from description with POS from Spacy library)
removed weird characters, numbers, spaces, stopwords
Cleaning
Standard Pandas cleaning
Cleaned the data by… See the full description on the dataset page: https://huggingface.co/datasets/ninar12/aesthetics-wiki.aesthetics_prompts_laiondownload with
huggingface-cli download Yuanzhi/aesthetics_prompts_laion --repo-type=dataset --local-dir aesthetics
Aesthetic6+, Aesthetic6.25+ & Aesthetic6.5+ prompts dataset for SiD-LSG training.
Filtered from dclure/laion-aesthetics-12m-umap.
All credits to the authors of SiD.
from datasets import load_dataset
import os
dataset_dir = 'aesthetics'
# ds = load_dataset("parquet", data_dir=dataset_dir, split='train')
ds = load_dataset("aesthetics", data_files='train.parquet')
# ds =… See the full description on the dataset page: https://huggingface.co/datasets/Yuanzhi/aesthetics_prompts_laion.aesthetics_v2_4.5Ko-LAION-Aesthetics-10M
LAION-Aesthetics 10M Dataset Card
Dataset details
Dataset type:
Laion aesthetic is a subset of laion5B that has been estimated by a model trained on top of clip embeddings to be aesthetic. The intended usage of this dataset is image generation
Paper or resources for more information:
https://laion.ai/blog/laion-aesthetics/
Acknowledgements
This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grants funded by the… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/Ko-LAION-Aesthetics-10M.laion-aesthetics-recap-qwen3p5-35b-a3b
LAION-Aesthetics recaptions with Qwen3.5-35B-A3B
Dataset laion-aesthetics: 24.290 Million caption rows.
This public caption-only repository contains 24,290,381 generated captions for 23,687,875 image assets and no image payload. It includes 602,506 additional distinct caption variants. Rows match BootsofLagrangian/laion-aesthetics-webp90-min256px-noresize through image_shard and image_member; URL and content hashes support independent reconciliation.
Captions were produced with… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/laion-aesthetics-recap-qwen3p5-35b-a3b.improved_aesthetics_6.5plus_clip_retrievallimingcv_LAION_Aesthetics_1024-sd-scripts-5000Each zip contains;
image.png/jpg/etc -> is the image
image.txt -> contains the caption
The captions are what I extracted from the json files at runtime. Hindsight says I should have kept the json but it is what it is for now.
I'll run a better one later. This one took quite a few hours as it was.
aesthetics_v2_4.75_filteredlaion_aesthetics_v2_6.0plusdeep-art-aesthetics-zh
Deep Art & Aesthetics Dialogue Dataset (Chinese)
深度艺术与审美对话数据集
Dataset Description
High-quality Chinese art and aesthetics dialogues covering aesthetic theory, visual art analysis, symbolism, and design philosophy.
高质量中文艺术与审美对话,涵盖美学理论、视觉艺术分析、象征主义、设计哲学等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata: Source platform, topic tags… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-art-aesthetics-zh.
