datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION_Aesthetics_1024LAION_Aesthetics_512laion_improved_aesthetics_6.5plus_with_imagesLaion-Aesthetics-High-Resolution-GoT
Laion-Aesthetics-High-Resolution-GoT
Paper
Dataset Description
The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information.
Key Features
Size: 3.77 million samples
Modalities: Image, Text, and Grounding Annotations
Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.Laion_aesthetics_5plus_1024_33MAVA-aesthetics-10pct-min50-10bins
AVA Aesthetics 10% Subset (min50, 10 bins)
This dataset is a curated 10% subset of the AVA Aesthetics Dataset (or the original AVA dataset as described in Murray et al., 2012). It includes images that have at least 50 total votes and have been stratified into 10 bins based on their computed mean aesthetic scores.
Dataset Overview
Dataset Name: AVA Aesthetics 10% Subset (min50, 10 bins)
Subset Size: 10% of the original AVA dataset (after filtering for a minimum of 50… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-aesthetics-10pct-min50-10bins.improved_aesthetics_4.5plus-ultra-hrversion https://git-lfs.github.com/spec/v1
oid sha256:98b45ea81164d1e1a1dd82255207053b15cd6c69d922a1c5cf3387ce604d4b74
size 28
improved_aesthetics_6.5pluslaion-aesthetics-12m-umap
LAION-Aesthetics :: CLIP → UMAP
This dataset is a CLIP (text) → UMAP embedding of the LAION-Aesthetics dataset - specifically the improved_aesthetics_6plus version, which filters the full dataset to images with scores of > 6 under the "aesthetic" filtering model.
Thanks LAION for this amazing corpus!
The dataset here includes coordinates for 3x separate UMAP fits using different values for the n_neighbors parameter - 10, 30, and 60 - which are broken out as separate columns with… See the full description on the dataset page: https://huggingface.co/datasets/dclure/laion-aesthetics-12m-umap.aesthetics_v2_4.75laion_aesthetics_sketchlaion_aesthetics_v2_6.5plusLaion_aesthetics_5plus_1024_33M_csvaesthetics-wiki
Introduction
This dataset is webscraped version of aesthetics-wiki. There are 1022 aesthetics captured.
Columns + dtype
title: str
description: str (raw representation, including \n because it could help in structuring data)
keywords_spacy: str (['NOUN', 'ADJ', 'VERB', 'NUM', 'PROPN'] keywords extracted from description with POS from Spacy library)
removed weird characters, numbers, spaces, stopwords
Cleaning
Standard Pandas cleaning
Cleaned the data by… See the full description on the dataset page: https://huggingface.co/datasets/ninar12/aesthetics-wiki.aesthetics_prompts_laiondownload with
huggingface-cli download Yuanzhi/aesthetics_prompts_laion --repo-type=dataset --local-dir aesthetics
Aesthetic6+, Aesthetic6.25+ & Aesthetic6.5+ prompts dataset for SiD-LSG training.
Filtered from dclure/laion-aesthetics-12m-umap.
All credits to the authors of SiD.
from datasets import load_dataset
import os
dataset_dir = 'aesthetics'
# ds = load_dataset("parquet", data_dir=dataset_dir, split='train')
ds = load_dataset("aesthetics", data_files='train.parquet')
# ds =… See the full description on the dataset page: https://huggingface.co/datasets/Yuanzhi/aesthetics_prompts_laion.aesthetics_v2_4.5Ko-LAION-Aesthetics-10M
LAION-Aesthetics 10M Dataset Card
Dataset details
Dataset type:
Laion aesthetic is a subset of laion5B that has been estimated by a model trained on top of clip embeddings to be aesthetic. The intended usage of this dataset is image generation
Paper or resources for more information:
https://laion.ai/blog/laion-aesthetics/
Acknowledgements
This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grants funded by the… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/Ko-LAION-Aesthetics-10M.improved_aesthetics_6.5plus_clip_retrievallimingcv_LAION_Aesthetics_1024-sd-scripts-5000Each zip contains;
image.png/jpg/etc -> is the image
image.txt -> contains the caption
The captions are what I extracted from the json files at runtime. Hindsight says I should have kept the json but it is what it is for now.
I'll run a better one later. This one took quite a few hours as it was.
aesthetics_v2_4.75_filteredlaion_aesthetics_v2_6.0plusdeep-art-aesthetics-zh
Deep Art & Aesthetics Dialogue Dataset (Chinese)
深度艺术与审美对话数据集
Dataset Description
High-quality Chinese art and aesthetics dialogues covering aesthetic theory, visual art analysis, symbolism, and design philosophy.
高质量中文艺术与审美对话,涵盖美学理论、视觉艺术分析、象征主义、设计哲学等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata: Source platform, topic tags… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-art-aesthetics-zh.LAION_Aesthetics_1024_bucketed_512
LAION Aesthetics 1024 Bucketed 512 Captioned
This is a captioned bucketed-shards export of images from limingcv/LAION_Aesthetics_1024.
Images were filtered and resized/cropped into SDXL-style aspect-ratio buckets at a 512 base resolution, without upsampling. The export contains 382,144 images across 397 uncompressed WebDataset-style tar shards.
The .txt files now contain model-generated captions, not the original LAION web-scrape alt text or surrounding page text. Captions were… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_512.Aesthetics_X_Phone_720p_Images_Rec_Captioned_16_9
Aesthetics X Image Dataset
Overview
This dataset contains high-quality aesthetic images collected from Twitter user @aestheticsguyy. The collection features visually pleasing digital artwork, wallpapers, and photography with a focus on visual appeal and design inspiration.
Dataset Contents
• Image files in JPEG/PNG format• High-resolution wallpaper collections• Thematically organized visual content
Collection Methodology
Images were gathered from… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Aesthetics_X_Phone_720p_Images_Rec_Captioned_16_9.aestheticssemiautomatic-aestheticsaesthetic_scorelaion_aesthetics_v2_6.25plusaesthetics_6_5plusanti_aesthetics_dataset_test
