datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion2b-en-aesthetic-square-human
Overview
This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset.
It has at least "a man" or "a woman" in it.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 8k images here.
Details
It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption.
I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.Harley_Street_Institute_Aesthetic_FAQ
Harley Street Institute Aesthetic FAQ — Nepali Q&A Dataset
1. Overview
This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about aesthetic medicine (एस्थेटिक मेडिसिन) — specifically aimed at doctors, nurses, dentists, and other healthcare professionals considering or building a career in medical aesthetics (botox, dermal fillers, injectable treatments, aesthetic training… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Harley_Street_Institute_Aesthetic_FAQ.Healthy_Choice_Aesthetic_Hospital_FAQ
Healthy Choice Aesthetic Hospital FAQ — Pure Nepali Q&A Dataset
1. Overview
This dataset is a pure Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about Healthy Choice Aesthetic Hospital — a real aesthetic/cosmetic and general specialty hospital located in Kathmandu, Nepal. Topics span hospital directions/location, hair transplant, skin treatments, plastic surgery, laser procedures… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Healthy_Choice_Aesthetic_Hospital_FAQ.aesthetic-corpus
InspiredHub Aesthetic Corpus — Open Layer
A structured dataset of 81,911 artworks, 11,568 books, 220 composers, and 146 philosophy pages from the world's major museums and archives.
Overview
Collection
Items
Description
Artworks
81,911
Paintings, calligraphy, sculpture, ceramics from 15+ museums
Books
11,568
Classic literature, philosophy, poetry (840 philosophy texts)
Composers
220
Classical composers with works catalog
Philosophy Pages
146… See the full description on the dataset page: https://huggingface.co/datasets/InspiredHub/aesthetic-corpus.AestheticEMOpenNiji-Dataset-Aesthetic-FinetuneUsed in quality tuning for OpenNiji
improved_aesthetics_6.5plus_clip_retrievallaion2b-en-aesthetic-square-cleaned
Overview
A subset of our opendiffusionai/laion2b-en-aesthetic-square, which is itself a subset of the widely known "Laion2b-en-aesthetic" dataset.
However, the original had only the website alt-tags for captions.
I have added decent AI captioning, via the "moondream2" model.
Additionally, I have stripped out a bunch of watermarked junk, and weeded out 20k duplicate images.
When I use it for training, I will be trimming out additional things like non-realistic images, painting… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-cleaned.deep-art-aesthetics-zh
Deep Art & Aesthetics Dialogue Dataset (Chinese)
深度艺术与审美对话数据集
Dataset Description
High-quality Chinese art and aesthetics dialogues covering aesthetic theory, visual art analysis, symbolism, and design philosophy.
高质量中文艺术与审美对话,涵盖美学理论、视觉艺术分析、象征主义、设计哲学等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata: Source platform, topic tags… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-art-aesthetics-zh.opentts-uk-aesthetics
Aesthetics of Open Text-to-Speech for 🇺🇦 Ukrainian dataset
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
What is it?
This dataset contains metrics for https://huggingface.co/datasets/Yehor/opentts-uk dataset retrieved by https://github.com/facebookresearch/audiobox-aesthetics
How metrics calculated?
You can find a… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/opentts-uk-aesthetics.Audio-aesthetics-scoreramanv-aesthetic-scoreslaion2b-en-aesthetic-square
Contents
This is a pretty raw filter of https://huggingface.co/datasets/laion/laion2B-en-aesthetic
I just filtered for "is image perfectly square, AND is image at least 1024x1024 pixels"
Approximate image count is a bit over 300k
Updated 2025/01/24
I just found out there are a bunch of watermarked sites in here.
So much for aesthetically chosen :(
So I filtered out a bunch of the "stock image" sites, just by looking at url strings.
Update 2025/01/31… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square.
