datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
artist-styles
artist-styles
Static gallery of artist styles. Plain HTML/JS (index.html) reading from
data/artists.json and images/ — no build step, no dependencies.
Running with Docker
docker.sh runs the site in a python:3.12-slim container serving the project
directory with python3 scripts/server.py (a stdlib-only server: static files
plus a small favorites API). The container is named artist-styles, restarts
automatically (--restart=always), and serves on port 7803 by… See the full description on the dataset page: https://huggingface.co/datasets/jtreminio/artist-styles.indian-art-styles
🎨 Indian Art Styles Dataset
A comprehensive image classification dataset covering 34 distinct Indian painting and art styles with 27,139 images in total. This dataset is designed for training Vision Transformer (ViT) and CNN-based classifiers to recognize traditional Indian art styles.
Dataset Overview
Style
Region
Medium
Image Count
aipan
Uttarakhand
floor/wall painting
6
bengal_school
West Bengal
painting
1287
bhil
Madhya Pradesh / Rajasthan /… See the full description on the dataset page: https://huggingface.co/datasets/Divya0001/indian-art-styles.Styles_Lorasfacade-styles
Facade — synthetic architectural style corpus
1,000 generated reference plates across 20 architectural styles, with
structured attribute labels, generated style readings, and a precomputed
retrieval index.
Built for a visual-similarity search task: photograph a building, retrieve the
closest reference plates, get a reading of the style. Live app:
Facade
What this is
Every plate is generated from a sampled attribute specification — style,
massing, material, window… See the full description on the dataset page: https://huggingface.co/datasets/Jonathandav/facade-styles.artistic-styles-dataset
Professional Illustrations Dataset - 30,000 images
The dataset comprises 30,000 high-quality artistic images spanning 20 distinct artistic styles and movements. It specifically designed for advancing research in artwork generation, style transfer, and the classification of visual arts.
By leveraging this dataset, researchers and developers can push the boundaries of generating images, creating new artistic creations, and conducting aesthetic evaluations. - Get the data
The… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/artistic-styles-dataset.Tauri-RL-Stylesstyles
Styled Image Dataset Generated with FLUX.1-dev and LoRAs from the community
Access the generation scripts here.
Dataset Description
This dataset contains 60,000 text-image-pairs. The images are generated by adding trained LoRA weights to the diffusion transformer model black-forest-labs/FLUX.1-dev. The images were created using 6 different style models, with each style having its own set of 10,000 images. Each style includes 10,000 captions sampled from the… See the full description on the dataset page: https://huggingface.co/datasets/rezashkv/styles.AIGI-styles
Midjourney Art Style Separation & Translation Dataset
This dataset consists of 192 high-quality images generated using Midjourney, carefully clustered into 12 distinct, style-separated categories. It includes 40 text prompt files corresponding to key images, detailing the exact prompts used to achieve each style's signature aesthetic.
The primary purpose of this dataset is to facilitate style translation and prompt engineering. It serves as a visual and textual training set for… See the full description on the dataset page: https://huggingface.co/datasets/DayK0n/AIGI-styles.StyleTransfer-Reward-StyleScore
StyleTransfer-Reward-StyleScore
Complete Reward Model training data from StyleScore evaluation pipeline.
📊 Data Contents (~292GB total)
Core Reward Data
Archive
Size
Description
style_images.tar.part_*
123GB
Winner images (高质量风格化结果)
genref_wds_content.tar.part_*
122GB
Content images (原始内容图)
loser_images.tar
14GB
Loser images (低质量对比)
Pair Configuration
Archive
Size
Description
cnt_sty_pairs_cfg.tar
3.4GB
100k pairs… See the full description on the dataset page: https://huggingface.co/datasets/mohan2/StyleTransfer-Reward-StyleScore.corpus-of-diverse-styles
Dataset Card for Corpus of Diverse Styles
Disclaimer
I am not the original author of the paper that presents the Corpus of Diverse Styles. I uploaded the dataset to HuggingFace as a convenience.
Dataset Summary
A new benchmark dataset that contains 15M
sentences from 11 diverse styles.
To create CDS, we obtain data from existing academic
research datasets and public APIs or online collections
like Project Gutenberg. We choose
styles that are easy for human… See the full description on the dataset page: https://huggingface.co/datasets/billray110/corpus-of-diverse-styles.stylesage-imagesnsfw-prompt-pack-core-styles-wildcards
NSFW Prompt Pack Core — Styles & Wildcards
Style and wildcard pack for Pony Diffusion / Illustrious / NoobAI.Compatible with Dynamic Prompts extension and Style Organizer.
Contents
Prompt styles (CSV format)
Wildcard files
Compatible with
Pony Diffusion V6 XL
Illustrious XL
NoobAI XL
Links
CivitAI: https://civitai.com/models/2379088
Style Organizer extension: https://civitai.com/models/2393177
More: https://civitai.com/user/Nyx_x
tabi-stylesnai-shared-styles
NAI shared styles
아카라이브 AI 그림 채널의 그림체 공유 게시글에서 수집한 예제 이미지와 NovelAI 생성 메타데이터입니다. V4.5·V5 그림체를 탐색하고 원본 게시글, 프롬프트와 생성 설정을 함께 확인하기 위한 자료입니다. 성인 대상 자료를 포함합니다.
파일과 사용법
NAImakeArtistGroup_shared_images_20260914_webp.zip: 품질 90 WebP 이미지와 연결·무결성 검증용 manifest.json.
arca_style_seed.sqlite: 게시글 출처, 이미지별 프롬프트·네거티브·모델·생성 설정을 담은 메타데이터 전용 SQLite DB.
archive_info.json: ZIP 크기, 파일 수, SHA-256, 이미지 원본 대비 용량 및 검증 결과.
NAImakeArtistGroup v0.1.10 이상에서 공유 그림체 수집 → Hugging Face에서 빠르게 받기를… See the full description on the dataset page: https://huggingface.co/datasets/okawaritsuika/nai-shared-styles.ramanv-image-real-stylesQwen_Styles_Lorafashion-styles
Fashion Styles
A structured taxonomy of fashion style labels for outfit analysis, visual style classification, retrieval, and LLM-based style judging.
The dataset contains 294 canonical style records. Each record pairs a human-readable style name and description with practical recognition metadata: visual indicators, color logic, silhouettes, common contexts, cultural or regional anchors, aesthetic moods, formality, mainstreamness, temporal references, and classifier guidance.… See the full description on the dataset page: https://huggingface.co/datasets/tolgayan/fashion-styles.describe_document_styles_no_predefined_styles_test
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
intel-stylesheet-javascript
Intel Ark Frontend Assets (CSS & JS)
This dataset contains the enterprise frontend assets (Stylesheets and JavaScript files) extracted from Intel Ark (ark.intel.com).
🎯 Primary Use Case
This dataset is specifically structured for pre-training and fine-tuning AI coding assistants and web-navigating agents. By analyzing production-grade code, models can learn how modern enterprise infrastructure (like Adobe Experience Manager) maps DOM elements to CSS rules… See the full description on the dataset page: https://huggingface.co/datasets/sphita/intel-stylesheet-javascript.Sinhala-writing-stylesclusterd_authors_style_desc_no_predefined_styles
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
Tauri-RL-Styles-V2three_styles_prompted_250_512x512
Dataset Card for "three_styles_prompted_250_512x512"
More Information needed
three_styles_prompted_all_512x512
Dataset Card for "three_styles_prompted_all_512x512"
More Information needed
Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles
ID
King-TTS-241
Duration
8.56 hours
Speakers
100 People
Labeling Details
Pronunciation, Rhythm, Breath sounds marked with {hx}
Language
Chinese
Description
Two styles: Deep and uplifting; covers a variety of product categories including food, clothing, beauty, personal care, electronics, and home goods.
URL… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles.building_styles_datasetfull-html-stying-dataset-nested-element-styles
Full HTML Tailwind Nested Element Styles
kogai/full-html-stying-dataset-nested-element-styles contains phase1_dataset_10k_1_elem_in_1_elem_style.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms for simple single-element inputs that often include nested inline structure.
Schema
input_html: source HTML before Tailwind classes are added.
output_html: transformed HTML with Tailwind utility classes applied.
style_phrase: normalized… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-nested-element-styles.Painting_Stylesfull-html-stying-dataset-motion-geometry-styles
Full HTML Tailwind Motion Geometry Styles
kogai/full-html-stying-dataset-motion-geometry-styles contains phase1_dataset_10k_motion_geometry_style.jsonl, a JSONL dataset with 10000 synthetic examples. Tailwind styling transforms emphasizing motion, geometry, and expressive visual direction for full HTML pages.
Schema
input_html: source HTML before Tailwind classes are added.
output_html: transformed HTML with Tailwind utility classes applied.
style_phrase:… See the full description on the dataset page: https://huggingface.co/datasets/kogai/full-html-stying-dataset-motion-geometry-styles.41605-People-Multiple-Styles-Video-Data-Sample
Description
41605 People - Multiple Style Video Data, which includes multiple styles video data of 41605 person IDs in various environment. The person IDs in this dataset covers different skin colors like white/yellow/brown/black and different ages like young/middle-age/old-age. The resolution and duration of each video is no less than 1080p and 10s, and all videos have audio. This dataset can be used for character consistent video generation and digital human generation tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/41605-People-Multiple-Styles-Video-Data-Sample.
