datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tauri-RL-Stylesstyles
Styled Image Dataset Generated with FLUX.1-dev and LoRAs from the community
Access the generation scripts here.
Dataset Description
This dataset contains 60,000 text-image-pairs. The images are generated by adding trained LoRA weights to the diffusion transformer model black-forest-labs/FLUX.1-dev. The images were created using 6 different style models, with each style having its own set of 10,000 images. Each style includes 10,000 captions sampled from the… See the full description on the dataset page: https://huggingface.co/datasets/rezashkv/styles.AIGI-styles
Midjourney Art Style Separation & Translation Dataset
This dataset consists of 192 high-quality images generated using Midjourney, carefully clustered into 12 distinct, style-separated categories. It includes 40 text prompt files corresponding to key images, detailing the exact prompts used to achieve each style's signature aesthetic.
The primary purpose of this dataset is to facilitate style translation and prompt engineering. It serves as a visual and textual training set for… See the full description on the dataset page: https://huggingface.co/datasets/DayK0n/AIGI-styles.corpus-of-diverse-styles
Dataset Card for Corpus of Diverse Styles
Disclaimer
I am not the original author of the paper that presents the Corpus of Diverse Styles. I uploaded the dataset to HuggingFace as a convenience.
Dataset Summary
A new benchmark dataset that contains 15M
sentences from 11 diverse styles.
To create CDS, we obtain data from existing academic
research datasets and public APIs or online collections
like Project Gutenberg. We choose
styles that are easy for human… See the full description on the dataset page: https://huggingface.co/datasets/billray110/corpus-of-diverse-styles.nai-shared-styles
NAI shared styles
아카라이브 AI 그림 채널의 그림체 공유 게시글에서 수집한 예제 이미지와 NovelAI 생성 메타데이터입니다. V4.5·V5 그림체를 탐색하고 원본 게시글, 프롬프트와 생성 설정을 함께 확인하기 위한 자료입니다. 성인 대상 자료를 포함합니다.
파일과 사용법
NAImakeArtistGroup_shared_images_20260914_webp.zip: 품질 90 WebP 이미지와 연결·무결성 검증용 manifest.json.
arca_style_seed.sqlite: 게시글 출처, 이미지별 프롬프트·네거티브·모델·생성 설정을 담은 메타데이터 전용 SQLite DB.
archive_info.json: ZIP 크기, 파일 수, SHA-256, 이미지 원본 대비 용량 및 검증 결과.
NAImakeArtistGroup v0.1.10 이상에서 공유 그림체 수집 → Hugging Face에서 빠르게 받기를… See the full description on the dataset page: https://huggingface.co/datasets/okawaritsuika/nai-shared-styles.ramanv-image-real-stylesQwen_Styles_Lorafashion-styles
Fashion Styles
A structured taxonomy of fashion style labels for outfit analysis, visual style classification, retrieval, and LLM-based style judging.
The dataset contains 294 canonical style records. Each record pairs a human-readable style name and description with practical recognition metadata: visual indicators, color logic, silhouettes, common contexts, cultural or regional anchors, aesthetic moods, formality, mainstreamness, temporal references, and classifier guidance.… See the full description on the dataset page: https://huggingface.co/datasets/tolgayan/fashion-styles.describe_document_styles_no_predefined_styles_test
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
intel-stylesheet-javascript
Intel Ark Frontend Assets (CSS & JS)
This dataset contains the enterprise frontend assets (Stylesheets and JavaScript files) extracted from Intel Ark (ark.intel.com).
🎯 Primary Use Case
This dataset is specifically structured for pre-training and fine-tuning AI coding assistants and web-navigating agents. By analyzing production-grade code, models can learn how modern enterprise infrastructure (like Adobe Experience Manager) maps DOM elements to CSS rules… See the full description on the dataset page: https://huggingface.co/datasets/sphita/intel-stylesheet-javascript.Sinhala-writing-stylesclusterd_authors_style_desc_no_predefined_styles
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
Tauri-RL-Styles-V2three_styles_prompted_250_512x512
Dataset Card for "three_styles_prompted_250_512x512"
More Information needed
three_styles_prompted_all_512x512
Dataset Card for "three_styles_prompted_all_512x512"
More Information needed
building_styles_datasetStyleSet
StyleSet
WARNING: This dataset contains some profane words.
A spoken language benchmark for evaluating speaking-style-related speech generationReleased in our paper, Audio-Aware Large Language Models as Judges for Speaking Styles
This dataset is released by NTU Speech Lab under the MIT license.
Tasks
Voice Style Instruction Following
Reproduce a given sentence verbatim.
Match specified prosodic styles (emotion, volume, pace, emphasis, pitch, non-verbal cues).… See the full description on the dataset page: https://huggingface.co/datasets/dcml0714/StyleSet.three_stylesvietnamese-author-styles-paraphrasedqwen_image_Illustrious_Styles_lora_neta_art_samplesthree_styles_prompted
Dataset Card for "three_styles_prompted"
More Information needed
Different-styles-of-ad-textsthree_styles_coded
Dataset Card for "three_styles_coded"
More Information needed
three_styles_10rand
Dataset Card for "three_styles_10rand"
More Information needed
anime-art-stylesthree_styles_prompted_250_512x512_50perclass_proposed
Dataset Card for "three_styles_prompted_250_512x512_50perclass_proposed"
More Information needed
three_styles_prompted_all_512x512_excluded_training
Dataset Card for "three_styles_prompted_all_512x512_excluded_training"
More Information needed
stylesthree_styles_prompted_500
Dataset Card for "three_styles_prompted_500"
More Information needed
three_styles_prompted_250_512x512_50perclass_identity
Dataset Card for "three_styles_prompted_250_512x512_50perclass_identity"
More Information needed
