datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_instructions_122k_alpaca_stylemultilingual-pl-bertAttribution: Wikipedia.org
dolly-15k-oai-style
Dataset Card for "dolly-15k-oai-style"
More Information needed
artist-styles
artist-styles
Static gallery of artist styles. Plain HTML/JS (index.html) reading from
data/artists.json and images/ — no build step, no dependencies.
Running with Docker
docker.sh runs the site in a python:3.12-slim container serving the project
directory with python3 scripts/server.py (a stdlib-only server: static files
plus a small favorites API). The container is named artist-styles, restarts
automatically (--restart=always), and serves on port 7803 by… See the full description on the dataset page: https://huggingface.co/datasets/jtreminio/artist-styles.guanaco-sharegpt-style
Dataset Card for "guanaco-sharegpt-style"
More Information needed
style-dpo
gijl style dataset (multi-type)
Generated by meta-models/Muse-Glimmer-30B through a tool-using scouting loop over real sources (Stack Exchange, GitHub, OSV, Hacker News, arXiv, Wikipedia, web). Splits are a deterministic hash of the record id (90/5/5); derived records inherit their parent's split. Synthetic, model-written, not human-verified. Every rejected response is intentionally poor and must never be used as an example of good behavior.
config
folder
train
validation… See the full description on the dataset page: https://huggingface.co/datasets/gijl/style-dpo.StyleTransferData
LUMIC Dataset
This is the dataset fpr LUMIC (insert github link/paper)
Dataset Details
This dataset consists of 2 different datasets. The first is the JUMP Pilot Dataset (link), which we subset by only using the 24H treatment group; the "Style Transfer" dataset was collected for this project and consists of 5 different cell type (HeLa, A549, HEK293T, 3T3, RPTE) treated with 61 different compounds.
All of the images have already been preprocessed using sklearn's… See the full description on the dataset page: https://huggingface.co/datasets/azhung/StyleTransferData.style-transfered-2013descriptiveness-sentiment-trl-style
TRL's Sentiment and Descriptiveness Preference Dataset
The dataset comes from https://arxiv.org/abs/1909.08593, one of the earliest RLHF work from OpenAI.
We preprocess the dataset using our standard prompt, chosen, rejected format.
Reproduce this dataset
Download the descriptiveness_sentiment.py from the https://huggingface.co/datasets/trl-internal-testing/descriptiveness-sentiment-trl-style/tree/0.1.0.
Run python examples/datasets/descriptiveness_sentiment.py… See the full description on the dataset page: https://huggingface.co/datasets/trl-internal-testing/descriptiveness-sentiment-trl-style.GarmageSet
GarmageSet
GarmageSet is a large-scale professionally-curated garment dataset introduced in the paper GarmageNet: A Multimodal Generative Framework for Sewing Pattern Design and Generic Garment Modeling.
It is built to support training and evaluation of methods that automate the creation of 2D sewing patterns, the construction of sewing relationships, and the synthesis of 3D garment initializations compatible with physics-based simulation.
This dataset facilitates versatile… See the full description on the dataset page: https://huggingface.co/datasets/Style3D/GarmageSet.indian-art-styles
🎨 Indian Art Styles Dataset
A comprehensive image classification dataset covering 34 distinct Indian painting and art styles with 27,139 images in total. This dataset is designed for training Vision Transformer (ViT) and CNN-based classifiers to recognize traditional Indian art styles.
Dataset Overview
Style
Region
Medium
Image Count
aipan
Uttarakhand
floor/wall painting
6
bengal_school
West Bengal
painting
1287
bhil
Madhya Pradesh / Rajasthan /… See the full description on the dataset page: https://huggingface.co/datasets/Divya0001/indian-art-styles.ramanv-image-real-style-editorialSADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation
Dataset Card for Dataset SADC
There is evidence that the driving style of an
autonomous vehicle is important to increase the acceptance
and trust of the passengers. The driving situation has been
found to have a significant influence on human driving behavior.
However, current driving style models only partially incorporate
driving environment information, limiting the alignment between
an agent and the given situation.
Therefore, we propose a dataset for situation-aware… See the full description on the dataset page: https://huggingface.co/datasets/jHaselberger/SADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation.Browsecomp-stylenepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.nn-3k-unify-styleGPT_Watercolor_Anime_Style_Images
GPT Watercolor Anime Style Images
Dataset Description
This is a synthetic GPT-generated Watercolor Anime Style image dataset. It contains 120 image-caption pairs with transparent watercolor washes, soft ink linework, pale paper texture, muted colors, and traditional anime illustration scenes.
The images focus on soft watercolor washes, expressive linework, gentle lighting, quiet interiors, nature scenes, village streets, character studies, and calm storybook… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Watercolor_Anime_Style_Images.StyleGender-Datasetfree_style_lora_meta
Free Style LoRA Meta
This dataset contains metadata and demo images for LoRA (Low-Rank Adaptation) models evaluated across three base model architectures. It serves as a reference for understanding LoRA training quality, visual style/content characteristics, and evaluation configurations.
Data Structure
free_style_lora_meta/
├── flux/ # FLUX-based LoRA evaluations
│ └── {lora_id}/
│ ├── {lora_id}.json # Main metadata… See the full description on the dataset page: https://huggingface.co/datasets/Blue2Giant/free_style_lora_meta.multilingual-phonemes-10k-alpha
Multilingual Phonemes 10K Alpha
This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows.
Languages
We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.GPT_Monet_Style_Images
GPT Monet Style Images
Dataset Description
This is a synthetic GPT-generated Monet-style image dataset. It contains 100 image-caption pairs with impressionist lighting, soft broken color, painterly atmosphere, and garden, water, street, interior, and still-life compositions.
The images focus on luminous outdoor light, loose brushwork, atmospheric color, reflective water, flowers, fields, cozy scenes, cafes, village streets, and calm impressionist subjects.… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Monet_Style_Images.Styles_Lorasolympiad_style_integer_math_problems
Olympiad Math Corpus
Version: v2.1.1
Release date: 2026-05-03
59,486 synthetically generated olympiad-style math problems with verified integer answers and formal computation graphs.
Loading
from datasets import load_dataset
ds = load_dataset("mihailgribov/olympiad_style_integer_math_problems", split="train")
lemma_applicability is stored as list[{lemma, status}] rather than a sparse dict (required for Arrow-based consumers). To convert to a dict for local use:… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_problems.mindbotz-style-sheets-v2
MINDBOTZ Style Sheets v2
20 professional character style/reference sheets, generated end-to-end by the Sonic Forage pipeline: an on-device prompt-expander LLM (Qwen3-VL inside Krea2) + Krea2 Turbo (int8) on a consumer RTX 4070, orchestrated remotely via SSH tunnel + ComfyUI API.
Every sheet follows a professional concept-art layout: character turnaround (front/side/¾), expression callout row, color palette swatch strip, props, clean background — each in a different art medium… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/mindbotz-style-sheets-v2.nebius__SWE-bench-extra__style-2__fs-oraclestyleSADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation
Dataset Card for Dataset SADC
There is evidence that the driving style of an
autonomous vehicle is important to increase the acceptance
and trust of the passengers. The driving situation has been
found to have a significant influence on human driving behavior.
However, current driving style models only partially incorporate
driving environment information, limiting the alignment between
an agent and the given situation.
Therefore, we propose a dataset for situation-aware driving… See the full description on the dataset page: https://huggingface.co/datasets/zzqasdfsdf/SADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation.muril-nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.human-style-preferences-images
Rapidata Image Generation Preference Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human preference datasets for text-to-image models, this release contains over 1,200,000 human preference… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-style-preferences-images.OmniStyle-150k
OmniStyle-150K Dataset
OmniStyle-150K is a high-quality triplet dataset specifically designed to support generalizable, controllable, and high-resolution image style transfer. Each triplet includes a content image, a style reference image, and the corresponding stylized result.
📦 Dataset Structure
OmniStyle-150K/: Stylized result images
content/: Original content images
style/: Style reference images
Each file in the OmniStyle-150K/ folder is named using the… See the full description on the dataset page: https://huggingface.co/datasets/StyleXX/OmniStyle-150k.
