datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FLUX-Reason-6M
FLUX-Reason-6M
FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems.
This dataset contains:
6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model.
20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.flux_generatedfluxloraflux_vgg50k_inv28_infer28_uncondIDTrue700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.5_synt_flux_street_selected_single_validated_1011Flux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.flux_sae_imagesFlux-2-pro_t2i_human_preference
Rapidata Flux 2 Pro Preference
This T2I dataset contains over ~400'000 human responses from over ~50'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Flux 2 Pro (version from 25.11.25) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux-2-pro_t2i_human_preference.flux-schnell-teacher-latents
Flux Schnell Teacher Latents
Pre-computed latents, decoded images, and text embeddings from FLUX.1-schnell for distillation and research.
Usage
from datasets import load_dataset
# Load specific subset
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_512")
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_2_512")
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_3_512")
Subsets
Config… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-schnell-teacher-latents.flux_datasetRSGPNet
Citation
If you find it useful, please consider citing:
@article{wang2026rsgpnet,
title={RSGPNet: Geometric Prompting for Remote Sensing Open-Vocabulary Semantic Segmentation},
author={Wang, Shanwen and Sun, Xin and Wang, Sirui and Zhu, Xiao Xiang},
journal={arXiv preprint arXiv:2606.28410},
year={2026}
}
Acknowledgments
We sincerely thank the authors of SAM3 and SegEarth‑OV3 for their excellent open‑source work, and we also thank the contributors of the… See the full description on the dataset page: https://huggingface.co/datasets/fluorites/RSGPNet.flux_generationsPersona-Fluxed-10k-2608
Persona Fluxed 10k
Synthetic persona portraits rendered with FLUX.2-klein-4b (8-step, 1024x1024) from the NVIDIA Nemotron-Personas-* datasets.
Each persona is grounded in real-world demographic, geographic and
personality-trait distributions for its country (CC BY 4.0 source; no real people).
Currently Nemotron-Personas exist for:
USA — English
Japan — Japanese
India — English, Hindi
Brazil — Portuguese
Singapore — English
France — French
Korea — Korean
El Salvador — Spanish… See the full description on the dataset page: https://huggingface.co/datasets/retowyss/Persona-Fluxed-10k-2608.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.flux-kontext-ipa-datasetFlux2-Image
Flux2-Image
This dataset is used for the Flux2-from-scratch project to train the Flux2 transformer from scratch.
Dataset structure
.
├── images_1/
│ └── a4ad6f1d46fb4ef3bcda56f05eacc537.png...
├── images_2/
│ └── 000121ce53464a63ad43ff134e979c7a.png...
├── images_3/
│ └── a47dc3302bb64b2ebfc7d525c8f0a88e.png...
├── images_4/
│ └── 0053ea2a8c764123b263d88c217f2995.png...
├── images_5/
│ └── 23fb74d40d8e4635914b4a399ee27b71.png...
└── data.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/Flux2-Image.fluxdev_controlnet_16klaion_fluxshchnell_generatedOpenSeeSimE-Fluid
OpenSeeSimE-Fluid: Engineering Simulation Visual Question Answering Benchmark
Dataset Summary
OpenSeeSimE-Fluid is a large-scale benchmark dataset for evaluating vision-language models on computational fluid dynamics (CFD) simulation interpretation tasks. It contains approximately 98,000 question-answer pairs across parametrically-varied fluid simulations including turbulent flow, heat transfer, and complex flow patterns.
Purpose
While vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid.Synthetic-Character-Dataset-for-FLUX-LORA
Synthetic Human Character Dataset (Flux LoRA)
Dataset Overview
This dataset is a fully synthetic human character dataset generated entirely using AI pipelines. It contains no real individuals and is not derived from any real-world person or identity.
The dataset is intended for Flux LoRA training and related diffusion-based character modeling tasks.
Total images: 42
Format: PNG
All images are fully captioned
Captions are stored in a single file where each caption… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/Synthetic-Character-Dataset-for-FLUX-LORA.Flux-workflows
1. FLUX Installation Tutorial on ComfyUI/ForgeUI (Native/GGUF/NF4 variant): 👇
https://www.stablediffusiontutorials.com/2024/08/flux-installation.html
2. Flux Kontext Installation + Workflow (Native/GGUF):👇
https://www.stablediffusiontutorials.com/2025/06/flux-kontext.html
3. Flux Kontext Dev LoRA Training on Windows/Linux:👇
https://www.stablediffusiontutorials.com/2025/06/train-flux-kontext-lora.html
4. Train Flux Dev/Schnell LoRA on Windows/Linux:… See the full description on the dataset page: https://huggingface.co/datasets/stablediffusiontutorials/Flux-workflows.OpenSeeSimE-Fluid-Small
OpenSeeSimE-Fluid-Small
A stratified 10% subset of cmudrc/OpenSeeSimE-Fluid for evaluating vision-language models at a reduced compute footprint while preserving the joint distribution of simulation type, question type, media type, and question id.
Subset Provenance
Parent dataset: cmudrc/OpenSeeSimE-Fluid (98,326 rows total)
Rows in this subset: 9,881 (10.05% of parent)
Source classes: Bent Pipe, Converging Nozzle, Heat Exchanger, Heat Sink, Mixing Pipe
Parquet shards:… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid-Small.5_synt_flux_street_selected_multi_v1_10225_synt_flux_street_selectedflux_pet_costume5_synt_flux_landmark_missing_imgs_validated117k_human_coherence_flux1.0_V_flux1.1Blueberry
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 340k human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/117k_human_preferences_flux1.0_V_flux1.1Blueberry
Link to the Text-2-Image Alignment dataset: https://huggingface.co/datasets/Rapidata/117k_human_alignment_flux1.0_V_flux1.1Blueberry
It was collected in ~2 Days using the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/117k_human_coherence_flux1.0_V_flux1.1Blueberry.flux_10k_captions
Dataset Card for "flux_10k_captions"
More Information needed
restoredit_fluxmail me @akshankrithick305@gmail.com for damage generation code
RestorEdit FLUX
Synthetic photo restoration dataset with 6 damage types.
Column
Description
original
Real damaged photo
clean
FLUX-generated clean version
scratches
v0: Opaque white/silver lines
crevices
v1: Cream paper tears
blob_tears
v2: Soft-edge tears
blur
v3: Gaussian/motion blur
burns
v4: Dark charred areas
jpeg
v5: Compression artifacts
