datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flux-attn-tsFLUX-Reason-6M
FLUX-Reason-6M
FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems.
This dataset contains:
6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model.
20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.fluidgym-dataflux-ffn-tsflux_generatedFluxVLAData
FluxVLA Engine 🚀
FluxVLA Engine is an integrated engineering platform designed for embodied intelligence applications. It follows the core design principles of unified configuration, standardized interfaces, module decoupling, and deployability, forming a complete engineering loop from data collection to real-world deployment. With a focus on building a "standardized industrial-academic-research foundation," FluxVLA significantly lowers the engineering threshold for VLA (Visual… See the full description on the dataset page: https://huggingface.co/datasets/limxdynamics/FluxVLAData.flux-2-klein-modelsflux.2-dev-synthetic-2M
Flux2.dev Synthetic: 2.2M Text-to-Image Pairs at 512×512
Dataset Summary
This dataset contains ~2.2 million large-scale synthetic image–caption pairs generated using the FLUX.2-dev diffusion model:
Model: black-forest-labs/FLUX.2-dev
Caption source: Text2Image-2M
Total samples: 2,282,665 (571 shards × ~4000 samples)
Resolution: 512 × 512
Image format: PNG (lossless)
Total shards: 571
Samples per shard: 4000 (last shard: 2665)
Total size: ~865 GB
Each sample consists of:… See the full description on the dataset page: https://huggingface.co/datasets/KempnerInstituteAI/flux.2-dev-synthetic-2M.flux-dev-modelsi1-fluxreason-1024-resolution-1m-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the fluxreason dataset at 1024×1024 resolution. Concretely, we only retain raw images with a shorter edge of at… See the full description on the dataset page: https://huggingface.co/datasets/i1-datasets/i1-fluxreason-1024-resolution-1m-tfrecord.i1-fluxreason-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the fluxreason dataset at 256×256 resolution.
It also serves as an example of what a dataset processed using… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/i1-fluxreason-tfrecord.musan
MUSAN: A Music, Speech, and Noise Corpus
MUSAN is a corpus of music, speech, and noise recordings designed for training models for voice activity detection and music/speech discrimination. This is a comprehensive collection suitable for various audio processing tasks.
Dataset Structure
The dataset is organized into three main categories:
1. Music (~42 hours)
Subcategories: Classical, Pop/Rock, Jazz, and more
Sources: Free Music Archive, Jamendo, and others… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/musan.datacorp-cnpj-datai1-fluxreason-512-resolution-1m-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the fluxreason dataset at 512×512 resolution. Concretely, we only retain raw images with a shorter edge of at… See the full description on the dataset page: https://huggingface.co/datasets/i1-datasets/i1-fluxreason-512-resolution-1m-tfrecord.FLUX.2-klein-base-9B_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with FLUX.2 [klein] 9B Base.
NOTE: The Base is not intended for image generation, so do not use these images to judge the quality of the model.
Base is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on FLUX.2 [klein] 9B Base
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/FLUX.2-klein-base-9B_samples_Best_of.Hunyuan3D-FLUX-Gen
Orient Anything V2 Dataset
Project Page | Paper | GitHub
Orient Anything V2 is an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. This dataset repository supports the model by providing assets for orientation estimation, 6DoF pose estimation, and object symmetry recognition.
Data Preparation
You can download the absolute orientation, relative rotation, and symm-orientation test datasets using the… See the full description on the dataset page: https://huggingface.co/datasets/Viglong/Hunyuan3D-FLUX-Gen.fluxlorafluent_speech_commands_synth
Dataset Card for "fluent_speech_commands_synth"
More Information needed
curated-danbooru-2026-512px-flux2-vaeflux_vgg50k_inv28_infer28_uncondIDTruepdm3-ht-20260528-flux2-vae-latents-public
PDM-3-HT FLUX.2 VAE latents for ImageNet-256 train
Public research artifact for PDM-3-HT VAE-backend experiments. This repository contains latent cache shards only. It intentionally does not contain raw ImageNet images, ADM-cropped uint8 images, PAE latents, PAE checkpoints, or training checkpoints.
Source and preprocessing
Source dataset: ImageNet-1k train via ILSVRC/imagenet-1k; access requires accepting the upstream ImageNet terms.
Image preprocessing before VAE… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/pdm3-ht-20260528-flux2-vae-latents-public.fleurs-full
FLEURS Full - Test Set for ASR Benchmarking
Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio.
Languages (30)
Asian Languages (13)
Code
Language
Samples
cmn_hans_cn
Chinese (Mandarin)
945
yue_hant_hk
Cantonese
819
ja_jp
Japanese
650
ko_kr
Korean
382
vi_vn
Vietnamese
857
th_th
Thai
1,021
id_id
Indonesian
687
ms_my
Malay
749
hi_in
Hindi
418
ar_eg
Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.mjnj_flux32FLUX.2-klein-base-9B_samplesThis dataset is a highly diverse set of high quality images generated with FLUX.2 [klein] 9B Base.
NOTE: The Base is not intended for image generation, so do not use these images to judge the quality of the model.
Base is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on FLUX.2 [klein] 9B Base
Quality testing
Data source
The images were created in ComfyUI… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/FLUX.2-klein-base-9B_samples.flux1-backup-202501700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.5_synt_flux_street_selected_single_validated_1011curated-danbooru-2026-256px-flux2-vaeflux1-backup-202507Sanjay_edu_fluxbase
