CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LucasFang /FLUX-Reason-6M FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative models. This dataset was created to bridge the performance gap between open-source and leading closed-source text-to-image systems. This dataset contains: 6 million high-quality, reasoning-focused images synthesized by the state-of-the-art FLUX.1-dev model. 20 million bilingual (English and Chinese) descriptions, providing a rich… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/FLUX-Reason-6M.image1M<n<10M113 likes9.8k downloads8mo agoHugging Face02lehduong /flux_generatedgatedimage1M<n<10M9 likes4.4k downloads1y agoHugging Face03fluowai /datacorp-cnpj-datatext10M<n<100M2 likes1.6k downloads6mo agoHugging Face04stablellama /FLUX.2-klein-base-9B_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with FLUX.2 [klein] 9B Base. NOTE: The Base is not intended for image generation, so do not use these images to judge the quality of the model. Base is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on FLUX.2 [klein] 9B Base Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/FLUX.2-klein-base-9B_samples_Best_of.texttext-to-image1K<n<10K2 likes1.4k downloads8mo agoHugging Face05Viglong /Hunyuan3D-FLUX-Gen Orient Anything V2 Dataset Project Page | Paper | GitHub Orient Anything V2 is an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. This dataset repository supports the model by providing assets for orientation estimation, 6DoF pose estimation, and object symmetry recognition. Data Preparation You can download the absolute orientation, relative rotation, and symm-orientation test datasets using the… See the full description on the dataset page: https://huggingface.co/datasets/Viglong/Hunyuan3D-FLUX-Gen.textother100K<n<1M2 likes1.4k downloads9mo agoHugging Face06hvai /fluxloraimagen<1K2 likes1.4k downloads1y agoHugging Face07Codec-SUPERB /fluent_speech_commands_synth Dataset Card for "fluent_speech_commands_synth" More Information needed audio100K<n<1M1 likes1.4k downloads3y agoHugging Face08aipracticecafe /curated-danbooru-2026-512px-flux2-vaetabular100K<n<1M0 likes1.3k downloads25d agoHugging Face09recoilme /mjnj_flux32tabular100K<n<1M0 likes1.1k downloads7mo agoHugging Face10stablellama /FLUX.2-klein-base-9B_samplesThis dataset is a highly diverse set of high quality images generated with FLUX.2 [klein] 9B Base. NOTE: The Base is not intended for image generation, so do not use these images to judge the quality of the model. Base is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on FLUX.2 [klein] 9B Base Quality testing Data source The images were created in ComfyUI… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/FLUX.2-klein-base-9B_samples.texttext-to-image1K<n<10K1 likes1.1k downloads8mo agoHugging Face11Rapidata /700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3 NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset Rapidata Image Generation Preference Dataset This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment. Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.imagetext-to-image10K<n<100K20 likes867 downloads2y agoHugging Face12aipracticecafe-mirror /curated-danbooru-2026-256px-flux2-vaetabular100K<n<1M0 likes820 downloads29d agoHugging Face13FluidInference /THCHS-30-tests THCHS-30 Test Set THCHS-30 test split for Mandarin Chinese speech recognition benchmarking. Dataset Info Language: Mandarin Chinese (zh-CN) Samples: 2,495 Speakers: 10 Sample Rate: 16 kHz License: Apache 2.0 Usage from datasets import load_dataset # After uploading to HuggingFace dataset = load_dataset("your-username/thchs30-test") # Example print(dataset['train'][0]) # { # 'audio': {'array': [...], 'sampling_rate': 16000, 'path': 'audio/D11_750.wav'}, #… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/THCHS-30-tests.audio1K<n<10K0 likes723 downloads6mo agoHugging Face14Rapidata /Flux_SD3_MJ_Dalle_Human_Alignment_Dataset NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset Rapidata Image Generation Alignment Dataset This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment. Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.imagetext-to-image10K<n<100K16 likes694 downloads2y agoHugging Face15FluidInference /ami-corpus-mirror AMI Corpus Mirror Mirror of the subset of the AMI Meeting Corpus used by FluidAudio diarization benchmarks. Hosted here so CI and local benchmark runs do not depend on the availability of the upstream groups.inf.ed.ac.uk server (see FluidAudio#752). Contents annotations/ami_public_manual_1.6.2.zip — AMI public manual annotations v1.6.2 (repackaged from the official archive; identical content, including segments/, words/, corpusResources/meetings.xml)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/ami-corpus-mirror.audio1K<n<10K0 likes680 downloads3mo agoHugging Face16jskresearch /HUGGER-Unified-Gravity-Fluid-Framework 🌍 H.U.G.G.E.R: Heuristic Universal Grid & Gravity Equilibrium Rendering Tensor This repository serves as an open academic archive and tensor-specification benchmark for generalized tensor standards, designed to resolve non-linear computational collapse and topological pole singularities in high-performance CFD and planetary atmospheric models. It acts as the Macroscopic Gravitational Backbone, perfectly entangled with the microscopic Topological Zero Tensor (TZT)… See the full description on the dataset page: https://huggingface.co/datasets/jskresearch/HUGGER-Unified-Gravity-Fluid-Framework.documentroboticsn<1K1 likes551 downloads4h agoHugging Face17Rapidata /Flux-2-pro_t2i_human_preference Rapidata Flux 2 Pro Preference This T2I dataset contains over ~400'000 human responses from over ~50'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Flux 2 Pro (version from 25.11.25) across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux-2-pro_t2i_human_preference.imagetext-to-image10K<n<100K15 likes548 downloads10mo agoHugging Face18Flux9665 /BibleMMSThe Dataset associated with the Paper "Meta Learning Text-to-Speech Synthesis in over 7000 Languages" by Florian Lux, Sarina Meyer, Lyonel Behringer, Frank Zalkow, Phat Do, Matt Coler, Emanuël A. P. Habets and Ngoc Thang Vu (Interspeech 2024). We generate 2000 spoken utterances per language using the subsets of the eBible dataset [1] that are under free licenses as the text input to the MMS TTS models [2]. The languages associated with the following ISO-639-3 codes are represented in this… See the full description on the dataset page: https://huggingface.co/datasets/Flux9665/BibleMMS.audiotext-to-speech100K<n<1M82 likes536 downloads2y agoHugging Face19flunardelli /mmlu Dataset Card for MMLU Dataset Summary Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021). This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks… See the full description on the dataset page: https://huggingface.co/datasets/flunardelli/mmlu.textquestion-answering100K<n<1M0 likes531 downloads2y agoHugging Face20k-mktr /improved-flux-prompts-photoreal-portrait Photo Portrait Prompt Dataset for FLUX Overview This dataset contains a curated collection of prompts specifically designed for generating photo portraits using FLUX.1, an advanced text-to-image model. These prompts are crafted to produce high-quality, lifelike portraits by leveraging sophisticated prompting techniques and best practices. Latest Version Improved on October 3, 2024. This version has undergone curation and improvement. What is new? Cleaned up… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/improved-flux-prompts-photoreal-portrait.texttext-classification10K<n<100K115 likes502 downloads2y agoHugging Face21AbstractPhil /flux-schnell-teacher-latents Flux Schnell Teacher Latents Pre-computed latents, decoded images, and text embeddings from FLUX.1-schnell for distillation and research. Usage from datasets import load_dataset # Load specific subset ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_512") ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_2_512") ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_3_512") Subsets Config… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-schnell-teacher-latents.imageimage-to-image100K<n<1M0 likes470 downloads8mo agoHugging Face22juliensimon /goes-xray-flux GOES Solar X-Ray Flux (1-Minute) Credit: NASA/SDO Part of a dataset collection on Hugging Face. Dataset description Solar soft X-ray flux from the GOES X-Ray Sensor (XRS), the operational backbone of solar flare monitoring. Updated daily from NOAA SWPC, growing incrementally at 1-minute cadence. The GOES (Geostationary Operational Environmental Satellite) X-Ray Sensor measures the Sun's soft X-ray irradiance in two wavelength bands: a "short" 0.05-0.4 nm… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/goes-xray-flux.tabulartime-series-forecasting100K<n<1M0 likes459 downloads21h agoHugging Face23CodecSR /fluent_speech_commands_femaleaudio10K<n<100K1 likes445 downloads2y agoHugging Face24proteinglm /fluorescence_prediction Dataset Card for Fluorescence Prediction Dataset Dataset Summary The Fluorescence Prediction task focuses on predicting the fluorescence intensity of green fluorescent protein mutants, a crucial function in biology that allows researchers to infer the presence of proteins within cell lines and living organisms. This regression task utilizes training and evaluation datasets that feature mutants with three or fewer mutations, contrasting the testing dataset, which comprises… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/fluorescence_prediction.texttext-classification10K<n<100K0 likes403 downloads2y agoHugging Face25retowyss /Persona-Fluxed-10k-2608 Persona Fluxed 10k Synthetic persona portraits rendered with FLUX.2-klein-4b (8-step, 1024x1024) from the NVIDIA Nemotron-Personas-* datasets. Each persona is grounded in real-world demographic, geographic and personality-trait distributions for its country (CC BY 4.0 source; no real people). Currently Nemotron-Personas exist for: USA — English Japan — Japanese India — English, Hindi Brazil — Portuguese Singapore — English France — French Korea — Korean El Salvador — Spanish… See the full description on the dataset page: https://huggingface.co/datasets/retowyss/Persona-Fluxed-10k-2608.imagetext-to-image10K<n<100K0 likes400 downloads1mo agoHugging Face26Rapidata /Flux_SD3_MJ_Dalle_Human_Coherence_Dataset NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset Rapidata Image Generation Coherence Dataset This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment. Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3 Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.imagequestion-answering10K<n<100K12 likes389 downloads2y agoHugging Face27obvious-research /flux-kontext-ipa-datasetimage10K<n<100K2 likes387 downloads1y agoHugging Face28kadirnar /fluxdev_controlnet_16kimage10K<n<100K30 likes368 downloads2y agoHugging Face29fluxae /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of… See the full description on the dataset page: https://huggingface.co/datasets/fluxae/Time-Series-Library.tabulartime-series-forecasting1M<n<10M0 likes329 downloads2mo agoHugging Face30cmudrc /OpenSeeSimE-Fluid OpenSeeSimE-Fluid: Engineering Simulation Visual Question Answering Benchmark Dataset Summary OpenSeeSimE-Fluid is a large-scale benchmark dataset for evaluating vision-language models on computational fluid dynamics (CFD) simulation interpretation tasks. It contains approximately 98,000 question-answer pairs across parametrically-varied fluid simulations including turbulent flow, heat transfer, and complex flow patterns. Purpose While vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/OpenSeeSimE-Fluid.imagevisual-question-answering10K<n<100K0 likes320 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.