CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Vchitect /Vchitect_T2V_DataVerse Vchitect-T2V-Dataverse Vchitect Team1  1Shanghai Artificial Intelligence Laboratory  Paper | Project Page | Data Overview The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models. It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.texttext-to-video1M<n<10M11 likes53k downloads2y agoHugging Face02ma-xu /fine-t2i Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv] by Xu Ma, Yitian Zhang, Qihua Dong, Yun Fu Northeastern Univeristy Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient). 🆕 What's New [2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️ [2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.imageimage-to-text100K<n<1M120 likes19k downloads7mo agoHugging Face03lioooox /T2I-CoReBench-Images T2I-CoReBench-Images 📖 Overview T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities. This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images.imagetext-to-image10K<n<100K5 likes10k downloads7mo agoHugging Face04arijitghosh /T2I-ImageNet-Normalimage1M<n<10M3 likes1.3k downloads1y agoHugging Face05NilanE /Vchitect_T2V_DataVerse_256p_8fps_wdshttps://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse resampled to 256p. Intended for training https://github.com/NilanEkanayake/TiTok-Video text100K<n<1M0 likes797 downloads1y agoHugging Face06CodeGoat24 /UnifiedReward-2.0-T2X-score-data Dataset Summary UnifiedReward-2.0-T2X-score-data is added for our UnifiedReward-2.0-qwen-[3b/7b/32b/72b] training. This dataset enables UnifiedReward-2.0 introducing several new capabilities: Pairwise scoring for image and video generation assessment on Alignment, Coherence, Style dimensions. Pointwise scoring for image and video generation assessment on Alignment, Coherence/Physics, Style dimensions. Welcome to try the latest version, and the inference code is available at… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-2.0-T2X-score-data.image100K<n<1M0 likes761 downloads1y agoHugging Face07nyu-dice-lab /imagenetpp-laion-t2iDataset Card for ImageNet++'s LAION Text-to-Image Split image100K<n<1M0 likes231 downloads2y agoHugging Face08arijitghosh /T2I-ImageNet-CutMiximage1M<n<10M0 likes178 downloads1y agoHugging Face09SPRIGHT-T2I /spright_coco Dataset Description SPRIGHT (SPatially RIGHT) is the first spatially focused, large scale vision-language dataset. It was built by re-captioning ∼6 million images from 4 widely-used datasets: CC12M Segment Anything COCO Validation LAION Aesthetics This repository contains the re-captioned data from COCO-Validation Set, while the data from CC12 and Segment Anything is present here. We do not release images from LAION, as the parent images are currently private. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SPRIGHT-T2I/spright_coco.image10K<n<100K5 likes164 downloads2y agoHugging Face10Higobeatz /t2adata3audio10K<n<100K0 likes122 downloads2y agoHugging Face11SPRIGHT-T2I /18_obj_444 Dataset Description This dataset contains the 444 images that we used for training our model - https://huggingface.co/SPRIGHT-T2I/spright-t2i-sd2. This contains the samples of this subset related to the Segment Anything images. We will release the LAION images, when the parent images are made public again. Our training and validation set are a subset of the SPRIGHT dataset, and consists of 444 and 50 images respectively, randomly sampled in a 50:50 split between LAION-Aesthetics and… See the full description on the dataset page: https://huggingface.co/datasets/SPRIGHT-T2I/18_obj_444.imagen<1K1 likes46 downloads2y agoHugging Face12jayw /t2v-gen-evaltextn<1K4 likes29 downloads3y agoHugging Face13kexul /MDM_t2mtextn<1K0 likes29 downloads2y agoHugging Face14nyu-dice-lab /imagenetpp-laionnet-t2iimage100K<n<1M0 likes16 downloads2y agoHugging Face15zsh12787 /t2i_dataimage100K<n<1M0 likes16 downloads11mo agoHugging Face16Sakeoffellow001 /T2i_Factualbenchimage1K<n<10K1 likes10 downloads1y agoHugging Face17Jayce-Ping /T2IS-dataimage10K<n<100K0 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.