CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gasstation /gs-images-v2image100K<n<1M1 likes7.8k downloads8mo agoHugging Face02UCSC-VLAA /GPT-Image-Edit-1.5M GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset 📃Arxiv | 🌐 Project Page | 💻Github GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1. 📣 News [2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download. [2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imageimage-to-image1M<n<10M90 likes6.8k downloads1y agoHugging Face03gaunernst /glint360k-wds-gz Glint360K This dataset is introduced in the Partial FC paper https://arxiv.org/abs/2010.05222. There are 17,091,657 images and 360,232 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112. This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_. The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 1,385 shards in total.… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/glint360k-wds-gz.imageimage-classification10M<n<100M5 likes5.6k downloads2y agoHugging Face04gdsu /csd_filesimagen<1K0 likes4k downloads1y agoHugging Face05gaunernst /ffhq-1024-wds Flickr-Faces-HQ Dataset (FFHQ) - 1024x1024 This is a reupload of FFHQ-1024. Refer to the original dataset repo for more information https://github.com/NVlabs/ffhq-dataset Specifically, this is the images1024x1024 set - faces are aligned and cropped to 1024x1024. Original PNG files were transcoded to WEBP losslessly to save space and packed to WebDataset format for ease of streaming. The original filenames are kept (with different file extension) so that you can match against… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ffhq-1024-wds.image10K<n<100K1 likes3.8k downloads2y agoHugging Face06ll-13 /GLH-BridgeThe GLH-Bridge dataset is a large-scale dataset for bridge detection in large-size VHR remote sensing images. More details can be found in Learning to Holistically Detect Bridges from Large-Size VHR Remote Sensing Imagery (TPAMI2024), paper link: https://ieeexplore.ieee.org/document/10509806 . image1K<n<10K4 likes2.6k downloads2y agoHugging Face07gaunernst /ms1mv3-wds MS-Celeb-1M (v3) This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper. There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112. This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds.imageimage-classification100K<n<1M0 likes1.9k downloads2y agoHugging Face08blowing-up-groundhogs /font-square-pretrain-20M 📚 Citation If you use this dataset in your research, please cite these papers: @article{pippi2023evaluating, title={Evaluating Synthetic Pre-Training for Handwriting Processing Tasks}, author={Pippi, Vittorio and Cascianelli, Silvia and Baraldi, Lorenzo and Cucchiara, Rita}, journal={Pattern Recognition Letters}, year={2023}, publisher={Elsevier} } @InProceedings{pippi2025zeroshot, author = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-pretrain-20M.image10M<n<100M0 likes1.8k downloads6mo agoHugging Face09yayoimizuha /Glint360k Dataset Card for Glint360K Citiation by InsightFace Repository We clean, merge, and release the largest and cleanest face recognition dataset Glint360K, which contains 17091657 images of 360232 individuals. By employing the Patial FC training strategy, baseline models trained on Glint360K can easily achieve state-of-the-art performance. Detailed evaluation results on the large-scale test set (e.g. IFRT, IJB-C and Megaface) are as follows: Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yayoimizuha/Glint360k.imageimage-feature-extraction10M<n<100M5 likes1.6k downloads1y agoHugging Face10GoldenCity /dan-webp-newimage1M<n<10M0 likes1.4k downloads1y agoHugging Face11blowing-up-groundhogs /font-square-v2 Accessing the font-square-v2 Dataset on Hugging Face The font-square-v2 dataset is hosted on Hugging Face at blowing-up-groundhogs/font-square-v2. It is stored in WebDataset format, with tar files organized as follows: tars/train/: Contains {000..499}.tar shards for the main training split. tars/fine_tune/: Contains {000..049}.tar shards for fine-tuning. Each tar file contains multiple samples, where each sample includes: An RGB image (.rgb.png) A black-and-white image (.bw.png)… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-v2.image1M<n<10M6 likes1.1k downloads1y agoHugging Face12jirong /grit_2mimage100K<n<1M2 likes908 downloads3y agoHugging Face13paralym /mint-1t-html-images-gte6-sample Size: 6769158 images sampled from Mint-1t-html Criteria: Data entries with greater than or equal to 6 images (gte6) image1M<n<10M0 likes871 downloads2y agoHugging Face14gdsu /sdxl_images_easy_prompts-artists-seed1image10K<n<100K0 likes869 downloads2y agoHugging Face15diffusion-cot /GenRef-wds GenRef-1M We provide 1M high-quality triplets of the form (flawed image, high-quality image, reflection) collected across multiple domains using our scalable pipeline from [1]. We used this dataset to train our reflection tuning model. To know the details of the dataset creation pipeline, please refer to Section 3.2 of [1]. Project Page: https://diffusion-cot.github.io/reflection2perfection Dataset loading We provide the dataset in the webdataset format for fast… See the full description on the dataset page: https://huggingface.co/datasets/diffusion-cot/GenRef-wds.imagetext-to-image1M<n<10M15 likes831 downloads1y agoHugging Face16roitberg-group /testdataimagen<1K0 likes636 downloads1y agoHugging Face17liaolw /ObjaverseXL_github_rendersimage1M<n<10M0 likes633 downloads9mo agoHugging Face18haifan-gong /PPIRD PPIRD: Patent-Product Image Retrieval Dataset PPIRD is the dataset released with the NeurIPS 2025 paper: Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval PPIRD is designed for Patent-Product Image Retrieval (PPIR), where a model retrieves relevant patent images from a large patent gallery given a product image query. This setting is useful for studying patent infringement search, open-set image retrieval, cross-domain visual matching, and… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/PPIRD.imageimage-to-image10M<n<100M0 likes610 downloads4mo agoHugging Face19yann111 /GlobalGeoTree GlobalGeoTree Dataset GlobalGeoTree is a comprehensive global dataset for tree species classification, comprising 6.3 million geolocated tree occurrences spanning 275 families, 2,734 genera, and 21,001 species across hierarchical taxonomic levels. Each sample is paired with Sentinel-2 image time series and 27 auxiliary environmental variables. Dataset Structure This repository contains three main components: 1. GlobalGeoTree-6M Training dataset with around 6M… See the full description on the dataset page: https://huggingface.co/datasets/yann111/GlobalGeoTree.text1M<n<10M14 likes572 downloads11mo agoHugging Face20LLLebin /GameIR GameIR Image restoration techniques such as super-resolution and image synthesis are used in products like NVIDIA's DLSS but are less understood by the public when applied to gaming. This is due to a shortage of relevant ground-truth training data for gaming, which differs from typical content with its distinct, sharp low-resolution images. In this case, we develop GameIR, a large-scale high-quality computer-synthesized ground-truth dataset to fill in the blanks, targeting at 2… See the full description on the dataset page: https://huggingface.co/datasets/LLLebin/GameIR.imageimage-to-image100K<n<1M1 likes547 downloads2y agoHugging Face21gaunernst /webface4m-wds-gzimage1M<n<10M6 likes477 downloads2y agoHugging Face22diffusion-cot /GenRef-CoT GenRef-CoT We provide 227K high-quality CoT reflections which were used to train our Qwen-based reflection generation model in ReflectionFlow [1]. To know the details of the dataset creation pipeline, please refer to Section 3.2 of [1]. Dataset loading We provide the dataset in the webdataset format for fast dataloading and streaming. We recommend downloading the repository locally for faster I/O: from huggingface_hub import snapshot_download local_dir =… See the full description on the dataset page: https://huggingface.co/datasets/diffusion-cot/GenRef-CoT.image100K<n<1M3 likes471 downloads1y agoHugging Face23andito /google-landmarksimage1M<n<10M1 likes459 downloads1y agoHugging Face24smz8599 /GUI-AIMA-multiturnimage100K<n<1M1 likes399 downloads7mo agoHugging Face25xwk123 /Mobile-GUI-Worldmodel-SFT Mobile-GUI-Worldmodel-SFT This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text. Repository Layout . ├── GUI-agent-main/ # Data annotation scripts and examples ├── eval/ # Evaluation assets │ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.documentimage-to-text10K<n<100K9 likes382 downloads5mo agoHugging Face26TrackingTeam /yuxuan_good_dataset_dtimage1M<n<10M0 likes312 downloads1mo agoHugging Face27gazeshift /VRGaze Dataset sample VRGaze dataset sample: image1M<n<10M0 likes300 downloads1y agoHugging Face28Zhongyuan /xl-genimage100K<n<1M0 likes289 downloads3y agoHugging Face29clip-benchmark /wds_gtsrbimage10K<n<100K0 likes286 downloads4y agoHugging Face30ek826 /imagenet-gen-sd1.5 ImageNet Generated using Stable Diffusion v1.5 The following repository mimics the size and class structure of the original ImageNet database. The classes can be found in the classes.txt file. This dataset contains approximately 1300 images per class over 1000 classes for a total of 1.3 million images. Here is an excerpt from classes.txt: 0 tench, Tinca tinca 1 goldfish, Carassius auratus 2 great white shark, white shark, man-eater, man-eating shark, Carcharodon caharias 3 tiger… See the full description on the dataset page: https://huggingface.co/datasets/ek826/imagenet-gen-sd1.5.imageimage-classification1M<n<10M5 likes228 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.