CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01biglam /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.imageimage-classification1M<n<10M64 likes6.7k downloads1mo agoHugging Face02Faizaniqbal /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.imageimage-classification1M<n<10M0 likes2.4k downloads1mo agoHugging Face03Rapidata /human-style-preferences-images Rapidata Image Generation Preference Dataset This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview One of the largest human preference datasets for text-to-image models, this release contains over 1,200,000 human preference… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-style-preferences-images.imagetext-to-image10K<n<100K29 likes707 downloads2y agoHugging Face04DigiGreen /Crop_Disease_Images Crop Disease Expert Annotations 1,026 expert annotations over 989 crop photographs, covering pests, diseases and nutrient deficiencies across 74 crop types. The images are included in this repository. Every image is a photo taken by a smallholder farmer on their own plot and sent to Farmer.Chat, an AI advisory service run by Digital Green. Agronomists then reviewed each photo on Digital Green's annotation platform. Nothing here is scraped, staged, or lab-photographed.… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Crop_Disease_Images.imageimage-classification1K<n<10K5 likes615 downloads2mo agoHugging Face05PestoRosso /lamoda-fashion-product-images High-Resolution Fashion Product Images This dataset is a highly optimized, high-resolution subset of the popular Fashion Product Images Dataset originally hosted on Kaggle. It contains thousands of unique e-commerce fashion products, combining high-resolution product images with multiple descriptive label attributes. All low-resolution thumbnails and anomalies have been aggressively filtered out. Every image in this dataset has a minimum resolution of 640px on its shortest… See the full description on the dataset page: https://huggingface.co/datasets/PestoRosso/lamoda-fashion-product-images.imageimage-classification10K<n<100K1 likes577 downloads3mo agoHugging Face06thomasht86 /road-images-and-embeddings Norwegian Road Images with Embeddings (Trondheim Area) A dataset of 34,908 road images from the Trondheim region of Norway (~40km radius), captured by Statens vegvesen (Norwegian Public Roads Administration) in 2025. Each image is paired with rich geospatial metadata, nearest address information, and a 3072-dimensional image embedding from Google's gemini-embedding-2-preview model. Dataset Structure Each example contains: Field Type Description image Image… See the full description on the dataset page: https://huggingface.co/datasets/thomasht86/road-images-and-embeddings.imageimage-feature-extraction1K<n<10K1 likes541 downloads6mo agoHugging Face07shravya11 /truck-images Truck Detection and Counting Dataset This repository contains multiple computer vision datasets for truck detection, counting, and classification. 1. Raw Truck Images (Root folder) Number of images: 466 Format: Unannotated images (JPEG/PNG) License: CC BY 4.0 2. Truck Counting Dataset Number of images: 1769 Format: YOLOv8 format (images, labels, and data.yaml) Location: truck_counting_yolov8/ folder License: CC BY 4.0 (Roboflow export) 3. Trucks… See the full description on the dataset page: https://huggingface.co/datasets/shravya11/truck-images.imageobject-detection1K<n<10K0 likes515 downloads4mo agoHugging Face08Thermostatic /frontier-synthetic-images-2026 Frontier Synthetic Images — Deduplicated Research Corpus This is a training-only corpus of 40,290 exact-deduplicated AI-generated images from recent and frontier generators. It normalizes four provenance-pinned sources into one row-per-image schema for image-forensics research. It is not an evaluation benchmark and should not be used to report detector accuracy after training on it. Sources and licensing Qwen/Qwen-Image-Bench at… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/frontier-synthetic-images-2026.imageimage-classification10K<n<100K0 likes514 downloads1mo agoHugging Face09Hassan881 /detection-images Hassan881/detection-images Staging dataset for Berhan XAI (BurhanXAI) AI-generated / edited-media detection. Field Value Training contract v0.3.0 (firefly top-up applied; 6/6 M2 providers) Held-out eval (M2) v0.3.0-eval Held-out eval (legacy) v0.2.1-eval Status published Classes real, ai_generated, ai_edited Seed 42 from datasets import load_dataset ds = load_dataset("Hassan881/detection-images", name="v0.3.0") ev =… See the full description on the dataset page: https://huggingface.co/datasets/Hassan881/detection-images.imageimage-classification10K<n<100K0 likes423 downloads5d agoHugging Face10stochastic /random_streetview_images_pano_v0.0.2 Dataset Card for panoramic street view images (v.0.0.2) Dataset Summary The random streetview images dataset are labeled, panoramic images scraped from randomstreetview.com. Each image shows a location accessible by Google Streetview that has been roughly combined to provide ~360 degree view of a single location. The dataset was designed with the intent to geolocate an image purely based on its visual content. Supported Tasks and Leaderboards None as of now!… See the full description on the dataset page: https://huggingface.co/datasets/stochastic/random_streetview_images_pano_v0.0.2.imageimage-classification10K<n<100K29 likes373 downloads4y agoHugging Face11HawkFranklin-Research /SCIN-Dermatology-Raw-Images SCIN-Dermatology-Raw-Images This dataset contains 6,517 patient-submitted photographs organized into 3,061 clinical cases of common skin diseases. The source images are curated from the public Google Skin Condition Image Network (SCIN) corpus, cleansed of quality and gradability conflicts, and paired with complete patient-reported demographics, clinical symptoms, and dermatologist gradings. Dataset Structure This repository follows the standard Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/SCIN-Dermatology-Raw-Images.imageimage-classification1K<n<10K0 likes372 downloads3mo agoHugging Face12gtfintechlab /ipo-images SEC IPO Filing Image Dataset A large-scale, labeled dataset of 76,000+ images extracted from U.S. IPO registration statements (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026. Every image has been classified through a multi-stage pipeline: initial detection with YOLOv8, followed by verification from an ensemble of 8 Vision-Language Models (VLMs). Chart images include additional structured metadata describing chart type, visual properties, and content.… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-images.imageimage-classification10K<n<100K7 likes339 downloads7mo agoHugging Face13simonko912 /crawl-images-2A simple dataset, from random sites crawled images First few images are basic sites like youtube, google, etc then it slowly transformes to more random at the end imageimage-classification10K<n<100K1 likes283 downloads2mo agoHugging Face14cipher982 /wine-images-126k Wine Images Dataset 126K A comprehensive dataset of 107,821 wine bottle images linked to the Wine Text Dataset 126K. This companion dataset provides high-quality wine bottle images for computer vision, multimodal machine learning, and wine recognition tasks. Dataset Description This dataset contains wine bottle images scraped from wine retailer websites. Each image is linked to detailed wine information (descriptions, pricing, categories, regions) via stable IDs that… See the full description on the dataset page: https://huggingface.co/datasets/cipher982/wine-images-126k.imageimage-classification100K<n<1M2 likes274 downloads8mo agoHugging Face15GangHitman /fashion-recommendation-images High-Resolution Fashion Product Images This dataset is a highly optimized, high-resolution subset of the popular Fashion Product Images Dataset originally hosted on Kaggle. It contains thousands of unique e-commerce fashion products, combining high-resolution product images with multiple descriptive label attributes. All low-resolution thumbnails and anomalies have been aggressively filtered out. Every image in this dataset has a minimum resolution of 640px on its shortest… See the full description on the dataset page: https://huggingface.co/datasets/GangHitman/fashion-recommendation-images.imageimage-classification10K<n<100K0 likes231 downloads1mo agoHugging Face16OpenMed /multicare-images MultiCaRe: Open-Source Clinical Case Dataset MultiCaRe is an open-source, multimodal clinical case dataset built from the PubMed Central Open Access (OA) Case Report articles. It aggregates de-identified, open-access case narratives, figure images, captions, and rich article metadata across diverse specialties (radiology, pathology, surgery, ophthalmology, etc.). The data is normalized so images, cases, and articles can be joined via stable IDs. Source and process: OA case reports… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/multicare-images.imageimage-classification100K<n<1M6 likes211 downloads1y agoHugging Face17simonko912 /crawl-imagesA simple dataset, from random sites crawled images First few images are basic sites like youtube, google, etc then it slowly transformes to more random at the end imageimage-classification10K<n<100K1 likes205 downloads5mo agoHugging Face18SoyVitou /62k-images-khmer-printed-dataset 62k Khmer-English Printed Dataset This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS). Installation Prerequisites Before cloning this repository, make sure you have Git LFS installed: Install Git LFS Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.imagetext-generation10K<n<100K2 likes201 downloads2y agoHugging Face19nyuuzyou /rule34lol-images-part2 Dataset Card for rule34lol-images-part2 Dataset Summary This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 77,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files (except the last archive). This is Part 2 of 2 for the complete rule34lol-images dataset. Part 1 can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part2.imageimage-classification100K<n<1M5 likes196 downloads2y agoHugging Face20SY95 /WCCA-AK-images WCCA-AK: Wearable Cultural Collection Archive - André Kim Dataset Description WCCA-AK is a large-scale dataset of 3D scans and multi-view images capturing 100 haute couture garments by André Kim (1962–2010), one of Korea's most iconic fashion designers. This dataset bridges computer vision research and cultural heritage preservation, enabling both faithful documentation of historical artifacts and generative exploration of artistic vision. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SY95/WCCA-AK-images.imageimage-to-3d1K<n<10K0 likes160 downloads1y agoHugging Face21BrainCause /Concept_Targeted_Causal_Images Dataset Card for Concept-Targeted Causal Images Dataset Summary Concept-Targeted Causal Images is a concept-centric image dataset designed for studying causal visual representations in the brain. For each concept, the dataset contains three complementary image types: Positive images that clearly depict the target concept Semantic negatives that are visually or semantically related to the concept, but do not satisfy it Counterfactual edits created by editing… See the full description on the dataset page: https://huggingface.co/datasets/BrainCause/Concept_Targeted_Causal_Images.imageimage-classification100K<n<1M4 likes157 downloads4mo agoHugging Face22Sigurdur /isl-finepdfs-images Icelandic FinePDFs Images Dataset Description This dataset contains page-level images extracted from Icelandic PDFs in the HuggingFaceFW/finepdfs collection (isl_Latn subset). Each PDF page has been converted to a PNG image with associated metadata and structured OCR output generated using rednote-hilab/dots.ocr. Dataset Summary Language: Icelandic (is) Source: HuggingFaceFW/finepdfs (isl_Latn) OCR Model: rednote-hilab/dots.ocr Format: PNG images with… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/isl-finepdfs-images.imageimage-to-text1K<n<10K3 likes107 downloads7mo agoHugging Face23V4ldeLund /nationalmuseet-open-images Nationalmuseet Open Images This dataset is an independently harvested research dataset from Nationalmuseet Samlinger Online. It contains metadata and optionally WebDataset image shards for Nationalmuseet asset records whose rights.license is one of: Public Domain CC-BY No known rights Public Domain and CC-BY are the strict open-license subset. No known rights is kept as a separate license bucket because Nationalmuseet says this label means that, to their best assessment, the… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/nationalmuseet-open-images.imageimage-classification100K<n<1M0 likes81 downloads4mo agoHugging Face24ImageIN /ImageIn_annotations_resized_images Dataset Card for ImageIn_annotations_resized_images More Information needed imageimage-classification1K<n<10K0 likes72 downloads3y agoHugging Face25sahilur /hyper-kvasir-labeled-imagesHyperKvasir Labeled images In total, the dataset contains 10,662 labeled images stored using the JPEG format. The images can be found in the images folder. The classes, which each of the images belongto, correspond to the folder they are stored in (e.g., the ’polyp’ folder contains all polyp images, the ’barretts’ folder contains all images of Barrett’s esophagus, etc.). The number of images per class are not balanced, which is a general challenge in the medical field due to the fact that some… See the full description on the dataset page: https://huggingface.co/datasets/sahilur/hyper-kvasir-labeled-images.imageimage-classification10K<n<100K1 likes71 downloads8mo agoHugging Face26KhangTruong /Tampered-Images-Combined-Dataset KhangTruong/Merged (cross_test) This benchmark combines multi-generator AI image tampering and manipulation localization datasets. Every sample contains an image and a ground-truth binary mask indicating tampered regions. Columns id (string): Unique identifier for the sample. image (image): The image (authentic or manipulated). mask (image): Single-channel ground-truth mask. For real images, this is an all-black mask (0). For fake images, white pixels (255)… See the full description on the dataset page: https://huggingface.co/datasets/KhangTruong/Tampered-Images-Combined-Dataset.imageimage-segmentation1K<n<10K0 likes66 downloads4d agoHugging Face27nyuuzyou /clker-images Dataset Card for Clker.com Images Dataset Summary This dataset contains 140,313 public domain clipart images collected from Clker.com. Clker.com hosts user-shared vector clip art that is explicitly released into the public domain (CC0). The dataset includes the images themselves along with metadata such as titles and tags associated with each image. Languages The dataset is primarily monolingual: English (en): All image titles and tags are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/clker-images.imageimage-classification100K<n<1M2 likes57 downloads1y agoHugging Face28nyuuzyou /rule34lol-images-part1 Dataset Card for rule34lol-images-part1 Dataset Summary This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 196,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. This is Part 1 of 2 for the complete rule34lol-images dataset. Part 2 can be found here. Languages… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part1.imageimage-classification100K<n<1M5 likes56 downloads2y agoHugging Face29hwany79 /google-open-images-hair-style-dataset Google Open Images — Hair Style Dataset 🇺🇸 English | 🇰🇷 한국어 Overview This dataset is a curated custom subset of the Google Open Images V7 dataset, specifically filtered to include images of humans with various hair styles.It is intended for use in computer vision research and applications such as hair style classification, person detection, and instance segmentation. Split Purpose train Model training validation Model evaluation / hyperparameter tuning… See the full description on the dataset page: https://huggingface.co/datasets/hwany79/google-open-images-hair-style-dataset.imageimage-classification1K<n<10K0 likes51 downloads7mo agoHugging Face30spongus /milly-imagesA collection of images from a very silly cat, these are all from @fatfatmillycat in twitter. Intended to be used with stable-diffusion-v1-4 imagetext-to-imagen<1K3 likes49 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.