CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01imageomics /TreeOfLife-200M Dataset Card for TreeOfLife-200M If you are looking for the original release TreeOfLife-200M dataset, as used in training BioCLIP 2 and presented the paper, please see Revision a8f38b4. The dataset, as presented here, was used to train BioCLIP 2.5 Huge; it completes the dataset cleaning process and resolves an issue where Observation.org occurrences were not included in the training data. With 233 million images representing 933,798 taxa across the tree of life, TreeOfLife-200M… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M.imageimage-classification100M<n<1B41 likes17k downloads4mo agoHugging Face02evanarlian /imagenet_1k_resized_256 Dataset Card for "imagenet_1k_resized_256" Dataset summary The same ImageNet dataset but all the smaller side resized to 256. A lot of pretraining workflows contain resizing images to 256 and random cropping to 224x224, this is why 256 is chosen. The resized dataset can also be downloaded much faster and consume less space than the original one. See here for detailed readme. Dataset Structure Below is the example of one row of data. Note that the labels in… See the full description on the dataset page: https://huggingface.co/datasets/evanarlian/imagenet_1k_resized_256.imageimage-classification1M<n<10M31 likes14k downloads3y agoHugging Face03DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes8.5k downloads19d agoHugging Face04benjamin-paine /imagenet-1k-256x256 Repack Information This repository contains a complete repack of ILSVRC/imagenet-1k in Parquet format with the following data transformations: Images were center-cropped to square to the minimum height/width dimension. Images were then rescaled to 256x256 using Lanczos resampling. Dataset Card for ImageNet Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/imagenet-1k-256x256.imageimage-classification1M<n<10M24 likes7.3k downloads2y agoHugging Face05perturb-ai /efficientnet-v2-l-adv-dataset Perturb Adversarial Images Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the Perturb network. Each row is one clean image together with all of its verified adversarial versions: images that are imperceptibly different from the original (L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction. This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.imageimage-classification1K<n<10K0 likes3k downloads17m agoHugging Face06cassiekang /cub200_dataset Dataset Card for CUB_200_2011 Dataset Summary The Caltech-UCSD Birds 200-2011 dataset (CUB-200-2011) is an extended version of the original CUB-200 dataset, featuring photos of 200 bird species primarily from North America. This 2011 version significantly expands its predecessor by doubling the number of images per class and introducing new part location annotations, alongside collecting detailed natural language descriptions for each image through Amazon Mechanical Turk… See the full description on the dataset page: https://huggingface.co/datasets/cassiekang/cub200_dataset.imageimage-classification10K<n<100K19 likes2.7k downloads3y agoHugging Face07nateraw /rendered-sst2 Rendered SST-2 The Rendered SST-2 Dataset from Open AI. Rendered SST2 is an image classification dataset used to evaluate the models capability on optical character recognition. This dataset was generated by rendering sentences in the Standford Sentiment Treebank v2 dataset. This dataset contains two classes (positive and negative) and is divided in three splits: a train split containing 6920 images (3610 positive and 3310 negative), a validation split containing 872 images (444… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/rendered-sst2.imageimage-classification1K<n<10K0 likes2.4k downloads4y agoHugging Face08andropar /relaion2b-natural LAION-Natural: Naturalness Scores for ReLAION-2B (CCN 2025, Roth & Hebart) LAION-Natural is a large-scale naturalness scoring dataset covering 2.1 billion images from ReLAION-2B-en-research-safe. Each image receives a score predicting how "natural" or "photographic" it looks versus artificial/rendered content. At the recommended threshold of 0.7, the dataset identifies ~500 million natural photographs suitable for vision research, cognitive science, and model training. Also… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural.imageimage-classification1B<n<10B5 likes2.1k downloads6mo agoHugging Face09Donghyun99 /CUB-200-2011 Dataset Card for "CUB-200-2011 (CUBS)" This is a non-official CUB-200-2011 dataset for fine-grained Image Classification. If you want to download the official dataset, please refer to the here. imageimage-classification10K<n<100K1 likes1.7k downloads2y agoHugging Face10trojblue /danbooru2025-metadata 🎨 Danbooru 2025 Metadata Latest Post ID: 9,158,800 (as of Apr 16, 2025) 📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork. Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes: More consistent tag history tracking Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.imagetext-to-image1M<n<10M38 likes1.6k downloads1y agoHugging Face11andropar /relaion2b-natural-embeddings LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart) LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7). Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.tabularfeature-extraction100M<n<1B1 likes1.5k downloads6mo agoHugging Face12timm /plant-pathology-2021 Description Dataset from the Plant Pathology 2021 (FGVC8) Challenge. ' For Plant Pathology 2021-FGVC8, we have significantly increased the number of foliar disease images and added additional disease categories. This year’s dataset contains approximately 23,000 high-quality RGB images of apple foliar diseases, including a large expert-annotated disease dataset. This dataset reflects real field scenarios by representing non-homogeneous backgrounds of leaf images taken at… See the full description on the dataset page: https://huggingface.co/datasets/timm/plant-pathology-2021.imageimage-classification10K<n<100K12 likes1.2k downloads1d agoHugging Face13NeurIPS-1899-ED-2026 /EpiBench-NeurIPS2026 EpiBench Anonymous release for NeurIPS 2026 Evaluations & Datasets Track review (paper ID 1899). All methodology, ablations, and analyses are in the companion paper; this card lists only what reviewers and downstream users need to load the data. A 25,737-patient ILAE-aligned multimodal epilepsy benchmark derived from PubMed Central case reports + 192 EpiRAG textbook vignettes. 6 tasks: epilepsy_type, seizure_type, ez_localization, aed_response, surgery_outcome, status_epilepticus… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-1899-ED-2026/EpiBench-NeurIPS2026.imagetext-classification100K<n<1M1 likes1.2k downloads5mo agoHugging Face14Holasyb918 /imagenet-1k-adm-crop-256 ImageNet-1k ADM Crop 256 This dataset is a preprocessed version of ILSVRC/imagenet-1k with all images center-cropped to 256×256 pixels using the ADM (Ablated Diffusion Model) algorithm. 🎯 Purpose Optimized for training diffusion models and other generative models that require fixed-size square images. 📊 Dataset Details Split Images Files Size (approx) train 1,281,167 294 ~38 GB test 50,000 28 ~3.5 GB 🔧 Processing Method… See the full description on the dataset page: https://huggingface.co/datasets/Holasyb918/imagenet-1k-adm-crop-256.imageimage-classification1M<n<10M0 likes1.1k downloads9mo agoHugging Face15Rapidata /text-2-image-Rich-Human-Feedback Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days. If you get value from this dataset and would like to see more in the future, please consider liking it. Overview We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.imagetext-to-image10K<n<100K37 likes1.1k downloads11h agoHugging Face16FaroukMoc2 /jev-stage2-image-beans-pilot Beans: one natural question per image Open the corrected preview. natural_v4 is the recommended and default preview: 100 original images, 100 rows, one three-way condition-class Choice question per image. All targets come directly from the source labels column (34 angular leaf spot, 33 bean rust, 33 healthy). Original image bytes and source annotations are unchanged. Example question: “Which source-defined condition class describes the bean leaf?” Options: angular_leaf_spot… See the full description on the dataset page: https://huggingface.co/datasets/FaroukMoc2/jev-stage2-image-beans-pilot.imageimage-classificationn<1K0 likes857 downloads8d agoHugging Face17bentrevett /caltech-ucsd-birds-200-2011 Caltech-UCSD Birds-200-2011 (CUB-200-2011) This dataset contains the Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset, from here. Each example consists of an image, a label, and a bounding box. (The dataset also contains x/y locations of "parts", e.g. beak, right eye, left wing, throat, etc. and "attributes", e.g. beak shape, wing color, feather pattern. I have not included either of these. Contact me if you want me to add them.) Note: Some of these images are also in ImageNet!… See the full description on the dataset page: https://huggingface.co/datasets/bentrevett/caltech-ucsd-birds-200-2011.imageimage-classification10K<n<100K1 likes851 downloads3y agoHugging Face18dgorbatov /vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10 Trajectory Ranking Dataset This dataset contains trajectory ranking results for autonomous navigation scenarios. Dataset Statistics Total examples: 39558 Chunks processed: 40 Upload date: 2025-09-13T00:44:30.335177 Features Image data with terrain analysis Trajectory rankings and reasoning Quality and diversity analysis Terrain and trajectory descriptions imageimage-classification10K<n<100K0 likes745 downloads1y agoHugging Face19flwrlabs /fed-isic2019 Dataset Card for Fed-ISIC-2019 Federated version of ISIC-2019 Datasets (ISIC2019 challenge and the HAM1000 database). This implementation is derived based on the FLamby implementation. Dataset Details The dataset contains 23,247 images of skin lesions divided among 6 clients representing different data centers. The number of samples for training/testing per data center is displayed in the table below: center_id Train Test 0 9930 2483 1 3163 791 2 2691… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/fed-isic2019.imageimage-classification10K<n<100K1 likes710 downloads2y agoHugging Face20bezzam /DigiCam-CelebA-26KData is measured at 30 cm, as shown below. After downloading and installing LenslessPiCam, the simulated PSF can be obtained and compared with the measured one with the following command: python scripts/sim/digicam_psf.py \ huggingface_repo=bezzam/DigiCam-CelebA-26K \ sim.waveprop=False \ sim.deadspace=True \ digicam.gamma=2.2 \ digicam.ap_center="[58,76]" \ digicam.ap_shape="[19,25]" \ digicam.rotate=0 \ digicam.horizontal_shift=-60 \ digicam.vertical_shift=-80 For a… See the full description on the dataset page: https://huggingface.co/datasets/bezzam/DigiCam-CelebA-26K.imageimage-to-image10K<n<100K0 likes694 downloads2y agoHugging Face21XIANG-Shuai /MCD-2.6m MCD-2.6m MCD-2.6m is a collection of 2,604,450 agricultural and plant images distributed in 49 Parquet shards. It combines images of multiple crops collected across several institutions and field-imaging projects. The release contains one train split. Images are embedded in the Parquet files and can be decoded directly with the Hugging Face datasets library. Dataset Structure Each example contains exactly three columns: Column Type Description row_id… See the full description on the dataset page: https://huggingface.co/datasets/XIANG-Shuai/MCD-2.6m.imageimage-classification1M<n<10M0 likes693 downloads2mo agoHugging Face22gaohongfa /imagenet-1k-256x256 Repack Information This repository contains a complete repack of ILSVRC/imagenet-1k in Parquet format with the following data transformations: Images were center-cropped to square to the minimum height/width dimension. Images were then rescaled to 256x256 using Lanczos resampling. Dataset Card for ImageNet Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in… See the full description on the dataset page: https://huggingface.co/datasets/gaohongfa/imagenet-1k-256x256.imageimage-classification1M<n<10M1 likes689 downloads8mo agoHugging Face23yuanchenyang /imagenet-256-flux2-vae-latents ImageNet-256 FLUX.2 VAE Latents Pre-computed deterministic, model-facing encodings from the FLUX.2 VAE (black-forest-labs/FLUX.2-dev) for the full ImageNet-1K training set at 256x256 resolution, stored as Parquet shards. Each example includes latents for both the original and horizontally flipped image, enabling flip augmentation without re-encoding at training time. Dataset Description Each example contains: Column Shape Stored type Description… See the full description on the dataset page: https://huggingface.co/datasets/yuanchenyang/imagenet-256-flux2-vae-latents.image-classification1M<n<10M0 likes586 downloads2mo agoHugging Face24Rapidata /Flux-2-pro_t2i_human_preference Rapidata Flux 2 Pro Preference This T2I dataset contains over ~400'000 human responses from over ~50'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation. Evaluating Flux 2 Pro (version from 25.11.25) across three categories: preference, coherence, and alignment. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux-2-pro_t2i_human_preference.imagetext-to-image10K<n<100K15 likes548 downloads10mo agoHugging Face25novaia /world-heightmaps-256px World Heightmaps 256px This is a dataset of 256x256 Earth heightmaps generated from SRTM 1 Arc-Second Global. Each heightmap is labelled according to its latitude and longitude. There are 573,995 samples. It is the same as World Heightmaps 360px but downsampled to 256x256. Method Convert GeoTIFFs into PNGs with Rasterio. import rasterio import matplotlib.pyplot as plt import os input_directory = '...' output_directory = '...' file_list =… See the full description on the dataset page: https://huggingface.co/datasets/novaia/world-heightmaps-256px.imageimage-classification100K<n<1M1 likes524 downloads3y agoHugging Face26ODELIA-AI /ODELIA-Challenge-2025gated ODELIA Challenge Dataset This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning. The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.tabularimage-classification1K<n<10K11 likes465 downloads11mo agoHugging Face27Thermostatic /frontier-synthetic-images-2026 Frontier Synthetic Images — Deduplicated Research Corpus This is a training-only corpus of 40,290 exact-deduplicated AI-generated images from recent and frontier generators. It normalizes four provenance-pinned sources into one row-per-image schema for image-forensics research. It is not an evaluation benchmark and should not be used to report detector accuracy after training on it. Sources and licensing Qwen/Qwen-Image-Bench at… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/frontier-synthetic-images-2026.imageimage-classification10K<n<100K0 likes460 downloads1mo agoHugging Face28nateraw /country211 Dataset Card for Country211 The Country 211 Dataset from OpenAI. This dataset was built by filtering the images from the YFCC100m dataset that have GPS coordinate corresponding to a ISO-3166 country code. The dataset is balanced by sampling 150 train images, 50 validation images, and 100 test images images for each country. imageimage-classification10K<n<100K6 likes442 downloads4y agoHugging Face29dbabnigg /botanical-vision-256 Botanical Vision Fine-grained flowering-plant classification dataset: 407,759 research-grade iNaturalist photos across 4,094 species (all flowering plants with at least 2,000 observations). Built for Advanced Computer Vision (UChicago ADSP 32023). Images are downscaled so the long edge is at most 256px (a smaller, Colab-friendly build of the full-resolution dataset). Splits split images train 285,136 val 61,288 test 61,335 Split is stratified… See the full description on the dataset page: https://huggingface.co/datasets/dbabnigg/botanical-vision-256.imageimage-classification100K<n<1M0 likes432 downloads3mo agoHugging Face30weikaih /ego4d-random-views-20k Ego4D Random Views Dataset This dataset contains 20,000 random view frames sampled from the Ego4D dataset using a high-performance multi-process generation system. Dataset Overview Total Images: 20,000 high-quality frames Image Format: PNG (1024×1024 resolution) Source: Ego4D v2 dataset (52,665+ video files) Sampling Method: Multi-process random sampling with maximum diversity Generation Time: 797.57 seconds (~13 minutes) Generation Speed: 25.08 frames/second… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/ego4d-random-views-20k.imageimage-classification10K<n<100K0 likes428 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.