CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01imageomics /TreeOfLife-200M Dataset Card for TreeOfLife-200M If you are looking for the original release TreeOfLife-200M dataset, as used in training BioCLIP 2 and presented the paper, please see Revision a8f38b4. The dataset, as presented here, was used to train BioCLIP 2.5 Huge; it completes the dataset cleaning process and resolves an issue where Observation.org occurrences were not included in the training data. With 233 million images representing 933,798 taxa across the tree of life, TreeOfLife-200M… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M.imageimage-classification100M<n<1B41 likes17k downloads4mo agoHugging Face02timm /oxford-iiit-pet The Oxford-IIIT Pet Dataset Description A 37 category pet dataset with roughly 200 images for each class. The images have a large variations in scale, pose and lighting. This instance of the dataset uses standard label ordering and includes the standard train/test splits. Trimaps and bbox are not included, but there is an image_id field that can be used to reference those annotations from official metadata. Website: https://www.robots.ox.ac.uk/~vgg/data/pets/… See the full description on the dataset page: https://huggingface.co/datasets/timm/oxford-iiit-pet.imageimage-classification1K<n<10K9 likes14k downloads3y agoHugging Face03DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes8.5k downloads19d agoHugging Face04timm /imagenet-22k-wdsgated Dataset Summary This is a copy of the full ImageNet dataset consisting of all of the original 21841 clases. It also contains labels in a separate field for the '12k' subset described at at (https://github.com/rwightman/imagenet-12k, https://huggingface.co/datasets/timm/imagenet-12k-wds) This dataset is from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-22k-wds.imageimage-classification100K<n<1M14 likes6.9k downloads3y agoHugging Face05torchgeo /eurosatRedistributed without modification from https://github.com/phelber/EuroSAT. EuroSAT100 is a subset of EuroSATallBands containing only 100 images. It is intended for tutorials and demonstrations, not for benchmarking. imageimage-classification10K<n<100K2 likes5.3k downloads2y agoHugging Face06TLAIM /TAIX-Ray TAIX-Ray Dataset TAIX-Ray is a comprehensive dataset of approximately 200k bedside chest radiographs from around 50k intensive care patients at University Hospital Aachen, Germany, collected between 2010 and 2024. Trained radiologists provided structured reports at the time of acquisition, assessing key findings such as cardiomegaly, pulmonary congestion, pleural effusion, pulmonary opacities, and atelectasis on an ordinal scale. Code & Details The code for data… See the full description on the dataset page: https://huggingface.co/datasets/TLAIM/TAIX-Ray.imageimage-classification100K<n<1M4 likes4.6k downloads5mo agoHugging Face07timm /resisc45 Description RESISC45 dataset is a publicly available benchmark for Remote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This dataset contains 31,500 images, covering 45 scene classes with 700 images in each class. The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/resisc45.imageimage-classification10K<n<100K7 likes4.3k downloads3y agoHugging Face08imageomics /TreeOfLife-10M Dataset Card for TreeOfLife-10M Dataset Summary With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M.documentimage-classification1M<n<10M50 likes4.2k downloads8mo agoHugging Face09TempoFunk /webvid-10Mtexttext-to-video10M<n<100M98 likes4.2k downloads3y agoHugging Face10timm /imagenet-1k-wdsgated Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated. 💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.imageimage-classification10K<n<100K35 likes3.9k downloads3y agoHugging Face11timm /imagenet-12k-wdsgated Dataset Summary This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm. The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-12k-wds.imageimage-classification100K<n<1M10 likes3.7k downloads3y agoHugging Face12tumor-vqa /DeepTumorVQA_2.0 DeepTumorVQA v2 3D abdominal-CT diagnostic Visual Question Answering benchmark with 42 clinical subtypes and 438K total QA pairs (10K curated benchmark + 428K training pool). Includes pre-extracted 2D and video modalities, 20K agent training trajectories with tool-use traces, and a paper-locked leaderboard. Resources 📄 Paper (arXiv) https://arxiv.org/abs/2605.09679 💻 Code (GitHub) https://github.com/Schuture/DeepTumorVQA 🤗 Dataset (this… See the full description on the dataset page: https://huggingface.co/datasets/tumor-vqa/DeepTumorVQA_2.0.imagevisual-question-answering10K<n<100K6 likes3.5k downloads4mo agoHugging Face13nebula /OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World. Code: https://github.com/iamwangyabin/OpenSDI imageimage-classification100K<n<1M2 likes3.2k downloads2y agoHugging Face14imageomics /IDLE-OO-Camera-Traps Dataset Card for IDLE-OO Camera Traps IDLE-OO Camera Traps is a 5-dataset benchmark of camera trap images from the Labeled Information Library of Alexandria: Biology and Conservation (LILA BC) with a total of 2,586 images for species classification. Each of the 5 benchmarks is balanced to have the same number of images for each species within it (between 310 and 1120 images), representing between 16 and 39 species. Supported Tasks and Leaderboards Image… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/IDLE-OO-Camera-Traps.imageimage-classification1K<n<10K2 likes2.8k downloads28d agoHugging Face15nebula /OpenSDI_test OpenSDI: Spotting Diffusion-Generated Images in the Open World This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper: Project Page: https://iamwangyabin.github.io/OpenSDI/ OpenSDID Dataset Highlights: User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs. Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.imageimage-classification100K<n<1M1 likes2.8k downloads2y agoHugging Face16timm /eurosat-rgb EuroSat (RGB) Description A dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples. This is the RGB version of the dataset with visible bands encoded as JPEG images. The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/eurosat-rgb.imageimage-classification10K<n<100K1 likes2.6k downloads3y agoHugging Face17ganghyunnnn /GSD-Sensitivity-Taxonomy-Labels GSD-Sensitivity Taxonomy: Task Labels for Remote Sensing VQA Per-task D / M1 / M2 taxonomy labels, inter-annotator agreement (IAA) data, and evaluation traces for four public RS-VQA benchmarks. Companion to *G. Park and D.-H. Lee, "Identifying the Measurement Gap in Remote Sensing VQA with a GSD-Sensitive Taxonomy," IEEE Geosci. Remote Sens. Lett., 2026* — accepted, DOI to follow. Code: github.com/ganghyunnnn/GSD-Sensitivity-Taxonomy ⚠️ This dataset contains annotations and… See the full description on the dataset page: https://huggingface.co/datasets/ganghyunnnn/GSD-Sensitivity-Taxonomy-Labels.textvisual-question-answeringn<1K0 likes1.8k downloads1mo agoHugging Face18BGLab /BioTrove-Train BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity Description See the BioTrove dataset card on HuggingFace to access the main BioTrove dataset (161.9M) BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of software… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove-Train.imageimage-classification100M<n<1B3 likes1.8k downloads1y agoHugging Face19EOA-team /SwissCrop25 SwissCrop25 A national benchmark dataset for operational crop mapping in Switzerland, providing Sentinel-2 time series, daily temperature data, and parcel-level crop type labels across seven growing seasons (2019–2025). Introduced in: SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping (TerraBytes II Workshop, ECCV 2026) — [Paper] [Code] [Team] Highlights Nationwide coverage of Switzerland (41,285 km²) Seven growing seasons (2019–2025) 73… See the full description on the dataset page: https://huggingface.co/datasets/EOA-team/SwissCrop25.tabularimage-segmentationn<1K3 likes1.6k downloads15d agoHugging Face20trojblue /danbooru2025-metadata 🎨 Danbooru 2025 Metadata Latest Post ID: 9,158,800 (as of Apr 16, 2025) 📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork. Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes: More consistent tag history tracking Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.imagetext-to-image1M<n<10M38 likes1.6k downloads1y agoHugging Face21semi-truths /Semi-Truths Semi Truths Dataset: A Large-Scale Dataset for Testing Robustness of AI-Generated Image Detectors (NeurIPS 2024 Track Datasets & Benchmarks Track) Recent efforts have developed AI-generated image detectors claiming robustness against various augmentations, but their effectiveness remains unclear. Can these systems detect varying degrees of augmentation? To address these questions, we introduce Semi-Truths, featuring 27, 600 real images, 223, 400 masks, and 1, 472, 700… See the full description on the dataset page: https://huggingface.co/datasets/semi-truths/Semi-Truths.imageimage-classification1M<n<10M9 likes1.6k downloads2y agoHugging Face22Bohan22 /MLS-Bench-Tasks MLS-Bench Tasks MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.texttext-generationn<1K8 likes1.5k downloads5mo agoHugging Face23tirtho149 /SAGE SAGE: Scalable Agentic Grounded Evaluation for Crop Disease Diagnosis Paper: arXiv:2605.09768 Arshad, Roy, Shen, Elango, Chiranjeevi, A. K. Singh, Ganapathysubramanian, Hegde, A. Singh, Sarkar. ~801,807 images across 335 crops / 1,251 disease classes, each enriched with source-grounded visual symptom knowledge: every visual description is extracted verbatim from an authoritative web page (university extension / APS / CABI) with a source URL and a supporting quote, and… See the full description on the dataset page: https://huggingface.co/datasets/tirtho149/SAGE.textimage-classification100K<n<1M2 likes1.5k downloads4d agoHugging Face24tpremoli /CelebA-attrs CelebA-128x128 CelebA with attrs at 128x128 resolution. Dataset Information The attributes are binary attributes. The dataset is already split into train/test/validation sets. Citation @inproceedings{liu2015faceattributes, title = {Deep Learning Face Attributes in the Wild}, author = {Liu, Ziwei and Luo, Ping and Wang, Xiaogang and Tang, Xiaoou}, booktitle = {Proceedings of International Conference on Computer Vision (ICCV)}, month = {December}, year… See the full description on the dataset page: https://huggingface.co/datasets/tpremoli/CelebA-attrs.imagefeature-extraction100K<n<1M10 likes1.4k downloads3y agoHugging Face25owkin /plism-dataset-tiles PLISM dataset The Pathology Images of Scanners and Mobilephones (PLISM) dataset was created by (Ochi et al., 2024) for the evaluation of AI models’ robustness to inter-institutional domain shifts. All histopathological specimens used in creating the PLISM dataset were sourced from patients who were diagnosed and underwent surgery at the University of Tokyo Hospital between 1955 and 2018. PLISM-wsi consists in a group of consecutive slides digitized under 7 different scanners and… See the full description on the dataset page: https://huggingface.co/datasets/owkin/plism-dataset-tiles.imageimage-feature-extraction1M<n<10M9 likes1.3k downloads2y agoHugging Face26timm /plant-pathology-2021 Description Dataset from the Plant Pathology 2021 (FGVC8) Challenge. ' For Plant Pathology 2021-FGVC8, we have significantly increased the number of foliar disease images and added additional disease categories. This year’s dataset contains approximately 23,000 high-quality RGB images of apple foliar diseases, including a large expert-annotated disease dataset. This dataset reflects real field scenarios by representing non-homogeneous backgrounds of leaf images taken at… See the full description on the dataset page: https://huggingface.co/datasets/timm/plant-pathology-2021.imageimage-classification10K<n<100K12 likes1.2k downloads19h agoHugging Face27purvanshi /TASTE TASTE: Human Preferences for Design-Quality Image Comparison This dataset is the human-evaluation corpus released alongside the TASTE preference model. It contains panel rankings of generated images across multiple quality dimensions — both aesthetic (does the image look good?) and description-faithfulness (does the image match what the prompt describes?) — plus a per-image hallucination judgement. Quick stats Table Rows Notes prompts.parquet ~200 one… See the full description on the dataset page: https://huggingface.co/datasets/purvanshi/TASTE.imageimage-classification10K<n<100K9 likes1.2k downloads4mo agoHugging Face28taesiri /imagenet-hard-4K Dataset Card for "Imagenet-Hard-4K" Project Page - Paper - Github ImageNet-Hard-4K is 4K version of the original ImageNet-Hard dataset, which is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard-4K.imageimage-classification1K<n<10K7 likes1.2k downloads11mo agoHugging Face29Yunncheng /Mirage-Test 🌊 Mirage-Test Dataset Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models. It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics. The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity. 📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.imageimage-classification10K<n<100K3 likes1.2k downloads10mo agoHugging Face30Schrodin-purrrrr /toothbrush-v2-dataset Toothbrushing Detection Dataset (v2) Video and image data for detecting toothbrushing behavior, collected for a Raspberry Pi Zero 2W toothbrush-detection project (toothbrush_v2). A single-class object detector is trained on this data to output [x, y, w, h, confidence] for the toothbrush in frame. Dataset structure Files are packed into tar shards (rather than uploaded individually) to stay within the Hub's per-repo file-count guidelines. To reconstruct the… See the full description on the dataset page: https://huggingface.co/datasets/Schrodin-purrrrr/toothbrush-v2-dataset.imageobject-detection10K<n<100K0 likes1.2k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.