CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Harvard-Edge /Wake-Vision Dataset Card for Wake Vision Dataset Description "Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally, it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age, subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.imageimage-classification1M<n<10M11 likes2k downloads10mo agoHugging Face02hardiksharma6555 /forgespectrum-114k ForgeSpectrum (v3) — AI-Generated Image Detection with Reasoning Traces ForgeSpectrum is a multi-domain corpus for AI-generated / manipulated image detection, annotated by Gemini-2.5-Pro with structured forensic reasoning traces (<fast>/<planning>/<reasoning>/<reflection>/<conclusion> patterns) plus per-image attributes and suspicious-region notes. v3 — what changed v3 is the cleaned, balanced release: 3 domains: faces, scenes, id_cards (docs and scene_text… See the full description on the dataset page: https://huggingface.co/datasets/hardiksharma6555/forgespectrum-114k.imageimage-classification10K<n<100K0 likes1.6k downloads4mo agoHugging Face03harpreetsahota /celebamask-hq Dataset Card for celebamask-hq This is a FiftyOne dataset with 30000 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/celebamask-hq") # Launch the App session = fo.launch_app(dataset) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/celebamask-hq.imageimage-classification10K<n<100K0 likes1.6k downloads1y agoHugging Face04taesiri /imagenet-hard-4K Dataset Card for "Imagenet-Hard-4K" Project Page - Paper - Github ImageNet-Hard-4K is 4K version of the original ImageNet-Hard dataset, which is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard-4K.imageimage-classification1K<n<10K7 likes1.2k downloads11mo agoHugging Face05Hariprasath5128 /marine-animals-multimodal-dataset Marine Animals Multimodal Dataset 🐋 A comprehensive multimodal dataset combining audio recordings and images of 32 marine species. Dataset Summary Total samples: 24,911 Species: 32 Audio files: 1,357 unique recordings Images: 581 (309 matched + 272 from iNaturalist) Features species (string): Species name label (int32): Numeric label (0–31) audio (Audio): Audio recording of the species image (Image): Species image image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.audioaudio-classification10K<n<100K0 likes904 downloads10mo agoHugging Face06dynamic-maps /hard-intersection-multimodal-sample Hard Intersection Multimodal Samples Release Notes Release Description v1.0.0 Initial public release. v1.1.0 Added Unreal Engine assets.Fixed issues in the OpenDRIVE map data.Updated the README to improve documentation and usability. Dataset Summary Hard Intersection Multimodal Samples is a curated multimodal dataset of accident-prone urban intersection in Japan for autonomous driving research and development.It provides… See the full description on the dataset page: https://huggingface.co/datasets/dynamic-maps/hard-intersection-multimodal-sample.imageimage-to-3dn<1K6 likes873 downloads3mo agoHugging Face07harvardairobotics /FairVision Dataset Card: Harvard-FairVision Dataset Summary Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes. This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.imageimage-classification10K<n<100K0 likes820 downloads5mo agoHugging Face08onullusoy /harmful-contents Harmful-Contents Dataset A multi-label image dataset for harmful-content classification across eight PEGI-aligned categories.The dataset consists of 5,153 rights-cleared images, split into train/validation/test sets and annotated with both binary labels and mask fields for controlled negative sampling. Dataset Structure Harmful-Contents/ csv/ train.csv val.csv test.csv data/ train/*.jpg val/*.jpg test/*.jpg Each CSV contains: name, alcohol… See the full description on the dataset page: https://huggingface.co/datasets/onullusoy/harmful-contents.imageimage-classification1K<n<10K1 likes740 downloads8mo agoHugging Face09harvardairobotics /FairVLMed Dataset Card: Harvard-FairVLMed Dataset Summary Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language. This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.imageimage-classification10K<n<100K0 likes287 downloads5mo agoHugging Face10youssefhassan13 /exp03-l23-hardening Exp 03 / 03b — Klüver L2/3 hardening study (SDXL + SD 3.5) Pre-registered study from the Operating System Hypothesis project. Sweep classifier-free guidance across two architectures and score every output blind, on two independent rubrics, for how far object structure has come apart. The prediction was written down and committed before the run. The commit dates in the GitHub repo are the proof. What is here that is not on GitHub The 860 generated PNGs. Every text… See the full description on the dataset page: https://huggingface.co/datasets/youssefhassan13/exp03-l23-hardening.imagetext-to-imagen<1K0 likes257 downloads1mo agoHugging Face11Harisundar /PALL-VLM-data PALL-VLM-data — Dental Vision-Language Dataset The training dataset for Harisundar/PALL-VLM, a dental vision-language model. It contains 32,884 records over 52,461 images, formatted as image+text conversations for LLaVA-style instruction tuning. Curated by: Harisundar R Used by: Harisundar/PALL-VLM · PALL on GitHub Language: English Layout vlm_train/ ├── images/ # 52,461 dental images ├── train.jsonl # 29,667 records ├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.imageimage-text-to-text10K<n<100K1 likes181 downloads4mo agoHugging Face12taesiri /imagenet-hard Dataset Card for "ImageNet-Hard" Project Page - ArXiv - Paper - Github - Image Browser Dataset Summary ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their ability to… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard.imageimage-classification10K<n<100K12 likes160 downloads11mo agoHugging Face13KIST-HARILAB /MAVIS MAVIS (Micro-surgical Artificial Vascular anastomosIS) This dataset was presented in the paper: SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding. Dataset Overview MAVIS is a microsurgical dataset comprising 19 videos of artificial vascular anastomosis procedures performed by three expert microsurgeons at College of Medicine, Korea University, Republic of Korea.For each video frame, it provides: Pixel-level… See the full description on the dataset page: https://huggingface.co/datasets/KIST-HARILAB/MAVIS.imageimage-segmentation0 likes160 downloads8mo agoHugging Face14harvardairobotics /Fundus-CoT Glaucoma Expert Chain-of-Thought Ophthalmologist six-step reasoning reports for fundus photographs, each paired with a binary glaucoma label. 1,074 cases from LAG and Papila. Files file rows glaucoma / not train.jsonl 823 304 / 519 val.jsonl 92 46 / 46 test.jsonl 159 79 / 80 images/ 1,074 <source>_<id>.jpg Record schema { "id": "1689", "source": "LAG", "image": "LAG_1689.jpg", "split": "train", "final_diagnosis_GT":… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/Fundus-CoT.imageimage-classification1K<n<10K0 likes122 downloads2mo agoHugging Face15u-haru /danbooru-tags-20260518Danbooru Dataset collected with my script. Collected post ids: 1 ~ 11403815 Usage: from datasets import load_dataset dataset = load_dataset("u-haru/danbooru-tags-20260518", split="train") imagetext-to-image10M<n<100M3 likes107 downloads4mo agoHugging Face16harveymannering /ultrasound_images_diffusionThis dataset contains synthetic EDM2 generated images of ultrasound scan, as described in the paper A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets. Code: https://github.com/xfetus/fetal-ultrasound-edm2 The class labels correspond to the following labels: plane_classes = { 0: 'Other', 1: 'Maternal cervix', 2: 'Fetal abdomen', 3: 'Fetal brain', 4: 'Fetal femur', 5: 'Fetal thorax', } imageimage-classification10K<n<100K0 likes104 downloads2mo agoHugging Face17harpreetsahota /marvel-masterpieces-with-3dmesh Dataset Card for reconstructions Wait! Before you go, ❤️ the dataset! Let's get this trending! This is a FiftyOne dataset with 255 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces-with-3dmesh.3dimage-classificationn<1K3 likes70 downloads2y agoHugging Face18harshitajainn /coco-cmfd COCO-CMFD A synthetic copy-move forgery dataset generated from MS-COCO 2017, with source/target-separated ground truth for copy-move forgery detection (CMFD). Each sample takes one annotated COCO object, applies a mild affine transform, and pastes it elsewhere in the same image at a location that passes scene-plausibility checks (support surface, horizon band, perspective scale, occlusion). Ground truth is provided as a 3-class trimap, a binary mask, a 16 px patch-label grid… See the full description on the dataset page: https://huggingface.co/datasets/harshitajainn/coco-cmfd.imageimage-segmentation10K<n<100K1 likes66 downloads2mo agoHugging Face19harvardairobotics /FairGenMed Dataset Card: FairGenMed Dataset Summary FairGenMed is the first dataset for studying fairness in medical generative models. It provides detailed quantitative clinical measurements alongside demographic annotations to investigate the semantic correlation between text prompts and anatomical regions across demographic subgroups. The dataset supports both generative model evaluation and downstream classification tasks for glaucoma detection. This dataset accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairGenMed.imageimage-classification10K<n<100K0 likes65 downloads5mo agoHugging Face20harpreetsahota /visual_ai_at_neurips2025_jina_with_ocr Dataset Card for harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr This is a FiftyOne dataset with 1134 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr.imageimage-classification1K<n<10K0 likes60 downloads11mo agoHugging Face21biglam /index-cards-harvard-botany-metropolitan-flora Card File of the Flora of the Metropolitan Parks (Harvard Botany Libraries, 1894–1895) 4,574 botanical specimen index cards from the Harvard University Botany Libraries' Card File of the Flora of the Metropolitan Parks, 1894–1895 (bulk), compiled by Walter Deane (1848–1930). Records flora of the Metropolitan Park system around Boston — Middlesex Fells Reservation, Blue Hills, Norfolk County, and adjacent areas — with one card per specimen entry: species, locality, collection date… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-harvard-botany-metropolitan-flora.imageimage-to-text1K<n<10K0 likes54 downloads4mo agoHugging Face22harvardairobotics /reLAIONet Dataset Card: reLAIONet Dataset Summary reLAIONet is a manually proofread, web-sourced image classification benchmark aligned to ImageNet's 1,000-class label space. Sourced entirely from open web crawls (reLAION-400M) rather than Flickr, it provides a challenging out-of-distribution complement to ImageNet val and ImageNetV2 for evaluating discriminative and class-conditional generative models. This dataset was introduced in: Fair Benchmarking of Emerging One-Step… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/reLAIONet.imageimage-classification10K<n<100K2 likes47 downloads6mo agoHugging Face23harpreetsahota /marvel-masterpieces Dataset Card for marvel_masterpieces Wait! Before you go, ❤️ the dataset! Let's get this trending! This is a FiftyOne dataset with 255 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/marvel-masterpieces") #… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces.imageimage-classificationn<1K1 likes37 downloads2y agoHugging Face24harpreetsahota /cardd_workshop_post_03 Dataset Card for harpreetsahota/cardd_workshop_post_03 This is a FiftyOne dataset with 2816 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/cardd_workshop_post_03") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/cardd_workshop_post_03.imageimage-classification1K<n<10K0 likes37 downloads10mo agoHugging Face25harpreetsahota /cardd_workshop_post_01 Dataset Card for car_dd This is a FiftyOne dataset with 2816 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/cardd_workshop_post_01") # Launch the App session = fo.launch_app(dataset) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/cardd_workshop_post_01.imageimage-classification1K<n<10K0 likes33 downloads1y agoHugging Face26detectors /ssb_hard-ood Dataset Card for SSB (hard) for OOD Detection Dataset Details Dataset Description Original Dataset Authors: Sagar Vaze, Kai Han, Andrea Vedaldi, Andrew Zisserman OOD Split Authors: Julian Bitterwolf, Maximilian Müller, Matthias Hein Shared by: Eduardo Dadalto License: unknown Dataset Sources Original Dataset Paper: http://arxiv.org/abs/2110.06207v2 First OOD Application Paper: http://arxiv.org/abs/2306.00826v1 Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/detectors/ssb_hard-ood.imageimage-classificationn<1K0 likes30 downloads3y agoHugging Face27harpreetsahota /ESC-10 Dataset Card for esc-10 This is a FiftyOne dataset with 400 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/ESC-10") # Launch the App session = fo.launch_app(dataset) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/ESC-10.imageimage-classificationn<1K0 likes30 downloads2y agoHugging Face28harpreetsahota /isaac_on_images Dataset Card for harpreetsahota/isaac_on_images This is a FiftyOne dataset with 50 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/isaac_on_images") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/isaac_on_images.imageimage-classificationn<1K0 likes28 downloads11mo agoHugging Face29HarishBonu /IIIT-INDIC-HW-WORDS-Hindi IIIT-INDIC-HW-WORDS-Hindi Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images. Overview The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research… See the full description on the dataset page: https://huggingface.co/datasets/HarishBonu/IIIT-INDIC-HW-WORDS-Hindi.imageimage-to-text10K<n<100K0 likes26 downloads1mo agoHugging Face30harpreetsahota /cardd_workshop_post_threed Dataset Card for harpreetsahota/cardd_workshop_post_03 This is a FiftyOne dataset with 2816 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/cardd_workshop_post_threed") # Launch the App session =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/cardd_workshop_post_threed.imageimage-classification1K<n<10K0 likes25 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.