datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TN3K
TN3K: Thyroid Nodule Dataset for Segmentation and Classification
Overview
TN3K is a comprehensive open-access thyroid nodule dataset containing 3,493 thyroid ultrasound images with high-quality annotations for both segmentation and classification tasks. The dataset addresses the critical need for diverse, multi-center thyroid imaging data collected from various ultrasound devices and scanning views, reflecting real-world clinical scenarios.
Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/TN3K.google-streetview-images-by-country
Dataset Card for google streetview images by country
⚠️ There are still images that should be deleted, such as those with tags or those that didn't load correctly.
Dataset Structure
folder with the individual countries
images have the creation date and the map name in the file name.
Dataset Card Contact
use the community section
images per country
sports-cards
Digital Card Magazine Dataset
This dataset contains sports card images and their associated metadata for training machine learning models in card recognition, text extraction, and value estimation.
Dataset Description
Dataset Summary
A comprehensive collection of sports card images and metadata, including:
Front and back card images
OCR-extracted text with confidence scores
AI-analyzed card attributes
Card details (player, team, year, etc.)
Vision API labels… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/sports-cards.bakkhali-river-high-low-tide
Bakkhali River — High Tide vs Low Tide, Bangladesh
517 photographs of the Bakkhali River near Cox's Bazar, Bangladesh, documenting the same general stretch of river at high tide (264 images) and low tide (253 images). Captured across 10 separate sessions between 2 July and 15 August 2026.
This is not a frame-by-frame matched pair set — sessions were shot on different dates and the camera position varies within each session — but high- and low-tide frames come from the same short… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/bakkhali-river-high-low-tide.Gorilla-SPAC-Wild
Gorilla-SPAC-Wild: Large-Scale Video Dataset for Gorilla Re-Identification
Overview
Gorilla-SPAC-Wild is a comprehensive benchmark dataset for individual re-identification of Western Lowland Gorillas from camera trap footage in natural rainforest environments. This dataset addresses a critical bottleneck in conservation: automating the analysis of vast archives of camera trap video to track endangered gorilla populations non-invasively.
This dataset is part of the… See the full description on the dataset page: https://huggingface.co/datasets/gorilla-watch/Gorilla-SPAC-Wild.bakkhali-estuary-high-tide-sample
Bakkhali River Estuary — High Tide Boat Survey (Free Sample)
50 GPS-tagged coastal images from a single high-tide boat survey of the Bakkhali River
estuary, Khurushkul, Cox's Bazar, Bangladesh.
By Golam Rob — www.golamrob.com
✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)".
Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution.
📸 These 50 images are a small taste of a 200,000–300,000 image personal library of… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/bakkhali-estuary-high-tide-sample.water-hyacinth-flowers-floating-pond
Water Hyacinth Flowers — Khurushkul Pond, Bangladesh
150 ground-level photographs of water hyacinth (Eichhornia crassipes) blooming in a freshwater pond near Khurushkul, Cox's Bazar, Bangladesh. All frames were captured in a single session on 9 August 2026 (16:31–16:45 local time) during the monsoon season.
This is a follow-up survey of the same pond documented in the June 2026 water lily / water hyacinth dataset by the same photographer.
Contents
150 JPG images… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/water-hyacinth-flowers-floating-pond.golden-data-animals
Golden Data: Washin Animal Village
1554 high-quality animal images.
riverside-sunset-golden-hour
Riverside Sunset at Golden Hour — Bakkhali River, Bangladesh
100 photographs of sunset over the Bakkhali River at low tide, near Cox's Bazar, Bangladesh. All frames were captured in a single evening session on 13 July 2026 (17:59–18:24 local time, UTC+6) during golden hour.
Contents
100 JPG images, 2000×1333 px, filenames riverside-sunset-golden-hour-004.JPG through -108.JPG. Numbers are not fully sequential — 100 of 110 frames captured during the session were… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/riverside-sunset-golden-hour.Gorilla-Zoo-Berlin
Gorilla-Berlin-Zoo Dataset
Overview
The Gorilla-Berlin-Zoo dataset serves as a cross-domain evaluation benchmark dataset for gorilla re-identification systems, offering camera trap footage of Western Lowland Gorillas (Gorilla gorilla gorilla) in a controlled zoo environment.
The dataset is part of the GorillaWatch project, which introduces an end-to-end pipeline integrating detection, tracking, and re-identification for automated gorilla monitoring. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/gorilla-watch/Gorilla-Zoo-Berlin.khurushkul-pond-water-lily-sample
Pink Water Lily & Water Hyacinth — Khurushkul Pond, Bangladesh
100 GPS-tagged freshwater wetland images from a single pond survey in Khurushkul, Cox's Bazar, Bangladesh. By Golam Rob — www.golamrob.com
✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)". Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution.
📸 These 100 images are a small taste of a 200,000+ image personal library of coastal, tidal, and freshwater… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/khurushkul-pond-water-lily-sample.coastal-multitask-380
Coastal & Rural Bangladesh — Multi-Task Visual Dataset
379 field photographs (JPEG, native resolution as shot — see classification/metadata.csv for per-image width/height) collected on foot along the Bakkhali river embankment and surrounding villages/farmland near Cox's Bazar, Bangladesh, structured into three ML-task "levels": classification, semantic segmentation, and change detection.
Source: huggingface data 06 (Golam Rob / Tawhid Enterprise photo collection).… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/coastal-multitask-380.google-illustrations-full
Google Illustrations: Complete 4K Multi-Layer Archive (2,052 Avatars)
A comprehensive, uncompressed 4096x4096 archival dataset of the complete Google Account Illustrations library, featuring all 59 artist collections, 2,052 unique scenes, decomposed layer assets, color presets, and semantic metadata.
Legal Disclaimer and Copyright Notice
PLEASE READ CAREFULLY:
This repository is an independent archival and educational compilation provided strictly for… See the full description on the dataset page: https://huggingface.co/datasets/Anson10124/google-illustrations-full.Web_UI
Web UI Dataset
This dataset contains web pages, their screenshots across different devices, and images extracted from the web pages. Scrolling videos are stored separately in the 'video' folder. It is intended for use in machine learning tasks related to web design, computer vision, and data analysis.
Dataset Summary
Total web pages: 34
Total images: 492
Total screenshots: 102
Total videos: 34
Contents
For each web page, the dataset includes:
URL of the web… See the full description on the dataset page: https://huggingface.co/datasets/GoofyGoof/Web_UI.Gore-Blood-Dataset-v1.0
Gore Blood Dataset (Version 1.0)
Overview
The Gore Blood Dataset (Version 1.0) is a collection of images curated by NeuralShell specifically designed for training AI models, particularly for stable diffusion models. These images are intended to aid in the development and enhancement of machine learning models, leveraging the advancements in the field of computer vision and AI.
Dataset Information
Dataset Name: Gore-Blood-Dataset-v1.0
Creator: NeuralShell
Base… See the full description on the dataset page: https://huggingface.co/datasets/NeuralShell/Gore-Blood-Dataset-v1.0.mirage-news
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
[Paper]
[Github]
This dataset contains a total of 15,000 pieces of real or AI-generated multimodal news (image-caption pairs) -- a training set of 10,000 pairs, a validation set of 2,500 pairs, and five test sets of 500 pairs each. Four of the test sets are out-of-domain data from unseen news publishers and image generators to evaluate detector's generalization ability.
=== Data Source (News Publisher + Image Generator)… See the full description on the dataset page: https://huggingface.co/datasets/Gouge666/mirage-news.Good_TiresA-MNISTThe dataset is built on top of MNIST.
It consists from 130K of images in 10 classes - 120K training and 10K test samples.
The training set was augmented with additional 60K images.war-gov-uap-release-1
Department of War UAP Release 1 — structured corpus
The first tranche of declassified U.S. government records on Unidentified
Anomalous Phenomena (UAP / UFOs), released by the Department of War on
8 May 2026 under the Presidential Unsealing and Reporting System for
UAP Encounters (PURSUE) directive.
This dataset is a structured, machine-readable companion to the source
material at https://www.war.gov/UFO/. It pairs every original document
with VLM-extracted page text, cropped… See the full description on the dataset page: https://huggingface.co/datasets/MTSlive/war-gov-uap-release-1.testing-goldstandard-cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/jrw2989/testing-goldstandard-cuthill.Curated_GoldStandard_Hoyal_Cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Curated_GoldStandard_Hoyal_Cuthill.golf-courses
Dataset Summary: golf-course
This dataset (bethecloud/golf-courses) includes 21 unique images of golf courses pulled from Unsplash.
The dataset is a collection of photographs taken at various golf courses around the world. The images depict a variety of scenes, including fairways, greens, bunkers, water hazards, and clubhouse facilities. The images are high resolution and have been carefully selected to provide a diverse range of visual content for fine-tuning a machine learning… See the full description on the dataset page: https://huggingface.co/datasets/bethecloud/golf-courses.chars74k-eng-good
Chars74k
The "Good" subset of the "English" subset of the Chars74k dataset split into training and validation sets.
The validation set was created to match the label distribution of the training set.
62 classes (0-9, A-Z, a-z)
Dataset page: https://teodecampos.github.io/chars74k/
Paper describing the dataset:https://www.semanticscholar.org/paper/Character-Recognition-in-Natural-Images-Campos-Babu/dbbd5fdc09349bbfdee7aa7365a9d37716852b32
5 images where removed due to poor quality.… See the full description on the dataset page: https://huggingface.co/datasets/Lajdre/chars74k-eng-good.google-open-images-hair-style-dataset
Google Open Images — Hair Style Dataset
🇺🇸 English | 🇰🇷 한국어
Overview
This dataset is a curated custom subset of the Google Open Images V7 dataset, specifically filtered to include images of humans with various hair styles.It is intended for use in computer vision research and applications such as hair style classification, person detection, and instance segmentation.
Split
Purpose
train
Model training
validation
Model evaluation / hyperparameter tuning… See the full description on the dataset page: https://huggingface.co/datasets/hwany79/google-open-images-hair-style-dataset.ais-4sd
AIS-4SD
AIS-4SD (AI Summit - 4 Stable Diffusion models) is a collection of 4.000 images, generated using a set of Stability AI text-to-image diffusion models
Context
This dataset was developed during the development of a collaborative project between PEReN and VIGINUM for the AI Summit held in Paris in February 2025.
This open-source project aims at assessing generated images detectors performances and their robustness to different models and transformations.
The code is… See the full description on the dataset page: https://huggingface.co/datasets/peren-gouv/ais-4sd.lao-plate-dataset
Lao License Plate Provinces
Cropped Lao license-plate images labelled with the issuing province (18 classes),
plus the plate letters, digits, and color.
Task: image classification (province from a plate crop).
Splits are grouped by plate (province+letters+digits) so the same physical
plate never appears in two splits — honest eval. Seed 42, ~80/10/10.
label is a ClassLabel over the 18 canonical provinces; province is the Lao string.
reviewed: yes = consensus-verified (two… See the full description on the dataset page: https://huggingface.co/datasets/goftaidc/lao-plate-dataset.swgolden-foot-football-players
Golden Foot Football Players Dataset
Este dataset contiene imágenes de jugadores de fútbol que han sido nominados o premiados con el galardón Golden Foot. Las imágenes están organizadas por jugador y etiquetadas para tareas de clasificación.
Estructura
train: imágenes para entrenamiento
validation: imágenes para validación
test: imágenes para prueba
Licencia
Este dataset se distribuye bajo la licencia [Creative Commons Attribution 4.0 International(CC BY… See the full description on the dataset page: https://huggingface.co/datasets/aaronqg/golden-foot-football-players.IPL-Player-Detection-IITB-PML
IPL Player Detection Dataset — IITB PML Sem1
IPL cricket player detection dataset with cell-level team annotations and player counts. Created at IIT Bombay for the Python for Machine Learning (PML) Sem 1 project.
Dataset Overview
1005 images from IPL broadcast footage (800×600px)
8×8 grid annotation per image — each of 64 cells labeled with the IPL team present
Player count per image (0–20)
Train/Test split: 793 train / 212 test
10 IPL teams: CSK, DC, GT, KKR… See the full description on the dataset page: https://huggingface.co/datasets/goyaljai/IPL-Player-Detection-IITB-PML.ash_gourd_disease_classification
Ash Gourd Disease Classification
A dataset for disease classification of ash gourd.The dataset contains 2,676 images.Images per class:
Aphid: 140
Downy mildew: 1066
Healthy: 803
Leaf curl: 528
Leaf miner: 139
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{jahan2025dataset,
title={Dataset of Ash gourd plant leaf images for detection and classification},
author={Jahan, Nusrat and Hasan, Md… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/ash_gourd_disease_classification.
