datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TN3K
TN3K: Thyroid Nodule Dataset for Segmentation and Classification
Overview
TN3K is a comprehensive open-access thyroid nodule dataset containing 3,493 thyroid ultrasound images with high-quality annotations for both segmentation and classification tasks. The dataset addresses the critical need for diverse, multi-center thyroid imaging data collected from various ultrasound devices and scanning views, reflecting real-world clinical scenarios.
Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/haifan-gong/TN3K.google-streetview-images-by-country
Dataset Card for google streetview images by country
⚠️ There are still images that should be deleted, such as those with tags or those that didn't load correctly.
Dataset Structure
folder with the individual countries
images have the creation date and the map name in the file name.
Dataset Card Contact
use the community section
images per country
sports-cards
Digital Card Magazine Dataset
This dataset contains sports card images and their associated metadata for training machine learning models in card recognition, text extraction, and value estimation.
Dataset Description
Dataset Summary
A comprehensive collection of sports card images and metadata, including:
Front and back card images
OCR-extracted text with confidence scores
AI-analyzed card attributes
Card details (player, team, year, etc.)
Vision API labels… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/sports-cards.bakkhali-river-high-low-tide
Bakkhali River — High Tide vs Low Tide, Bangladesh
517 photographs of the Bakkhali River near Cox's Bazar, Bangladesh, documenting the same general stretch of river at high tide (264 images) and low tide (253 images). Captured across 10 separate sessions between 2 July and 15 August 2026.
This is not a frame-by-frame matched pair set — sessions were shot on different dates and the camera position varies within each session — but high- and low-tide frames come from the same short… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/bakkhali-river-high-low-tide.Gorilla-SPAC-Wild
Gorilla-SPAC-Wild: Large-Scale Video Dataset for Gorilla Re-Identification
Overview
Gorilla-SPAC-Wild is a comprehensive benchmark dataset for individual re-identification of Western Lowland Gorillas from camera trap footage in natural rainforest environments. This dataset addresses a critical bottleneck in conservation: automating the analysis of vast archives of camera trap video to track endangered gorilla populations non-invasively.
This dataset is part of the… See the full description on the dataset page: https://huggingface.co/datasets/gorilla-watch/Gorilla-SPAC-Wild.inline-digital-holography-v2
Dataset Card for Synthetic Inline Holographical Images
This dataset provides synthetic image triplets representing inline holographical imaging in a simulated environment. Each data sample consists of:
An object-domain field (ground truth),
Its corresponding forward-propagated hologram (the inline holographic pattern at the sensor plane, intensity),
The numerically reconstructed image (via angular spectrum method).
The dataset is intended to facilitate research in computational… See the full description on the dataset page: https://huggingface.co/datasets/gokhankocmarli/inline-digital-holography-v2.inline-digital-holography-v3
Dataset Card for Synthetic Inline Holographical Images v3 (224px Highly Diverse)
This dataset provides synthetic image triplets representing inline holographical imaging in a simulated environment. This version (v3) uses a native 224x224 resolution optimized for modern Vision Transformers (ViT, Swin) and contains 25,000 samples across 8 noise configurations.
Each data sample consists of:
An object-domain field (ground truth),
Its corresponding forward-propagated hologram (the… See the full description on the dataset page: https://huggingface.co/datasets/gokhankocmarli/inline-digital-holography-v3.bakkhali-estuary-high-tide-sample
Bakkhali River Estuary — High Tide Boat Survey (Free Sample)
50 GPS-tagged coastal images from a single high-tide boat survey of the Bakkhali River
estuary, Khurushkul, Cox's Bazar, Bangladesh.
By Golam Rob — www.golamrob.com
✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)".
Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution.
📸 These 50 images are a small taste of a 200,000–300,000 image personal library of… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/bakkhali-estuary-high-tide-sample.water-hyacinth-flowers-floating-pond
Water Hyacinth Flowers — Khurushkul Pond, Bangladesh
150 ground-level photographs of water hyacinth (Eichhornia crassipes) blooming in a freshwater pond near Khurushkul, Cox's Bazar, Bangladesh. All frames were captured in a single session on 9 August 2026 (16:31–16:45 local time) during the monsoon season.
This is a follow-up survey of the same pond documented in the June 2026 water lily / water hyacinth dataset by the same photographer.
Contents
150 JPG images… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/water-hyacinth-flowers-floating-pond.golden-data-animals
Golden Data: Washin Animal Village
1554 high-quality animal images.
BrainAge-Golden-Raw
BrainAge Golden Raw MRI Dataset
Curated collection of 6,152 healthy-brain T1-weighted MRI scans spanning
ages 0–86 years, assembled from 12 public neuroimaging datasets.
Splits
Split
Subjects
Age range
Size
Golden-0-to-25/
4,782
0 – 25 y
~42 GB
Golden-25plus/
1,370
25 – 86 y
~13 GB
Source datasets
BCP, Calgary, ds002726, ds000248, PTBP, IXI, MPI-Leipzig,
AOMIC-ID1000, NKI-Rockland, ABIDE-I, ABIDE-II, ADHD-200.
File… See the full description on the dataset page: https://huggingface.co/datasets/zareenz741/BrainAge-Golden-Raw.riverside-sunset-golden-hour
Riverside Sunset at Golden Hour — Bakkhali River, Bangladesh
100 photographs of sunset over the Bakkhali River at low tide, near Cox's Bazar, Bangladesh. All frames were captured in a single evening session on 13 July 2026 (17:59–18:24 local time, UTC+6) during golden hour.
Contents
100 JPG images, 2000×1333 px, filenames riverside-sunset-golden-hour-004.JPG through -108.JPG. Numbers are not fully sequential — 100 of 110 frames captured during the session were… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/riverside-sunset-golden-hour.Gorilla-Zoo-Berlin
Gorilla-Berlin-Zoo Dataset
Overview
The Gorilla-Berlin-Zoo dataset serves as a cross-domain evaluation benchmark dataset for gorilla re-identification systems, offering camera trap footage of Western Lowland Gorillas (Gorilla gorilla gorilla) in a controlled zoo environment.
The dataset is part of the GorillaWatch project, which introduces an end-to-end pipeline integrating detection, tracking, and re-identification for automated gorilla monitoring. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/gorilla-watch/Gorilla-Zoo-Berlin.khurushkul-pond-water-lily-sample
Pink Water Lily & Water Hyacinth — Khurushkul Pond, Bangladesh
100 GPS-tagged freshwater wetland images from a single pond survey in Khurushkul, Cox's Bazar, Bangladesh. By Golam Rob — www.golamrob.com
✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)". Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution.
📸 These 100 images are a small taste of a 200,000+ image personal library of coastal, tidal, and freshwater… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/khurushkul-pond-water-lily-sample.uijudge-bench
UIJudgeBench v0.4.0 (pre-release)
UIJudgeBench evaluates systems that judge web UI quality from frozen, versioned page
artifacts with machine-checkable ground truth. It covers accessibility, layout,
referring/computed-style questions, and a separate pairwise design-quality instrument.
This is the dataset release. The independently versioned Python harness is available on
PyPI, and the canonical source release is
GitHub v0.4.0. The PyPI
wheel intentionally does not bundle this… See the full description on the dataset page: https://huggingface.co/datasets/gojiberries/uijudge-bench.coastal-multitask-380
Coastal & Rural Bangladesh — Multi-Task Visual Dataset
379 field photographs (JPEG, native resolution as shot — see classification/metadata.csv for per-image width/height) collected on foot along the Bakkhali river embankment and surrounding villages/farmland near Cox's Bazar, Bangladesh, structured into three ML-task "levels": classification, semantic segmentation, and change detection.
Source: huggingface data 06 (Golam Rob / Tawhid Enterprise photo collection).… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/coastal-multitask-380.google-illustrations-full
Google Illustrations: Complete 4K Multi-Layer Archive (2,052 Avatars)
A comprehensive, uncompressed 4096x4096 archival dataset of the complete Google Account Illustrations library, featuring all 59 artist collections, 2,052 unique scenes, decomposed layer assets, color presets, and semantic metadata.
Legal Disclaimer and Copyright Notice
PLEASE READ CAREFULLY:
This repository is an independent archival and educational compilation provided strictly for… See the full description on the dataset page: https://huggingface.co/datasets/Anson10124/google-illustrations-full.quickdrawThe Quick Draw Dataset is a collection of 50 million drawings across 345 categories, contributed by players of the game Quick, Draw!.
The drawings were captured as timestamped vectors, tagged with metadata including what the player was asked to draw and in which country the player was located.Web_UI
Web UI Dataset
This dataset contains web pages, their screenshots across different devices, and images extracted from the web pages. Scrolling videos are stored separately in the 'video' folder. It is intended for use in machine learning tasks related to web design, computer vision, and data analysis.
Dataset Summary
Total web pages: 34
Total images: 492
Total screenshots: 102
Total videos: 34
Contents
For each web page, the dataset includes:
URL of the web… See the full description on the dataset page: https://huggingface.co/datasets/GoofyGoof/Web_UI.Gore-Blood-Dataset-v1.0
Gore Blood Dataset (Version 1.0)
Overview
The Gore Blood Dataset (Version 1.0) is a collection of images curated by NeuralShell specifically designed for training AI models, particularly for stable diffusion models. These images are intended to aid in the development and enhancement of machine learning models, leveraging the advancements in the field of computer vision and AI.
Dataset Information
Dataset Name: Gore-Blood-Dataset-v1.0
Creator: NeuralShell
Base… See the full description on the dataset page: https://huggingface.co/datasets/NeuralShell/Gore-Blood-Dataset-v1.0.mirage-news
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
[Paper]
[Github]
This dataset contains a total of 15,000 pieces of real or AI-generated multimodal news (image-caption pairs) -- a training set of 10,000 pairs, a validation set of 2,500 pairs, and five test sets of 500 pairs each. Four of the test sets are out-of-domain data from unseen news publishers and image generators to evaluate detector's generalization ability.
=== Data Source (News Publisher + Image Generator)… See the full description on the dataset page: https://huggingface.co/datasets/Gouge666/mirage-news.google-landmark-geo
Dataset Card for Geo Coordinate Augmented Google-Landmarks
Geo coordinates were added as data to a tar file's worth of images from the Google Landmark V2. Not all of the
images could be geo-tagged due to lack of coordinates on the image's wikimedia page.
Dataset Details
Dataset Description
Geo coordinates were added as data to a tar file's worth of images from the Google Landmark V2. There were many more images that could have
been downloaded but this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/google-landmark-geo.ChildGaitDecoding Children's Gait Behavior
ECCV 2026
Yifan Shen1,2,*,
Boyi Li1,*,
Meihuan Huang2,3,4,*,
Yuanzhe Liu1,*,
Xu Cao1,2,*,§,
Jinyang Jin1,
Zhengyuan Li1,
Anglin Liu5,
Junho Kim1,
Jingyuan Zhu2,
Fangzhou Lan2,
Jianguo Cao2,3,
Jintai Chen5,
Ismini Lourentzou1,
James M. Rehg1,†
1 University of Illinois Urbana-Champaign
2 PediaMed AI
3 Shenzhen Children's Hospital
4 Hong Kong Polytechnic University… See the full description on the dataset page: https://huggingface.co/datasets/gopichand143/ChildGait.Good_Tiresorislop-youtube-prepared-v2
Orislop YouTube prepared v2
This public research dataset contains compact NPZ v2 samples inside uncompressed
WebDataset tar shards. It has 29510 prepared samples from
8698 source videos; 57 sources were
recorded as terminal skips.
Label warning
The real and fake values are trusted scraper assumptions, not independently
verified ground truth. Pre-2021 videos are proxy-real and creator-disclosed altered
videos are proxy-fake. Do not report model agreement as… See the full description on the dataset page: https://huggingface.co/datasets/gonnerthetooner/orislop-youtube-prepared-v2.A-MNISTThe dataset is built on top of MNIST.
It consists from 130K of images in 10 classes - 120K training and 10K test samples.
The training set was augmented with additional 60K images.war-gov-uap-release-1
Department of War UAP Release 1 — structured corpus
The first tranche of declassified U.S. government records on Unidentified
Anomalous Phenomena (UAP / UFOs), released by the Department of War on
8 May 2026 under the Presidential Unsealing and Reporting System for
UAP Encounters (PURSUE) directive.
This dataset is a structured, machine-readable companion to the source
material at https://www.war.gov/UFO/. It pairs every original document
with VLM-extracted page text, cropped… See the full description on the dataset page: https://huggingface.co/datasets/MTSlive/war-gov-uap-release-1.testing-goldstandard-cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/jrw2989/testing-goldstandard-cuthill.Curated_GoldStandard_Hoyal_Cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Curated_GoldStandard_Hoyal_Cuthill.golf-courses
Dataset Summary: golf-course
This dataset (bethecloud/golf-courses) includes 21 unique images of golf courses pulled from Unsplash.
The dataset is a collection of photographs taken at various golf courses around the world. The images depict a variety of scenes, including fairways, greens, bunkers, water hazards, and clubhouse facilities. The images are high resolution and have been carefully selected to provide a diverse range of visual content for fine-tuning a machine learning… See the full description on the dataset page: https://huggingface.co/datasets/bethecloud/golf-courses.
