datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BioTrove
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
Description
See the BioTrove-Train dataset card on HuggingFace to access the samller BioTrove-Train dataset (40M)
BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove.BioTrove-Train
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
Description
See the BioTrove dataset card on HuggingFace to access the main BioTrove dataset (161.9M)
BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of software… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove-Train.biodex_sprint1
BioDex Aviary Birds — Sprint 1
Image-classification dataset for the 55 bird species held in the Parque aviary
(2026 catalogue). Built for training a species classifier that runs on photos
visitors and keepers take on-site.
75,082 images · 55 classes · 224×224 RGB JPEG · train / val / test splits ·
GPU data-augmentation on the training split.
Each column is one source image — top: original, below: its two augmented copies (rotation, lighting, motion blur, flip).
How… See the full description on the dataset page: https://huggingface.co/datasets/santianwandter/biodex_sprint1.Biomedica2025EvalSet
Biomedica 2025 Eval Set
Unified test snapshot of the BioMedica 2025 vision–language evaluation
suites used in AMInZeroShotOpenEvalAllTasks. Every row is a single image
with closed-ended options, the gold answer, and provenance fields that
point back to the original dataset.
Images are stored as original JPEG/PNG bytes (or JPEG-encoded arrays) inside
parquet so the Hugging Face dataset viewer is enabled
(~1959 MB download, 89941 examples).
Suites
config… See the full description on the dataset page: https://huggingface.co/datasets/Alejandro98/Biomedica2025EvalSet.BIOSCAN-30k
Dataset Card for BIOSCAN-30k
This is a FiftyOne dataset with 30000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/BIOSCAN-30k")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details
The BIOSCAN-5M… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/BIOSCAN-30k.bioscan-traits
Dataset Card for BIOSCAN-Traits
Dataset Details
Dataset Description
BIOSCAN-Traits is a trait-level annotation dataset for fine-grained insect imagery. Derived from BIOSCAN-5M, it provides morphology-centric natural language trait descriptions automatically generated by a two-stage pipeline: (1) a Sparse Autoencoder (SAE) trained on DINOv2 visual features identifies species-level salient visual parts (wings, legs, antennae, etc.), and (2) a Multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/bioscan-traits.minecraft-biomes
Minecraft Biomes (RGBD, pseudo-labeled)
Pseudo-labeled RGBD screenshots from Minecraft, covering 12 broad biome
categories. Each sample is an (RGB, depth) pair at 640×360 resolution.
Source
RGBD frames: 1908 paired (rgb, depth) samples from
zid8/syntheticMinecraftRGBD,
collected via MineRL.
Labels: Generated by a Gemma-3-4B model fine-tuned with LoRA on
the willowc/minecraft-biomes
dataset, then augmented with ~160 hand-selected MineRL ocean frames
to fix an… See the full description on the dataset page: https://huggingface.co/datasets/Wafik20/minecraft-biomes.Handwritten-Biology-Notes-Dataset
English Handwritten Biology Notes Dataset
This dataset contains high-resolution images of handwritten biology notes written in English. The collection includes labeled diagrams, definitions, explanations of biological processes, and annotated sketches. It supports AI research in handwriting recognition, diagram understanding, and document interpretation within the field of life sciences.
Contact
For queries or collaborations related to this dataset, contact:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Biology-Notes-Dataset.biometric-fingerprint-spoofing
Fingerprint Spoofing
The dataset comprises 5,000+ high-quality fingerprint images collected from 100 individuals (ten fingers per person) captured using multiple fingerprint scanners and biometric sensors. Designed for spoofing detection and liveness detection tasks, the fingerprint dataset provides labeled biometric data from different devices and fingers to train and evaluate biometric security and fingerprint recognition systems.
By utilizing this data, researchers and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/biometric-fingerprint-spoofing.2-000-Physics-Chemistry-And-Biology-Questions-With-Image-Explanations-Dataset
2,000 Physics, Chemistry, and Biology Questions with Image Explanations Dataset
This collection contains 2,000 high-quality physics, chemistry, and biology questions, featuring a core modality of original images paired with text explanations. The data covers various formats, including multiple-choice, fill-in-the-blanks, experimental, and calculation questions. Each record provides the original question image, precise OCR-extracted text, and detailed step-by-step textual… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/2-000-Physics-Chemistry-And-Biology-Questions-With-Image-Explanations-Dataset.human-iris-images-biometric
Human Iris Dataset
The dataset comprises 5,000+ high-quality iris images from 5,000+ individuals, captured for iris recognition and biometric tasks, with each person contributing left and right eye images to enable verification and identification algorithms. It supports classification of ocular structures, detection of attacks, and analysis of colors, textures, and unique irises for image processing and iris pattern analysis tasks.
By utilizing this dataset researchers and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-iris-images-biometric.NIH-CXR14-BiomedCLIP-Features
NIH-CXR14-BiomedCLIP-Features Dataset
This dataset is derived from the NIH Chest X-ray Dataset (NIH-CXR14) and processed using the BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 model from Microsoft. It contains image and text features extracted from chest X-ray images and their corresponding textual findings.
Dataset Description
The original NIH-CXR14 dataset comprises 112,120 chest X-ray images with disease labels from 30,805 unique patients. This processed dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yasintuncer/NIH-CXR14-BiomedCLIP-Features.body-measurements-image-dataset
Body Measurements Image Dataset - 13,000 Images
Dataset consists of 13,000+ standardized photos of 1,000+ people, offering a robust resource for body measurements estimation, human body analysis, and personalized sizing recommendations in e-commerce. Each subject is captured in front and side poses with paired 17+ anthropometric measurements, enabling precise shape estimation, weight prediction, and body characteristics detection.— Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/body-measurements-image-dataset.biodex
BioDex
Bird species image classification dataset: 55 species, ~28,752 images, built by downloading observation photos from GBIF and iNaturalist, supplemented with a small number of images from the Birds 525 Species Kaggle dataset for a few underrepresented species.
Dataset structure
imagefolder format: one directory per species, directory name = class label.
Single train split. Train/val/test splits are not defined yet — held back until data augmentation is… See the full description on the dataset page: https://huggingface.co/datasets/santianwandter/biodex.moth_biotrove
Dataset Card for moth_biotrove
This is a FiftyOne dataset with 1000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("pjramg/moth_biotrove")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/pjramg/moth_biotrove.open-palm-hand-images
Palm Dataset - 500,000 Images
Dataset comprises 500,000 high-quality images featuring diverse human hands, specifically designed for hand detection, palm recognition, and gesture analysis. It provides diverse training data with metadata on age, gender, and ethnicity for accurate computer vision model training.— Get the data
Dataset characteristics:
Characteristic
Data
Description
Open palm images designed for training and evaluating hand-based… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/open-palm-hand-images.Anti-Spoofing-Real-Videos
Face Anti Spoofing Dataset - 98 000+ files
Dataset features 98,000+ files of real photos and videos of people from 170+ countries, representing 70,000+ unique individuals. By leveraging this dataset, developers can enhance spoofing detection techniques, improve recognition systems, and deploy anti-spoofing algorithms capable of preventing fraud in deep learning-based solutions.- Get the data
Dataset characteristics:
Characteristic
Data
Description
Live… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/Anti-Spoofing-Real-Videos.schulz_bank_biotopes
Summary
Image classification dataset with biotope labels extracted from Meyer et al., 2022.
Images were extracted from six remotely operated vehicle (ROV) dives during two SponGES cruises conducted in the summers of 2017 and 2018. The ROV dives were performed by Ægir 6000 and the cruises were conducted on the RV G.O. Sars. The ROV dives traversed across various regions on the seamount, from the base of the seamount at 2700 m depth to the summit at 580 m depth. 600 images were… See the full description on the dataset page: https://huggingface.co/datasets/CGame1/schulz_bank_biotopes.biometric-fingerprint-spoofing
Fingerprint Spoofing Dataset 5,000 photos
Dataset contains 5,000+ high-quality fingerprint images capturing real fingerprints and multiple spoofing fingerprint attack types, including print and replay scenarios. Designed for spoofing detection and liveness detection tasks, the fingerprint dataset provides labeled biometric data from different devices and fingers to train and evaluate biometric security and fingerprint recognition systems.- Get the data
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/biometric-fingerprint-spoofing.autotrain-data-beccacp
AutoTrain Dataset for project: beccacp
Dataset Description
This dataset has been automatically processed by AutoTrain for project beccacp.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<1600x838 RGB PIL image>",
"target": 1
},
{
"image": "<1200x628 RGB PIL image>",
"target": 1
}]
Dataset Fields
The… See the full description on the dataset page: https://huggingface.co/datasets/Bioskop/autotrain-data-beccacp.hand-gesture-recognition
Movement Recognition Dataset - 10k+ videos
Dataset comprises 10,000+ videos of people demonstrating 5 distinct hand gestures, designed for advancing gesture recognition systems and improving recognition technology. It is ideal for training real-time recognition software, improving control systems, and exploring 3D models of hand pose and finger movements. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Each video shows a person… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/hand-gesture-recognition.ear-detection-dataset
Ear Detection - 14,000+ Images
The dataset comprises 14,000+ ear images from 2,000 unique individuals, paired with reference face photos and demographic labels. Designed for ear recognition and biometric identification, it helps research in human ear detection, recognition accuracy, and biometric systems.— Get the data
Dataset characteristics:
Characteristic
Data
Description
Images of an ear with a face photo
Data types
Image
Tasks
Ear‑biometric R&D… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/ear-detection-dataset.Selfie-and-ID-Dataset
Document Dataset - 65,000+ photos
The dataset contains 65,000+ photos of 5,000+ people from 40+ countries, featuring selfie images paired with identity documents for facial recognition, Know Your Customer (KYC), and re-identification tasks. It is designed to enhance verification systems, improve recognition models, and support biometric research. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of individuals and their… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/Selfie-and-ID-Dataset.kids-and-teens-selfie-dataset
Age Estimation - 6,000 Photos
Dataset contains 9,000 high-quality facial images of children and teenagers aged 7–15, designed for age estimation, facial recognition, and anti-spoofing research. Its primary application is supporting the development of robust age estimation models, improving facial analysis for younger demographics, and studying social media usage patterns. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/kids-and-teens-selfie-dataset.Biology-Images-DatasetDataset Description:
This Biology image dataset is part of a large-scale STEM image dataset containing 383,697 images, designed to support the development and training of advanced computer vision, educational AI, and multimodal learning systems.
Additionally, this dataset can be used in pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, improving model performance in scientific image understanding, visual reasoning, concept recognition… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Biology-Images-Dataset.biome-themed-pantanal-biopark-fish-tanks
Biome-Themed Pantanal Biopark Fish Tanks
This dataset contains 3654 images taken in the Pantanal Biopark.
A paper with further details and baseline results is to be published soon in PLOS ONE.
multi-material-fingerprint-spoofing
Fingerprint Spoofing Dataset 4,000 photos
The dataset contains 4,000+ fingerprint images from 100 individuals, captured with a ZKTeco ZK9500 optical scanner and including real fingerprints and spoofing attacks created with alginate, plasticine, and silicone materials. It includes metadata (gender, age, finger, hand, device) and supports biometric security research, presentation attack detection, spoof detection, and fingerprint recognition model training.- Get the data… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/multi-material-fingerprint-spoofing.biometric
DepositPhotos Biometric & Human Features Dataset (Sample)
Overview
This dataset is a curated sample of high-resolution Biometric, Facial, and Human Feature images sourced from the DepositPhotos library. This subset is designed for training, testing, and evaluating Computer Vision models for facial recognition, identity verification, emotional analysis, and healthcare AI applications.
This sample demonstrates the exact quality, diversity, and structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/Depositphotos/biometric.
