datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BioTrove
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
Description
See the BioTrove-Train dataset card on HuggingFace to access the samller BioTrove-Train dataset (40M)
BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove.biomedical_lectures_v2
Vidore Benchmark 2 - MIT Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_v2.MedThinkVQA
MedThinkVQA
MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning.
Links
GitHub: https://github.com/benluwang/MedThinkVQA
Leaderboard: https://benluwang.github.io/MedThinkVQA/
Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.HESRT
HESRT
Human-reviewed, accession-level spatial-omics artifacts with tissue images, expression matrices and source provenance.
Samples (GSMs)
Studies (GSEs)
Expression observations
Compressed packages
4,468
507
37,054,005
413.94 GB
Each sample is a self-contained, checksummed tar.zst package. Browse the catalog, select the accessions you need, then download those samples. The full collection contains approximately 647.72 GB of uncompressed member data; do not clone… See the full description on the dataset page: https://huggingface.co/datasets/Biogod/HESRT.biorXiv-pdf
BiorXiv Pdf
BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets.
BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.BioTrove-Train
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
Description
See the BioTrove dataset card on HuggingFace to access the main BioTrove dataset (161.9M)
BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of software… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove-Train.bioburden_labelled_dataBIOSCAN-1M
BIOSCAN-1M
Overview
The BIOSCAN-1M dataset offers researchers detailed information about insects, with each record containing four primary attributes:
DNA Barcode Sequence
Barcode Index Number (BIN)
Biological Taxonomy Classification
RGB image
Citation
If you make use of the BIOSCAN-1M dataset and/or its code repository, please cite the following paper:
cite as:
@inproceedings{gharaee2023step,
title={A Step Towards Worldwide Biodiversity… See the full description on the dataset page: https://huggingface.co/datasets/bioscan-ml/BIOSCAN-1M.vannamei-shrimp-biomass-dataset
Litopenaeus vannamei Shrimp Biomass Dataset (mirror)
This is a mirror of the original dataset published on Mendeley Data.
It is not my data — all credit goes to the original authors. Re-hosted
here under the terms of the CC BY 4.0 license for easier programmatic access
(Kaggle / Hugging Face datasets loading).
Original source
Ramírez-Coronel, F.J., Esquer-Miranda, E., Rodríguez-Elías, O.M.,
García-Hinostro, P., Parra-Salazar, G.C. (2024).
"A Litopenaeus vannamei… See the full description on the dataset page: https://huggingface.co/datasets/DeepanSadhukhan/vannamei-shrimp-biomass-dataset.guertin-mcro-forensic-corpus-biography
Guertin MCRO Forensic Corpus: Biography
Contents: 44 projects from Matthew Guertin's portfolio (2006–2023) with 2,615 media files; timeline-data.js lists the projects (title, year, media); story-biography/ holds the biography page's exhibits.
Layout: timeline_files/<project>/; project folders that repeat the same files keep one copy.
Provenance: the portfolio as published on mattguertin.com and on mncourtfraud.com (/story/).
Integrity: manifest.tsv lists every file with its… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-biography.biodex_sprint1
BioDex Aviary Birds — Sprint 1
Image-classification dataset for the 55 bird species held in the Parque aviary
(2026 catalogue). Built for training a species classifier that runs on photos
visitors and keepers take on-site.
75,082 images · 55 classes · 224×224 RGB JPEG · train / val / test splits ·
GPU data-augmentation on the training split.
Each column is one source image — top: original, below: its two augmented copies (rotation, lighting, motion blur, flip).
How… See the full description on the dataset page: https://huggingface.co/datasets/santianwandter/biodex_sprint1.Biomedica2025EvalSet
Biomedica 2025 Eval Set
Unified test snapshot of the BioMedica 2025 vision–language evaluation
suites used in AMInZeroShotOpenEvalAllTasks. Every row is a single image
with closed-ended options, the gold answer, and provenance fields that
point back to the original dataset.
Images are stored as original JPEG/PNG bytes (or JPEG-encoded arrays) inside
parquet so the Hugging Face dataset viewer is enabled
(~1959 MB download, 89941 examples).
Suites
config… See the full description on the dataset page: https://huggingface.co/datasets/Alejandro98/Biomedica2025EvalSet.BIOSCAN-30k
Dataset Card for BIOSCAN-30k
This is a FiftyOne dataset with 30000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/BIOSCAN-30k")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details
The BIOSCAN-5M… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/BIOSCAN-30k.BioD2bark-ambrosia-beetle-benchmark
Bark and Ambrosia Beetle Detection Benchmark
Version 2.0.1 · 14,491 images · 175 species · 21 tribes · 70 genera · COCO detection format
A specimen-disjoint, species-level object-detection benchmark for bark and ambrosia
beetles (Coleoptera: Curculionidae: Scolytinae and Platypodinae), derived from the
Bark and Ambrosia Gallery (https://barkandambrosiagallery.org/). Species
determinations are made or reviewed by taxonomists; individual specimens carry
bounding boxes.
This Zenodo… See the full description on the dataset page: https://huggingface.co/datasets/IBBI-bio/bark-ambrosia-beetle-benchmark.safestep-regional-bio-vision
SAFEstep Regional Biological-Hazard Vision
This package trains a small biological-hazard object detector for an offline
mobile safety app. Its initial visual classes are snake, mushroom, spider,
scorpion and jellyfish. It is deliberately not a venom, toxicity or edibility
classifier and never tells a user that a photographed organism is safe.
Possible snake detected. Keep your distance. Do not touch, trap, or approach
it. If anyone was bitten, follow the app's snakebite… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/safestep-regional-bio-vision.ibbi_ood_data
Dataset Card for IBBI Out-of-Distribution (OOD) Dataset
Dataset Summary
This dataset contains out-of-distribution (OOD) images of bark and ambrosia beetles, intended for evaluating the robustness and generalization capabilities of models from the ibbi Python package. The 121 species included in this dataset were not part of the original training or in-distribution test sets for the models in the ibbi package.
This dataset is crucial for testing how well models can:… See the full description on the dataset page: https://huggingface.co/datasets/IBBI-bio/ibbi_ood_data.genomic-bioimaging
Clustering phenotype populations by genome-wide RNAi and multiparametric imaging
52224 images of a high-content RNAi knockdown screening.
Author one-liner: "To cluster genes and predict function on a genome-wide scale, we measured the effects of 22 839 siRNA-mediated knockdowns on HeLa cells. Each siRNA effect was summarized by a phenotypic profile."
Dataset Details
Organism: human (Homo sapiens)
Cell type: HeLa
Imaging method: Fluorescence microscopy
Study… See the full description on the dataset page: https://huggingface.co/datasets/stefanches/genomic-bioimaging.bioscan-traits
Dataset Card for BIOSCAN-Traits
Dataset Details
Dataset Description
BIOSCAN-Traits is a trait-level annotation dataset for fine-grained insect imagery. Derived from BIOSCAN-5M, it provides morphology-centric natural language trait descriptions automatically generated by a two-stage pipeline: (1) a Sparse Autoencoder (SAE) trained on DINOv2 visual features identifies species-level salient visual parts (wings, legs, antennae, etc.), and (2) a Multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/bioscan-traits.biomasstersRehosted dataset from official repo. However, we have created tar balls instead of split zip archives, since they are easier to extract.
We also changed the nested directory structure of the train_features directory and renamed the .csv files.
The tar balls can be extracted with theses commands:
tar -xvf train_agbm.tar.gz
tar -xvf test_agbm.tar.gz
cat train_features.tar.gz* | tar xvfz -
cat test_features.tar.gz* | tar xvfz -
You can also check the accompanying dataloader in torchgeo.
If you… See the full description on the dataset page: https://huggingface.co/datasets/torchgeo/biomassters.hle-gold-bio-chem
Humanity's Last Exam (HLE) Bio/Chem Gold
Humanity’s Last Exam (HLE) is a challenging question-answering AI benchmark covering advanced academic fields including Math, Physics, Chemistry, Biology, Engineering, and Computer Science.
At FutureHouse, we audited the biology and chemistry subsets of HLE using a combination of expert human evaluators and our in-house research agent, and found that around 30% of the questions contain answers directly contradicted by peer-reviewed… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/hle-gold-bio-chem.biomedical_lectures_eng_v2
Vidore Benchmark 2 - MIT Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_eng_v2.sf_biodiv_accessBioVITA
Citation
@inproceedings{shinoda2026biovita,
title = {BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment},
author = {Risa Shinoda and Kaede Shiohara and Nakamasa Inoue and Kuniaki Saito and Hiroaki Santo and Fumio Okura},
booktitle = {CVPR},
year = {2026},
}
minecraft-biomes
Minecraft Biomes (RGBD, pseudo-labeled)
Pseudo-labeled RGBD screenshots from Minecraft, covering 12 broad biome
categories. Each sample is an (RGB, depth) pair at 640×360 resolution.
Source
RGBD frames: 1908 paired (rgb, depth) samples from
zid8/syntheticMinecraftRGBD,
collected via MineRL.
Labels: Generated by a Gemma-3-4B model fine-tuned with LoRA on
the willowc/minecraft-biomes
dataset, then augmented with ~160 hand-selected MineRL ocean frames
to fix an… See the full description on the dataset page: https://huggingface.co/datasets/Wafik20/minecraft-biomes.biolit-coastal-species-grounding-truthBioMedFlickr
BioMedFlickr
BioMedFlickr is a biomedical image–caption retrieval benchmark built from
public Flickr pathology / microscopy albums. Each example is a single image
paired with a cleaned English caption. Images are stored as original JPEGs inside parquet (~865 MB download). This dataset is the retrieval benchmark used in BIOMEDICA (Lozano et al.), published at CVPR 2025: https://cvpr.thecvf.com/virtual/2025/poster/33761
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/BioMedFlickr.livecellibbi_test_data
Dataset Card for IBBI Bark Beetle Testing Dataset
Dataset Summary
This dataset is the primary testing and benchmarking set for the ibbi Python package. It contains images of bark and ambrosia beetles used to evaluate the performance of object detection and classification models.
Note: While this dataset serves as the testing set for the ibbi package's evaluation functions, it is hosted on the Hugging Face Hub as the train split. You can access it using… See the full description on the dataset page: https://huggingface.co/datasets/IBBI-bio/ibbi_test_data.Handwritten-Biology-Notes-Dataset
English Handwritten Biology Notes Dataset
This dataset contains high-resolution images of handwritten biology notes written in English. The collection includes labeled diagrams, definitions, explanations of biological processes, and annotated sketches. It supports AI research in handwriting recognition, diagram understanding, and document interpretation within the field of life sciences.
Contact
For queries or collaborations related to this dataset, contact:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Biology-Notes-Dataset.
