datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.rlbenchfail_val_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.GeoChrono-Data
ChronoBench & ChronoInstruct
ChronoBench is a comprehensive, multi-dimensional, and multi-granularity benchmark for high-resolution long-temporal remote sensing understanding. It decomposes long-term remote sensing understanding into a four-level cognitive hierarchy — from Land Cover Perception through Temporal Recognition and Long-Term Memory to Spatio-Temporal Reasoning — comprising 12 sub-tasks and 17,689 rigorously validated QA pairs derived from 3,469 high-resolution… See the full description on the dataset page: https://huggingface.co/datasets/Davidup1/GeoChrono-Data.PALL-VLM-data
PALL-VLM-data — Dental Vision-Language Dataset
The training dataset for Harisundar/PALL-VLM,
a dental vision-language model. It contains 32,884 records over 52,461 images,
formatted as image+text conversations for LLaVA-style instruction tuning.
Curated by: Harisundar R
Used by: Harisundar/PALL-VLM · PALL on GitHub
Language: English
Layout
vlm_train/
├── images/ # 52,461 dental images
├── train.jsonl # 29,667 records
├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.Guardian-FailCoT-OOD-datasets
Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks
This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026):
UR5-Fail — our newly collected three-view real-robot benchmark.
RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023).
RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.bdv2fail_train_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.VinDR-CXR-VQA
VinDr-CXR-VQA Dataset
Dataset Description
VinDr-CXR-VQA is a large-scale chest X-ray Visual Question Answering (VQA) dataset designed for explainable medical AI with spatial grounding capabilities. The dataset combines natural language question-answer pairs with bounding box annotations and clinical reasoning explanations.
Key Features
🏥 4,394 chest X-ray images from VinDr-CXR
💬 17,597 question-answer pairs across 6 question types
📍 Spatial… See the full description on the dataset page: https://huggingface.co/datasets/Dangindev/VinDR-CXR-VQA.ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.ur5fail_train_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.ur5fail_val_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_val_dataset.bdv2fail_val_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_val_dataset.MM-Bench-E-CommerceThis is the HuggingFace repository of the paper named MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding in WSDM 2026 (oral).
In this paper, we argue that generative Multimodal Large Language Models (MLLMs) hold significant potential for improving product representation learning.
We propose the first generative MLLM-based model named MOON for product representation learning.
Furthermore, we contruct and publish a large-scale real-world… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/MM-Bench-E-Commerce.orion-dataset
Dataset
The ORION dataset is a curated collection of satellite imagery and triage labels used to fine-tune the VLM for orbital image classification. Images are fetched from SimSat's Mapbox API and paired with classification prompts and ground-truth labels.
Dataset Structure
images/
low_ocean_pacific_nemo.png
med_city_chicago.png
high_port_rotterdam.png
...
train_dataset.jsonl
val_dataset.jsonl
test_dataset.jsonl
images/: 512x512 RGB satellite images fetched from… See the full description on the dataset page: https://huggingface.co/datasets/Saransh-cpp/orion-dataset.bdv2fail_test_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.Weitaikang_bench_data
HeroFrame-Bench
A benchmark for key-frame selection that scores a selected frame directly,
without routing it through a question-answering model and without requiring it
to match a fixed reference set.
204 films, 2,031 frozen comparison chains, 1,970 learned criteria, 30,465
pairwise judgements.
What problem this addresses
A film is almost always encountered first as a single still: a cover, a
thumbnail, a poster. Producing that still from the film is the task… See the full description on the dataset page: https://huggingface.co/datasets/weitaikang/Weitaikang_bench_data.solar-flare-hmi-datasplitsThis dataset is intended to be used for training/testing solar flare forecasting models. It contains various data splits (in json format) of SDO/HMI magnetogram images compiled by
Boucheron, L.E., et al., 2023, Sci Data 10, 825, https://doi.org/10.1038/s41597-023-02628-8.
Splits "train", "val", "test" corresponds to the original data splits provided by Boucheron et al., while the other splits are created by downsampling the No-Flare and C flare class to obtain
more balanced splits and… See the full description on the dataset page: https://huggingface.co/datasets/inaf-oact-ai/solar-flare-hmi-datasplits.Pumpkin-Maturity-Grading-Dataset
Pumpkin Maturity Grading Dataset
The current agricultural industry faces challenges in quality control of crops, especially in judging the maturity of pumpkins. Traditional methods often rely on manual identification, which is inefficient and prone to errors. Although existing image recognition technology has made some progress, there is a lack of high-quality datasets specifically targeted at pumpkin maturity. This dataset aims to improve the accuracy of maturity assessment for… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pumpkin-Maturity-Grading-Dataset.lanternfly_research_dataset
Lantern Fly Research Dataset
This dataset contains human-verified spotted lanternfly sightings collected through the Lantern Fly Tracker app. Each entry includes high-quality photos, precise geolocation data, and comprehensive metadata for ecological research.
🎯 Purpose
This dataset supports:
Ecological research on spotted lanternfly distribution and spread patterns
Machine learning model training with verified, high-quality data
Temporal and spatial analysis of… See the full description on the dataset page: https://huggingface.co/datasets/rlogh/lanternfly_research_dataset.ourdream-database
OurDream Character Database
Complete database of 20,884 AI characters from ourdream.ai, categorized as Women (12,492) and Trans (8,392).
Dataset Structure
database/all_characters.json — Full character database with 20,884 entries
Each entry includes: id, displayId, name, gender, style, age, likeCount, messageCount, tags, shortDescription, thumbUrl, category
Categories
Category
Count
Women
12,492
Trans
8,392
Total
20,884… See the full description on the dataset page: https://huggingface.co/datasets/lcuifer0/ourdream-database.Raspberry-Variety-Classification-Dataset
Raspberry Variety Classification Dataset
The current agricultural industry faces challenges in managing the diversity of crop types, especially in raspberry cultivation and variety identification. Existing solutions often rely on manual identification, which is inefficient and prone to errors. This dataset aims to address the issue of low accuracy in variety classification by providing high-quality raspberry images, meeting the needs of intelligent agriculture. The dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Raspberry-Variety-Classification-Dataset.Pesticide-Spraying-Scene-Classification-Dataset
Pesticide Spraying Scene Classification Dataset
The current agricultural industry faces challenges such as low spraying efficiency and environmental pollution, especially during large-scale farmland spraying. Traditional methods rely on manual operations, which can lead to pesticide waste and uneven spraying. Existing solutions often lack efficient image recognition technology and cannot monitor spraying effectiveness and crop conditions in real-time. This dataset aims to support… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pesticide-Spraying-Scene-Classification-Dataset.Asparagus-Identification-Dataset
Asparagus Identification Dataset
The current agricultural industry faces challenges in efficient crop monitoring and recognition. Traditional manual detection methods are inefficient and prone to errors. Existing solutions often rely on empirical judgment without scientific model support. This dataset aims to provide diverse asparagus images to help train automatic recognition models, improving the accuracy and efficiency of crop monitoring. The dataset includes images of asparagus… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Asparagus-Identification-Dataset.Experiment-Design-Sketch-Image-Classification-Dataset
Experiment Design Sketch Image Classification Dataset
In the field of industrial manufacturing, the design process often relies on a large number of design sketches that need to be quickly converted into actual engineering designs during the subsequent manufacturing stages. However, manually processing these sketches is often time-consuming and prone to errors, currently relying mainly on manual labeling and conversion by designers, which is inefficient and unstable. Existing… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Experiment-Design-Sketch-Image-Classification-Dataset.Pumpkin-Maturity-Grading-Dataset
Pumpkin Maturity Grading Dataset
The current agricultural industry faces challenges in quality control of crops, especially in judging the maturity of pumpkins. Traditional methods often rely on manual identification, which is inefficient and prone to errors. Although existing image recognition technology has made some progress, there is a lack of high-quality datasets specifically targeted at pumpkin maturity. This dataset aims to improve the accuracy of maturity assessment for… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Pumpkin-Maturity-Grading-Dataset.Papaya-Tree-Recognition-Dataset
Papaya Tree Recognition Dataset
The current agricultural sector faces issues of low efficiency in crop recognition and management, especially against the backdrop of the growing development of smart agriculture. Traditional manual recognition methods are unable to meet the rapidly changing needs. Existing solutions often rely on image data from a single environment, lacking diversity and universality, which leads to poor model generalization. The Papaya Tree Recognition Dataset aims… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Papaya-Tree-Recognition-Dataset.Field-Pumpkin-Recognition-Dataset
Field Pumpkin Recognition Dataset
The current agriculture industry faces challenges such as low efficiency in crop recognition and difficulty in pest monitoring. Existing solutions often rely on manual identification, which is inefficient and prone to errors. This dataset aims to support the training of machine learning models by providing high-quality pumpkin image data, enhancing the accuracy and speed of pumpkin recognition. Data is primarily collected in the field using… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Field-Pumpkin-Recognition-Dataset.ui-design-audit-dataset
UI Design Audit Screenshot Benchmark v2.1
A reproducible synthetic benchmark of 3,000 mobile and web UI screenshots labeled across 12 design-risk categories.
Splits
train: 2,400
validation: 300
test: 300
Labels
small_touch_targets
low_contrast
action_overload
navigation_overload
form_friction
content_density
responsive_risk
modal_overuse
deep_scrolling
weak_hierarchy
interaction_overload
mobile_web_mismatch
Data creation
Every… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/ui-design-audit-dataset.
