datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet-1k-vl-enriched
Visualize on Visual Layer
Imagenet-1K-VL-Enriched
An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and label issues!
With this additional information, the ImageNet-1K dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues helps to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: The original filename of the image from… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched.scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
oxford-iiit-pet-vl-enriched
Visualize on Visual Layer
Oxford-IIIT-Pets-VL-Enriched
An enriched version of the Oxford IIIT Pets Dataset with image caption, bounding boxes, and label issues!
With this additional information, the Oxford IIIT Pet dataset can be extended to various tasks such as image retrieval or visual question answering.
The label issues help to curate a cleaner and leaner dataset.
Description
The dataset consists of 6 columns:
image_id: Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/visual-layer/oxford-iiit-pet-vl-enriched.scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/prabhats0605/scanned-images-dataset-for-ocr-and-vlm-finetuning.PALL-VLM-data
PALL-VLM-data — Dental Vision-Language Dataset
The training dataset for Harisundar/PALL-VLM,
a dental vision-language model. It contains 32,884 records over 52,461 images,
formatted as image+text conversations for LLaVA-style instruction tuning.
Curated by: Harisundar R
Used by: Harisundar/PALL-VLM · PALL on GitHub
Language: English
Layout
vlm_train/
├── images/ # 52,461 dental images
├── train.jsonl # 29,667 records
├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace.
The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels.
The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.citrus-disease-vlm-instruct
Citrus Disease VLM Instruct
An instruction-tuning dataset for training a small vision-language model (VLM) to look at a photo of a citrus leaf, fruit or shoot, name the disease, pest or nutrient deficiency, explain the cause and symptoms, and recommend both biological/organic and chemical management.
Every example pairs one image with a chat conversation in the format used by TRL's SFTTrainer for multimodal models (Qwen-VL, SmolVLM, Idefics, LLaVA and similar).
What… See the full description on the dataset page: https://huggingface.co/datasets/ML-Intern-lab/citrus-disease-vlm-instruct.iconclass-vlm-brillfull
Iconclass VLM — brill full labels
Training-ready VLM iconclass-classification dataset rebuilt from the fuller, cleaner
source labels in biglam/brill_iconclass
(CC0). Recovers labels lost to truncation in davanstrien/iconclass-vlm-sft.
Source images: same Brill Arkyves images as biglam/brill_iconclass, bytes passed through verbatim (no re-encode).
Labels: full Iconclass codes with operators (+n), key-combos :, and qualifiers (TEXT) kept intact. Empty/sentinel tokens stripped; ~5… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/iconclass-vlm-brillfull.medical-vlm-unlearning-corpus
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-corpus.pitvqa-unified-vlm
PitVQA Unified VLM Classification Dataset
Surgical workflow classification dataset for training vision-language models on pituitary surgery phase detection, step recognition, and instrument identification.
🔗 GitHub: https://github.com/matheus-rech/pit_project
🤖 Trained Model: mmrech/pitvqa-qwen2vl-unified
📄 Original Dataset: UCL Research Data Repository
Dataset Description
This dataset contains 5,184 surgical frames with classification annotations for surgical… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/pitvqa-unified-vlm.tomato-leaves-dataset
Tomato Leaves Dataset
Overview
This dataset contains images of tomato leaves categorized into different classes based on the type of disease or health condition. The dataset is divided into training, validation, and test sets, with a ratio of 8:1:1. The classes include various diseases as well as healthy leaves. The dataset includes both augmented and non-augmented images.
Dataset Structure
The dataset is organized into three main splits:
train
validation
test… See the full description on the dataset page: https://huggingface.co/datasets/Vlone571/tomato-leaves-dataset.medical-vlm-unlearning-incremental-subset
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-incremental-subset.GlobalRG-Retrieval
GlobalRG - Retrieval Across Universals Task
Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models' cultural inclusivity, but they have limited coverage of cultures and do not adequately assess cultural diversity across universal as well as culture-specific local concepts. We introduce the GlobalRG-Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/UBC-VL/GlobalRG-Retrieval.brain-mri-plane-aware-vlm
Brain MRI Plane-Aware VLM Dataset
Dataset Description
This dataset is derived from the BRISC 2025 dataset and has been processed specifically for training Vision-Language Models (VLMs) with plane-aware understanding of brain MRI scans.
Source Dataset
Original dataset: BRISC 2025
The BRISC 2025 dataset contains:
6,000 T1-weighted MRI images
Four tumor classes: Glioma, Meningioma, Pituitary Tumor, and No Tumor
Pixel-wise segmentation masks validated by… See the full description on the dataset page: https://huggingface.co/datasets/AhmadIshaqai/brain-mri-plane-aware-vlm.evals-eastrus-vl
evals-eastrus-vl
Independent evaluation dataset for the EstrusVision cattle estrus detection model. Contains ground-truth labels for measuring deployment readiness.
Contents
43 total samples (40 in-domain cattle vulval images, 3 out-of-domain)
Embedded image column (no external file dependencies)
Six symptom ground-truth labels per in-domain sample
out_of_domain flag for rejection testing
notes field with clinical observations
Label distribution (in-domain only)… See the full description on the dataset page: https://huggingface.co/datasets/prapaa/evals-eastrus-vl.brain-mri-plane-aware-vlm
Brain MRI Plane-Aware VLM Dataset
Dataset Description
This dataset is derived from the BRISC 2025 dataset and has been processed specifically for training Vision-Language Models (VLMs) with plane-aware understanding of brain MRI scans.
Source Dataset
Original dataset: BRISC 2025
The BRISC 2025 dataset contains:
6,000 T1-weighted MRI images
Four tumor classes: Glioma, Meningioma, Pituitary Tumor, and No Tumor
Pixel-wise segmentation masks validated… See the full description on the dataset page: https://huggingface.co/datasets/zjj30/brain-mri-plane-aware-vlm.eastrus-vl
eastrus-vl
prapaa/eastrus-vl is a self-contained vision-language dataset built from the labeled cattle vulval images in this repository. The Hugging Face dataset stores the image bytes inside Parquet shards, so it can be loaded anywhere without needing the original local file paths.
What is included
255 training examples
Embedded image column
Raw structured labels for estrus-related symptom analysis
A reusable prompt column for VL fine-tuning
Two text supervision… See the full description on the dataset page: https://huggingface.co/datasets/prapaa/eastrus-vl.mist-vlm-judges
MIST - Misleading-Image Stroop Test for VLM Judges
Anonymous release accompanying a double-blind submission. 200 potentially idiomatic English compounds annotated by two disjoint panels of three human annotators each, plus labels produced by 13 vision-language models under 4 prompting techniques, 3 image conditions, and 2 instruction variants.
Source and construction
We build on a public instruction-tuning release, UCSC-Admire/idiom-SFT-dataset-561, which extends… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/mist-vlm-judges.
