datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
picbreeder-vlm-archive
Picbreeder-VLM Archive
Every image evolved by the swarm of vision-language-model "breeders" in
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
(GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the
lineage graphs, and the analysis artifacts behind the paper and blog.
The original Picbreeder (Secretan et al., 2008) let crowds of
humans collaboratively evolve images from
CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.vlm_test_imagesBunch of random test cases for vision language in the wild.
GEOBench-VLM
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Summary
While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they fall short in addressing the unique demands of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, which is critical for applications such as environmental monitoring, urban planning, and disaster management. Some of the unique… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/GEOBench-VLM.VLMEvalKitvlm_evaluation_v1.0
Datacard
This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
The dataset structure is as follows:
vlm_evaluation_v1.0/
├── CommenSence/
├── add_condiment_common_sense/
├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.Latex-VLMMMEB-V3
MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models
🌐 Website |
GitHub |
🏆 Leaderboard |
📖 MMEB-V3 Paper |
📖 MMEB-V2 Paper |
📖 MMEB-V1 Paper |
🤗 Models
Introduction
MMEB-V3 is a comprehensive benchmark for evaluating omni-modality embedding models across text, image, video, audio, visual-document, and agent-centric retrieval scenarios.
Building upon MMEB-V1 and MMEB-V2, MMEB-V3 adds 111 new tasks, resulting in 190 evaluation tasks in… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MMEB-V3.VLM4Bio
Dataset Card for VLM4Bio
Instructions for downloading the dataset
Install Git LFS
Git clone the VLM4Bio repository to download all metadata and associated files
Run the following commands in a terminal:
git clone https://huggingface.co/datasets/imageomics/VLM4Bio
cd VLM4Bio
Downloading and processing bird images
To download the bird images, run the following command:
bash download_bird_images.sh
This should download the bird images inside datasets/Bird/images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/VLM4Bio.VLM-SubtleBench
VLM-SubtleBench
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/VLM-SubtleBench.MMLongBench-docVista
Dataset Card for "Vista"
"700.000 Vietnamese vision-language samples open-source dataset"
Dataset Overview
This dataset contains over 700,000 Vietnamese vision-language samples, created by Gemini Pro. We employed several prompt engineering techniques: few-shot learning, caption-based prompting and image-based prompting.
For the COCO dataset, we generated data using Llava-style prompts
For the ShareGPT4V dataset, we used translation prompts.
Caption-based prompting:… See the full description on the dataset page: https://huggingface.co/datasets/Vi-VLM/Vista.vlmsareblindArXiv - Website
ViDoSeek-page-fixedMMLongBench-page-fixedViDoSeekKoLLaVA-v1.5-Instruct-581k
KoLLaVA-v1.5-Instruct-581k
한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다.
데이터셋 정보
총 샘플 수: 435,093개
형식: ChatML 형식 (role: user/assistant, content: 텍스트)
이미지: COCO + GQA + Visual Genome 데이터셋
언어: 한국어
포함된 데이터셋
COCO 데이터: 362,953개 샘플
MS COCO 2017 이미지 기반
한국어 대화 데이터
GQA 데이터: 72,140개 샘플
GQA (Visual Question Answering) 이미지 기반
한국어 대화 데이터
Visual Genome 데이터: 포함
Visual Genome 이미지 기반
한국어 대화 데이터
제외된 데이터셋
EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.Nemotron-VLM-Dataset-v2from nvidia/Nemotron-VLM-Dataset-v2
samples are:
visual7w_telling_cot: 435299
plotqa_cot: 295354
wiki_ko: 200000
wiki_en: 200000
mulberry_cot_1: 189378
mulberry_cot_2: 102279
sparsetables: 100000
mantis_instruct_cot: 67714
llava_cot_100k: 63019
visual_web_instruct_cot: 47800
chartqa_cot: 45710
docvqa_cot: 36333
tabmwp_cot: 20305
infographicsvqa_cot: 19548
hiertext: 514
MVBench
MVBench
Forked from https://huggingface.co/datasets/OpenGVLab/MVBench for reproducibility.
Important Update
[18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference.
We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MVBench.scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.nano-omni-vlmvlm_wsi_patch
VLM WSI Patch
Closed-answer WSI patch VQA dataset with 48,766 metadata-eligible rows and
48,609 valid 1344×1344 RGB top-4 mosaics, picked from of PathGen. Rows without an image retain an
explicit processing rejection status. Each patch uses a WSI level-0 top-left coordinate and
672×672 pixels; backfill and patch substitution are excluded.
The image field is a repository-relative PNG path. metadata/manifest.jsonl and
metadata/manifest.parquet contain QA labels, top-4 coordinates… See the full description on the dataset page: https://huggingface.co/datasets/HCOOH/vlm_wsi_patch.vlms-are-biased
Vision Language Models are Biased
by
An Vo1*,
Khai-Nguyen Nguyen2*,
Mohammad Reza Taesiri3,
Vy Tuong Dang1,
Anh Totti Nguyen4†,
Daeyoung Kim1†
*Equal contribution †Equal advising
1KAIST, 2College of William and Mary, 3University of Alberta, 4Auburn University
TLDR: State-of-the-art Vision Language Models (VLMs) perform perfectly on counting tasks with original images but fail catastrophically (e.g., 100% → 17.05%… See the full description on the dataset page: https://huggingface.co/datasets/anvo25/vlms-are-biased.opht_vlms_lowvlm-info-loss-results
VLM Grounding Evaluation Results
Grounding evaluation results for vision-language models on robotics manipulation datasets.
Part of the vlm-info-loss project studying
how VLM connectors transform visual representations.
Background
Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation:
they sharpen dominant-object representations while compressing secondary-object category identity.
All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.nemotron-vlm-datamultilingual-vlm-reasoning
Multilingual VLM Visual Reasoning Benchmark
Paper: Do Multilingual VLMs Reason Equally? A Cross-Lingual Visual Reasoning Audit for Indian Languages
Overview
This dataset contains ~1,000 visual reasoning questions translated from English into 6 Indian languages:
Hindi (hi), Tamil (ta), Telugu (te), Bengali (bn), Kannada (kn), Marathi (mr).
Source benchmarks: MathVista (testmini), ScienceQA, MMMU
Translation: IndicTrans2 (AI4Bharat), verified against GPT-4o/Gemini… See the full description on the dataset page: https://huggingface.co/datasets/Swastikr/multilingual-vlm-reasoning.TransCity-VLM-dataset
TransCity-VLM Dataset
The TransCity-VLM Dataset provides multimodal smart-city data for traffic, energy, mobility, grid operation, urban context understanding, and map-grounded question answering. It supports the training and evaluation of vision-language models for urban prediction, decision support, conversational QA, and reasoning tasks.
The training data are available at this Hugging Face dataset repository.
Dataset Summary
Split
Rows / Files
test JSONL… See the full description on the dataset page: https://huggingface.co/datasets/TransCity-VLM/TransCity-VLM-dataset.myanmar-ocr-dataset-for-vlm
Myanmar OCR Dataset
A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts.
Subsets
Subset
Description
Details
single_font
Rendered with Pyidaungsu font only
437 books
multi_font
Rendered with 76 Myanmar fonts
3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.orpo-vlm-pairs-full
ORPO VLM Preference Pairs (Full)
This dataset contains two versions of vision-language preference pairs for training VLM models using ORPO, DPO, or similar preference-based alignment methods.
Dataset Description
File
Rows
Description
orpo_pairs.jsonl
67,754
Refined/filtered pairs (recommended)
orpo_pairs_all.jsonl
94,346
Full dataset before filtering
Images: 11,982 images
Format: JSONL + PNG images
Language: English
Task: Vision-language… See the full description on the dataset page: https://huggingface.co/datasets/mncai/orpo-vlm-pairs-full.unified-vlm-steering-emu35
Emu3.5 (BAAI): activation steering sweeps
Steered text and image generations from Emu3.5 (BAAI), one of the unified vision-language models in the unified-vlm-steering project. Steering adds alpha * v_hat (the per-layer unit difference-of-means vector) to the residual stream at every layer of a layer config.
34B dense autoregressive model, 64 decoder layers (0-indexed). Images are 32x32 IBQ tokens (512 px), generated on BAAI's patched vLLM engine with request-id-keyed CFG. The… See the full description on the dataset page: https://huggingface.co/datasets/saintsauce/unified-vlm-steering-emu35.
