datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VLM4Bio
Dataset Card for VLM4Bio
Instructions for downloading the dataset
Install Git LFS
Git clone the VLM4Bio repository to download all metadata and associated files
Run the following commands in a terminal:
git clone https://huggingface.co/datasets/imageomics/VLM4Bio
cd VLM4Bio
Downloading and processing bird images
To download the bird images, run the following command:
bash download_bird_images.sh
This should download the bird images inside datasets/Bird/images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/VLM4Bio.WN18RRWN18OmniMed_VLMFB15k
FB15k Dataset
The details of it can be got by this paper titled:
Translating Embeddings for Modeling Multi-relational Data
llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.RefCOCO-VLMEvalKitUMLSSEEDBench2
Dataset Card for Dataset Name
The SEEDBench2 evaluation dataset hosted by VLMEval (authorized by the author).
Dataset Details
Language(s) (NLP): English
License: Apache 2.0
Repository: https://github.com/AILab-CVC/SEED-Bench
Paper [optional]: https://arxiv.org/abs/2311.17092
Citation
@misc{li2023seedbench2,
title={SEED-Bench-2: Benchmarking Multimodal Large Language Models},
author={Bohao Li and Yuying Ge and Yixiao Ge and Guangzhi Wang and Rui… See the full description on the dataset page: https://huggingface.co/datasets/VLMEval/SEEDBench2.FB15k-237KinshipGMAI-MMBenchTallyQA-VLMEvalKitOmni3DBench-VLMEvalKitYAGO3-10vlmevalkit_filesVLMEvalKit_CVQA
CVQA for VLMEvalKit
Original dataset: ported to VLMEvalKit
From the original authors:
CVQA is a culturally diverse multilingual VQA benchmark consisting of over 10,000 questions from 39 country-language pairs. The questions in CVQA are written in both the native languages and English, and are categorized into 10 diverse categories.
{'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=2048x1536 at 0x7C3E0EBEEE00>,
'ID': '5919991144272485961_0',
'Subset':… See the full description on the dataset page: https://huggingface.co/datasets/timothycdc/VLMEvalKit_CVQA.vlwnc-if-vf-universal-class-nanofabricator-v1
Vaelorium Luminex / The Weave NooCathedral InfiLattice / Veyrglass Fabricator "VLWNC-IF-VF" - Universal Class
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific release: v1.0.0 · Hub packaging: hf.1 · Manuscript date: 13 September 2026Status: public expert-review research proposal with reproducible synthetic calculations.
Light-addressed physical compilation for heterogeneous fabrication: a proposed multi-cartridge “light printer in a box” combining… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/vlwnc-if-vf-universal-class-nanofabricator-v1.3DSRBench
3DSRBench Circular Evaluation Package
This upload contains the processed TSV required by PhysBrainEvalKit for 3DSRBench circular evaluation. The TSV embeds the evaluation images as base64 data, so no separate image archive is required.
Files in this repository
3dsrbench_v1_vlmevalkit_circular.tsv: processed evaluation data used by PhysBrainEvalKit.
compute_3drbench_results_circular.py: optional standalone result computation script.
.gitattributes: large-file… See the full description on the dataset page: https://huggingface.co/datasets/VLyb/3DSRBench.VietOnlineNews
VietOnlineNews: Vietnamese Online News Topic Classification Dataset
Dataset Description
VietOnlineNews is a Vietnamese online news dataset constructed for the task of single-label multi-class topic classification. Each sample corresponds to one news article and is assigned exactly one main topic label through the category field.
The dataset was collected from multiple Vietnamese online news sources and processed through a data cleaning pipeline to remove… See the full description on the dataset page: https://huggingface.co/datasets/VLUS06/VietOnlineNews.Mathverse_VLMEvalKitMedical_VLM_SycophancyThis the official data hosting repository for paper "EchoBench: Benchmarking Sycophancy in Medical
Large Vision Language Models".
============open-source_models============
For experiments on open-source models, our implementation is built upon the VLMEvalkit framework.
Navigate to the VLMEval directory
Set up the environment by running: "pip install -e ."
Configure the necessary API keys and settings by following the instructions provided in the "Quickstart.md" file of VLMEvalkit.
To… See the full description on the dataset page: https://huggingface.co/datasets/Botai666/Medical_VLM_Sycophancy.geneturingLiveXiv-VLMEvalKitcrpe_vlmevalkitgpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.rusentimentGSM8K-V-VLMEvalKitvlsp2016This is a copy instance of VLSP2016 dataset. Please obtain the permission and cite the original work when using it.
Original dataset: https://vlsp.org.vn/vlsp2016/eval/sa
VLSP2018-ABSA-Hotel
VLSP2018-ABSA-Hotel
Dataset Summary
The VLSP 2018 Hotel corpus is designed for Vietnamese Aspect-Based Sentiment Analysis (ABSA), covering two sub-tasks of Aspect Category Sentiment Analysis (ACSA):
Aspect Category Detection (ACD): identify which Aspect#Category pairs are present in each review.
Sentiment Polarity Classification (SPC): assign one of three sentiment labels (Positive, Negative, Neutral) to each detected Aspect#Category.
This unified CSV contains 5,600… See the full description on the dataset page: https://huggingface.co/datasets/visolex/VLSP2018-ABSA-Hotel.
