datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VLM4Bio
Dataset Card for VLM4Bio
Instructions for downloading the dataset
Install Git LFS
Git clone the VLM4Bio repository to download all metadata and associated files
Run the following commands in a terminal:
git clone https://huggingface.co/datasets/imageomics/VLM4Bio
cd VLM4Bio
Downloading and processing bird images
To download the bird images, run the following command:
bash download_bird_images.sh
This should download the bird images inside datasets/Bird/images… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/VLM4Bio.llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.3DSRBench
3DSRBench Circular Evaluation Package
This upload contains the processed TSV required by PhysBrainEvalKit for 3DSRBench circular evaluation. The TSV embeds the evaluation images as base64 data, so no separate image archive is required.
Files in this repository
3dsrbench_v1_vlmevalkit_circular.tsv: processed evaluation data used by PhysBrainEvalKit.
compute_3drbench_results_circular.py: optional standalone result computation script.
.gitattributes: large-file… See the full description on the dataset page: https://huggingface.co/datasets/VLyb/3DSRBench.gpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.tab-vlm
TAB-VLM: Temporal Anachronism Benchmark for Vision-Language Models
Paper: On the Cultural Anachronism and Temporal Reasoning in Vision Language Models (ACL 2026 Findings)
Authors: Mukul Ranjan, Prince Jha, Khushboo Kumari, Zhiqiang Shen
TAB-VLM is a benchmark for measuring cultural anachronism in Vision-Language Models — the tendency to misinterpret historical artifacts using temporally inappropriate concepts, materials, or cultural frameworks. The benchmark consists of 600… See the full description on the dataset page: https://huggingface.co/datasets/mukul54/tab-vlm.Vie-Scenic
