datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MSR-VTTClone from "friedrichor/MSR-VTT".
MSRVTT contains 10K video clips and 200K captions.
We adopt the standard 1K-A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text-Video Retrieval field.
Train:
train_7k: 7,010 videos, 140,200 captions
train_9k: 9,000 videos, 180,000 captions
Test:
test_1k: 1,000 videos, 1,000 captions
🌟 Citation
@inproceedings{xu2016msrvtt,
title={Msr-vtt: A large video description dataset… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSR-VTT.MMLongBench-docNExTQAVista
Dataset Card for "Vista"
"700.000 Vietnamese vision-language samples open-source dataset"
Dataset Overview
This dataset contains over 700,000 Vietnamese vision-language samples, created by Gemini Pro. We employed several prompt engineering techniques: few-shot learning, caption-based prompting and image-based prompting.
For the COCO dataset, we generated data using Llava-style prompts
For the ShareGPT4V dataset, we used translation prompts.
Caption-based prompting:… See the full description on the dataset page: https://huggingface.co/datasets/Vi-VLM/Vista.ViDoSeek-page-fixedtest-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.MMLongBench-page-fixedViDoSeekMVBench
MVBench
Forked from https://huggingface.co/datasets/OpenGVLab/MVBench for reproducibility.
Important Update
[18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference.
We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MVBench.vlm-info-loss-results
VLM Grounding Evaluation Results
Grounding evaluation results for vision-language models on robotics manipulation datasets.
Part of the vlm-info-loss project studying
how VLM connectors transform visual representations.
Background
Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation:
they sharpen dominant-object representations while compressing secondary-object category identity.
All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.reward-projection-goal-generalisation-vlmVCR-Bench
VCR-Bench ( A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning)
🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub
Dataset Details
As shown in the figure below, current video benchmarks often lack comprehensive annotations of CoT steps, focusing only on the accuracy of final answers during model evaluation while neglecting the quality of the reasoning process. This evaluation approach makes it difficult to comprehensively evaluate model’s… See the full description on the dataset page: https://huggingface.co/datasets/VLM-Reasoning/VCR-Bench.libero_90_lerobot_pathmask_vlm_labeledThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 3959,
"total_frames": 574571,
"total_tasks": 73,
"total_videos": 15836,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3959"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jesbu1/libero_90_lerobot_pathmask_vlm_labeled.cataract_surgery_vlm_eval
Cataract Surgery VLM Evaluation Dataset — Test
Self-contained, flat evaluation set for cataract-surgery video understanding.
No external dependencies — every record and its video live in this single folder.
At a glance
Item
Count
Notes
JSONL records
1004
one file = one record = one valid JSON line
Videos (.mp4)
518
flat, prefixed names; ~3.45 GB total
Split
Test only
held-out evaluation, no training overlap
Folder layout
flat
no subfolders; every… See the full description on the dataset page: https://huggingface.co/datasets/shahedm2001/cataract_surgery_vlm_eval.oda-recon-val-4Bdetails_VLM-Reasoner__Qwen2.5-VL-3B-Instruct-se-v1-80step
Dataset Card for Evaluation run of VLM-Reasoner/Qwen2.5-VL-3B-Instruct-se-v1-80step
Dataset automatically created during the evaluation run of model VLM-Reasoner/Qwen2.5-VL-3B-Instruct-se-v1-80step.
The dataset is composed of 4 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/VLM-Reasoner/details_VLM-Reasoner__Qwen2.5-VL-3B-Instruct-se-v1-80step.openvivqa-formating-vlm
OpenViVQA Formatting Dataset for VLM
A Vietnamese multimodal instruction-format dataset for training Vision Language Models (VLMs) on Visual Question Answering (VQA) tasks.
This dataset reformats OpenViVQA-style samples into conversational instruction-tuning format compatible with modern VLM training pipelines such as:
Qwen2-VL
LLaVA
InternVL
Phi-3 Vision
Idefics
SmolVLM
Dataset Structure
Each sample contains:
image: input image
conversations: multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Nhanvi282/openvivqa-formating-vlm.vlm-fix
vlmfix HF Upload
This folder contains local parquet exports prepared for Hugging Face upload.
Included configs:
vlm_fix: main, images
vlm_fix_text_only: main
vlm_fix_posttrain_d1: train, test
vlm_fix_posttrain_d2: train, test
vlm_fix_posttrain_d3: train, test
synth_legs_train: train
ViewSpatial-Bench-vlmevalMathverse_VLMEvalKitogiri-bokete-unsloth-vlm
Japanese Bokete Ogiri — Unsloth VLM format
YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。
各JSONLレコードは「1画像 + 1回答」です。
{
"messages": [
{"role": "user", "content": [
{"type": "image", "image": "images/124469.jpg"},
{"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"}
]},
{"role": "assistant", "content": [
{"type": "text", "text": "..."}
]}
]
}
Files
train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.details_Qwen__Qwen2.5-VL-3B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-VL-3B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-VL-3B-Instruct.
The dataset is composed of 4 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/VLM-Reasoner/details_Qwen__Qwen2.5-VL-3B-Instruct.dexjoco-vlm-labels-Astra
DexJoCo VLM labels — Astra v3
Six single-arm DexJoCo tasks, 600 episodes, 13,352 labeled stride-16 anchors
over 218,993 original frames. Model: gemini/gemini-3.8-flash via an
OpenAI-compatible LiteLLM endpoint. These are annotation sidecars, not robot
images, action trajectories, contact ground truth, or compression-speed labels.
Files
File
Purpose
labels/anchors.parquet
Actual VLM labels, grades A–D, conf, instruction, measured facts, camera names and… See the full description on the dataset page: https://huggingface.co/datasets/prehj/dexjoco-vlm-labels-Astra.VLM-GeoPrivacyBench
VLM-GeoPrivacy
📊 Repository | 📑 Paper (Accepted to ICLR 2026)
Dataset description
We introduce VLM-GeoPrivacy, the first benchmark that challenges VLMs to reason about image context and sharing intent to choose the contextually-appropriate level of location disclosure.
Our dataset consists of 1,200 real-world images richly annotated with context, sharing intent, and expected granularity, curated from general geolocation datasets including YFCC4k, YFCC26k… See the full description on the dataset page: https://huggingface.co/datasets/RayY/VLM-GeoPrivacyBench.GSM8K-V-VLMEvalKitpiper_vlmbase_w100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "piper",
"total_episodes": 79,
"total_frames": 14782,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:79"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jwwoo/piper_vlmbase_w100.TBStar-VLM-R2MMLongBenchdetails_._ckpt_Qwen2.5-VL-Instruct-3B-se-v1-460step
Dataset Card for Evaluation run of ._ckpt_Qwen2.5-VL-Instruct-3B-se-v1-460step
Dataset automatically created during the evaluation run of model ._ckpt_Qwen2.5-VL-Instruct-3B-se-v1-460step.
The dataset is composed of 4 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/VLM-Reasoner/details_._ckpt_Qwen2.5-VL-Instruct-3B-se-v1-460step.
