datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MSR-VTTClone from "friedrichor/MSR-VTT".
MSRVTT contains 10K video clips and 200K captions.
We adopt the standard 1K-A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text-Video Retrieval field.
Train:
train_7k: 7,010 videos, 140,200 captions
train_9k: 9,000 videos, 180,000 captions
Test:
test_1k: 1,000 videos, 1,000 captions
🌟 Citation
@inproceedings{xu2016msrvtt,
title={Msr-vtt: A large video description dataset… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSR-VTT.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.MVBench
MVBench
Forked from https://huggingface.co/datasets/OpenGVLab/MVBench for reproducibility.
Important Update
[18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference.
We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MVBench.vlm-info-loss-results
VLM Grounding Evaluation Results
Grounding evaluation results for vision-language models on robotics manipulation datasets.
Part of the vlm-info-loss project studying
how VLM connectors transform visual representations.
Background
Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation:
they sharpen dominant-object representations while compressing secondary-object category identity.
All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.Vlaserpixie
Pixie Dataset
This dataset contains data and pre-trained models for the paper Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels.
Project Page: https://pixie-3d.github.io/
Code: https://github.com/vlongle/pixie
Contents
checkpoints_continuous_mse/: Continuous material property prediction model checkpoints
checkpoints_discrete/: Discrete material classification model checkpoints
real_scene_data/: Real scene data for evaluation… See the full description on the dataset page: https://huggingface.co/datasets/vlongle/pixie.reward-projection-goal-generalisation-vlmClinSeek-Bench
ClinSeek-Bench
ClinSeek-Bench is the evaluation suite introduced in
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical
Reasoning. It evaluates clinical reasoning
under two paired settings with the same task definitions and answer labels:
Curated Input: the model answers from the evidence package provided by
the source benchmark.
Automated Evidence-Seeking: the curated context is removed, and the model
must retrieve evidence from raw clinical data using… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Bench.VCR-Bench
VCR-Bench ( A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning)
🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub
Dataset Details
As shown in the figure below, current video benchmarks often lack comprehensive annotations of CoT steps, focusing only on the accuracy of final answers during model evaluation while neglecting the quality of the reasoning process. This evaluation approach makes it difficult to comprehensively evaluate model’s… See the full description on the dataset page: https://huggingface.co/datasets/VLM-Reasoning/VCR-Bench.SID-VLN Datasets of Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale.
VLN-CE-R2R_easi
VLN-CE R2R Dataset for EASI
Vision-and-Language Navigation in Continuous Environments (VLN-CE) Room-to-Room
(R2R) benchmark, repackaged for the EASI
evaluation framework.
Task
An agent receives a natural language navigation instruction and must navigate
through a Matterport3D indoor environment to reach a goal location. The agent
uses discrete actions: STOP, MOVE_FORWARD (0.25m), TURN_LEFT (15 deg),
TURN_RIGHT (15 deg).
Success is measured when the agent stops within 3.0m… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/VLN-CE-R2R_easi.thinkflow-vla-features-b2oda-recon-val-4Bopenvivqa-formating-vlm
OpenViVQA Formatting Dataset for VLM
A Vietnamese multimodal instruction-format dataset for training Vision Language Models (VLMs) on Visual Question Answering (VQA) tasks.
This dataset reformats OpenViVQA-style samples into conversational instruction-tuning format compatible with modern VLM training pipelines such as:
Qwen2-VL
LLaVA
InternVL
Phi-3 Vision
Idefics
SmolVLM
Dataset Structure
Each sample contains:
image: input image
conversations: multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Nhanvi282/openvivqa-formating-vlm.tsiolkovsky-papers
Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus
Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the
Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky
(1857–1935), who derived the rocket equation and described the multistage rocket
decades before anyone could test either.
The archive had been scanned and put online, but without a catalogue you could
query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.robotrace-vla-robustness-traces
RoboTrace Evidence Bundle
This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines.
The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run.
What this bundle is for
Use this repository to inspect evidence from RoboTrace:
action-trace stability metrics
visual… See the full description on the dataset page: https://huggingface.co/datasets/i-am-shaurya05/robotrace-vla-robustness-traces.TBStar-VLM-R2vlabench_primitive_ft_lerobot_video
VLABench Primitive Tasks — LeRobot v3.0 (TsFile)
Apache TsFile version of VLABench/vlabench_primitive_ft_lerobot_video.
Overview
This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially. Compared with the v2.0 and the RLDS versions, this release stores the visual observations in a video-compressed format rather than as individual image files, giving better storage efficiency and data-loading… See the full description on the dataset page: https://huggingface.co/datasets/THULab/vlabench_primitive_ft_lerobot_video.VLN-CE-RxR_easi
VLN-CE RxR Dataset for EASI
Vision-and-Language Navigation in Continuous Environments (VLN-CE)
Room-across-Room (RxR) benchmark, repackaged for the
EASI evaluation framework.
Task
An agent receives a natural language navigation instruction (in English,
Hindi, or Telugu) and must navigate through a Matterport3D indoor environment
to reach a goal location. The agent uses discrete actions: STOP,
MOVE_FORWARD (0.25m), TURN_LEFT (30 deg), TURN_RIGHT (30 deg),
LOOK_UP (30 deg)… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/VLN-CE-RxR_easi.urban-vla-expert-v1
Urban VLA Expert v1
Urban VLA Expert v1 is a simulator dataset for language-conditioned urban driving. Each frame pairs a 256 x 256 front-camera image with ego state, a natural-language instruction, and continuous driving controls.
This is a small research dataset, not evidence that a policy is ready for a real vehicle. The expert is a deterministic simulator controller, and the language prompts are curated paraphrases rather than speech collected from drivers.
What… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/urban-vla-expert-v1.vllm-benchmark-payloads
vLLM Benchmark Payloads
Synthetic OpenAI-format chat-completion payloads for latency / throughput benchmarking
of a model served with vLLM's OpenAI-compatible server.
File
File
Records
Notes
payloads_1k_generic.jsonl
1,000
The benchmark dataset — one request body per line. Each has a ~12k-token system prompt + a dynamic per-record customer context.
sample_payload.json
1
One pretty-printed record, to inspect the format quickly.
All data is 100%… See the full description on the dataset page: https://huggingface.co/datasets/rohitjain28/vllm-benchmark-payloads.opensubs-collocations
OpenSubtitles Collocations
NPMI-scored bigram collocations extracted from the OpenSubtitles parallel corpus. Three languages, three relation types, ~43K bigrams total.
Languages & Corpus Size
Language
Code
Corpus lines
Bigrams
English
en
~100M
15,000
Dutch
nl
~105M
15,000
Serbian
sr
~50M
13,586
Relation Types
ADJ+NOUN — adjective-noun pairs: "slim contract", "kreditan kartica"
VERB+ADP — phrasal verbs / verb-preposition: "come on", "houden… See the full description on the dataset page: https://huggingface.co/datasets/vladvlasov256/opensubs-collocations.vlm_direction_testbedQwen__Qwen2-VL-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-VL-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-VL-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-VL-72B-Instruct-details.Qwen__Qwen2-VL-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2-VL-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-VL-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2-VL-7B-Instruct-details.ViLReward-73KProcess Reward Data for ViLBench: A Suite for Vision-Language Process Reward Modeling
Paper | Project Page
There are 73K vision-language process reward data sourcing from five training sets.
GenderBias-VLdeepscaler-teacher-sft-vllm-official-40k-clean-v2
DeepScaleR Teacher40k Clean v2
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
minimum official reward: 1.0
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe
Counts
raw examples: 40300
kept examples: 21727
train examples: 21292
val… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v2.ogiri-bokete-unsloth-vlm
Japanese Bokete Ogiri — Unsloth VLM format
YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。
各JSONLレコードは「1画像 + 1回答」です。
{
"messages": [
{"role": "user", "content": [
{"type": "image", "image": "images/124469.jpg"},
{"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"}
]},
{"role": "assistant", "content": [
{"type": "text", "text": "..."}
]}
]
}
Files
train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 8192
maximum response chars: 65000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v3-no-reward-filter.
