datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMS-VPR
MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments
Overview
MMS-VPR is the first large-scale multimodal street-level visual place recognition dataset featuring comprehensive integration of images, videos, and rich textual annotations with day–night coverage and a 7-year temporal span in dense pedestrian-only environments.
MMS-VPR comprises 110,529 images and 2,527 video clips… See the full description on the dataset page: https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR.MMSI-Video-Bench_lmmseval
MMSI-Video-Bench
A video-based spatial intelligence benchmark for evaluating Multimodal Large Language Models (MLLMs).
Dataset Description
MMSI-Video-Bench tests models on:
Spatial reasoning
Motion understanding
Planning and prediction
Cross-video reasoning
Dataset Structure
MMSI-Video-Bench/
├── data/
│ └── test-00000-of-00001.parquet # 1106 samples
├── frames.zip # Extracted video frames
├── ref_images.zip #… See the full description on the dataset page: https://huggingface.co/datasets/oscarqjh/MMSI-Video-Bench_lmmseval.MMSDataKhoPhimVTV
