datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.ccnewsThis dataset is the result of processing all WARC files in the CCNews Corpus, from the beginning (2016) to June of 2024.
The data has been cleaned and deduplicated, and language of articles have been detected and added. The process is similar to what HuggingFace's DataTrove does.
Overall, it contains about 600 million news articles in more than 100 languages from all around the globe.
For license information, please refer to CommonCrawl's Terms of Use.
Sample Python code to explore this… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/ccnews.image2struct-latex-v1
Image2Struct - Latex
Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo
License: Apache License Version 2.0, January 2004
Dataset description
Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images.
This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt:
Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.FUSION-Finetune-12M
FUSION-12M Dataset
Please see paper & website for more information:
https://arxiv.org/abs/2504.09925
https://github.com/starriver030515/FUSION
Overview
FUSION-12M is a large-scale, diverse multimodal instruction-tuning dataset used to train FUSION-3B and FUSION-8B models. It builds upon Cambrian-1 by significantly expanding both the quantity and variety of data, particularly in areas such as OCR, mathematical reasoning, and synthetic high-quality Q&A data. The goal is… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Finetune-12M.FUSION-Pretrain-10M
FUSION-10M Dataset
Please see paper & website for more information:
https://arxiv.org/abs/2504.09925
https://github.com/starriver030515/FUSION
Overview
FUSION-10M is a large-scale, high-quality dataset of image-caption pairs used to pretrain FUSION-3B and FUSION-8B models. It builds upon established datasets such as LLaVA, ShareGPT4, and PixelProse. In addition, we synthesize 2 million task-specific image-caption pairs to further enrich the dataset. The goal of… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Pretrain-10M.strata-insurance-corpus
Strata Insurance Corpus
A reproducible, fully synthetic, multi-format insurance document corpus for a fictional pan-European
property-&-casualty insurer, Meridian Mutual, shipped with a golden evaluation set produced by
construction. Built to exercise and benchmark document-RAG systems on enterprise-shaped data —
born-digital and scanned PDFs, Word documents, spreadsheets, and photos — with trustworthy ground truth.
Everything here is synthetic. No real persons, companies, or… See the full description on the dataset page: https://huggingface.co/datasets/NikolaiSachok/strata-insurance-corpus.STRIDE-QA-Dataset-Mini
STRIDE-QA-Dataset-Mini
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
⚠️ Note: STRIDE-QA-Dataset-Mini is provided as a preliminary version and does not fully match the format of the… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset-Mini.StreamingBench-Slice
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.stackoverflowVQA-filteredStreamGaze_v2
StreamGaze Dataset
StreamGaze is a comprehensive streaming video benchmark for evaluating MLLMs on gaze-based QA tasks across past, present, and future contexts.
Companion dataset: The EgoGazeVQA dataset is hosted separately at Peanuttoad/gaze_dataset.
📁 Dataset Structure
streamgaze/
├── metadata/
│ ├── egtea.csv # EGTEA fixation metadata
│ ├── egoexolearn.csv # EgoExoLearn fixation metadata
│ └── holoassist.csv # HoloAssist… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/StreamGaze_v2.StreetVision-10K
StreetVision-10K
Each sample contains:
A system prompt instructing the model to act as an OSINT/geospatial expert
A user message with a street-level photo and the instruction to determine coordinates
An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format
Format
Each line is a JSON array of ChatML messages:
[
{"role": "system", "content": "..."},
{"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.Multimodal-STEM-HLE-plus-plus
multimodal-STEM-HLE++
A high-value multimodal STEM dataset designed and empirically proven to push state-of-the-art LLMs beyond their current limits.
Explore the full multimodal-STEM-HLE++ dataset: https://go.turing.com/mm-stem-hle
Why This Dataset
Post-training with RL is now the primary driver of frontier model improvement. The bottleneck is finding data at the right difficulty for current SOTA models. MMLU is saturated (>90%). HLE, once considered unsolvable, is now… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Multimodal-STEM-HLE-plus-plus.StyleCruxGen
Dataset Card for StyleCruxGen
A visual dataset exploring the interplay between style, object, and environment using diffusion-generated imagery.
Dataset Details
Dataset Description
StyleCruxGen is a synthetic multi-style image dataset generated using the Stable Diffusion XL (SDXL) model.
It contains 15465 high-resolution images (1024px) featuring 30 distinct object-environment pairs rendered across photorealistic and 10 artistic styles.
The 30… See the full description on the dataset page: https://huggingface.co/datasets/bodhisattamaiti/StyleCruxGen.medix-medical-sft
MediX medical SFT project dataset
Prepared training data for the Medical Agent Assistant project, containing VQA, medical knowledge, clinical-vignette QA, and research-context QA. This is a project-specific mixture and split, not an official benchmark split or a clinician-validated dataset.
Split
VQA
Knowledge
Case
Context
Total
Train
1726
1193
1734
765
5418
Validation
93
150
169
91
503
Test
105
150
162
97
514
Sources and licenses
VQA:… See the full description on the dataset page: https://huggingface.co/datasets/starttoshow/medix-medical-sft.stackoverflowVQA-filtered-small
Dataset Card for "stackoverflowVQA-filtered-small"
More Information needed
emma_stone
EMMA Clone Dataset (Small Version)
EMMA Stone is a reduced version of the EMMA (Enhanced MultiModal reAsoning) benchmark with 8 samples per subject category, designed for quick testing and development.
This dataset contains:
Chemistry: 8 samples
Coding: 8 samples
Math: 8 samples
Physics: 8 samples
All: 32 samples (8 from each category)
Usage
Loading with datasets library
from datasets import load_dataset
# Load specific subject
chemistry_data =… See the full description on the dataset page: https://huggingface.co/datasets/winvswon78/emma_stone.VQA_MedStylExNet5k
Dataset Card for StylExNet5k
StylExNet5k is a multi-style synthetic image dataset consisting of 5,000 images across 100 everyday object categories, each rendered in 10 distinct artistic or representational styles and placed in varied real-world contextual environments.
It is intended to support evaluation and training of computer vision and vision–language models across style and context domains.
Dataset Details
Dataset Description
StylExNet5k is a synthetic… See the full description on the dataset page: https://huggingface.co/datasets/bodhisattamaiti/StylExNet5k.senufo-staffs-bernard-de-grunne-tefaf-2014
senufo-staffs-bernard-de-grunne-tefaf-2014
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
45
Avg chars/chunk
756
Avg images/chunk
1.49
Source files
1
Duplicates removed
0
Quality filtered
0
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean
string… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/senufo-staffs-bernard-de-grunne-tefaf-2014.
