datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-modal-Self-instruct
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Evaluation
Citation
You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip.
Dataset Description
Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.modal-semantics-reasoning
Modal Semantics Reasoning
Can a language model change its answer when the rules of modal logic change?
Each example contains the same premises and conclusion under two semantic
specifications. Only one rule about possible worlds or objects changes, and
the correct answer changes with it. Automated theorem provers verify every
label.
This dataset accompanies Same Formulas, Different Semantics: Do Language
Models Follow Modal Logic Specifications?
Dataset subsets… See the full description on the dataset page: https://huggingface.co/datasets/sileod/modal-semantics-reasoning.full-modality-data
Full Modality Dataset Statistics
Video Statistics
Total Videos: 28,472
Total Duration: 1422.33 hours
Average Duration: 179.84 seconds
Median Duration: 160.08 seconds
Duration Range: 10.04s - 1780.03s
QA Statistics
Total Questions: 1,444,526
Average Questions per Video: 50.7
Questions per Video Range: 14 - 450
Question Type Distribution
OE: 1,444,526 (100.0%)
Question Category Distribution
temporal: 96,873 (6.7%)
causal: 96,873… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/full-modality-data.modality-conflict-arbitration-v2
Modality-Conflict Arbitration Benchmark (v2)
A controlled benchmark for studying how a vision-language model arbitrates between
its two input channels when they disagree — and whether that choice tracks the
reliability of each channel.
Each row is a single conflict trial: an image of one math problem paired with the
text of a different problem. Because the two ground-truth answers are carried side by
side, the model's output alone tells you which modality it followed — no… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/modality-conflict-arbitration-v2.
