datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VideoMind
🔍VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
Dataset Description
VideoMind is a large-scale video-centric multimodal dataset that can be used to learn powerful and transferable text-video representations
for video understanding tasks such as video question answering and video retrieval. The VideoMind dataset contains 105K(5K test for
only) video samples, each of which is accompanied by audio, as well as systematic… See the full description on the dataset page: https://huggingface.co/datasets/DixinChen/VideoMind.morse-500
MORSE-500 Benchmark
🔥 News
May 15, 2025: We release MORSE-500, 500 programmatically generated videos across six reasoning categories: abstract, mathematical, physical, planning, spatial, and temporal, to stress-test multimodal reasoning. Frontier models including OpenAI o3 and Gemini 2.5 Pro score lower than… See the full description on the dataset page: https://huggingface.co/datasets/video-reasoning/morse-500.
