datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-o3-Video
Open-o3 Video
TL; DR: Open-o3 Video integrates explicit spatio-temporal evidence into video reasoning through curated STGR datasets and a two-stage SFT–RL training strategy, achieving state-of-the-art results on V-STAR and delivering verifiable, reliable reasoning for video understanding.
Data
To provide unified spatio-temporal supervision for grounded video reasoning, we build two datasets: STGR-CoT-30k for supervised fine-tuning and STGR-RL-36k for reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/marinero4972/Open-o3-Video.OpenVideo-Scene-Reasoning
OpenVideo-Scene-Reasoning
OpenVideo-Scene-Reasoning is a video understanding dataset containing 2,894 short video clips, where each sample consists of a 10-second video, five uniformly sampled frames, and a dense scene-level response describing the complete temporal sequence.
Rather than generating captions for individual frames independently, the responses are synthesized by jointly reasoning over the sampled frames to capture temporal progression, object interactions, actions… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenVideo-Scene-Reasoning.
