zeyun-zhong/RealTimeVideo-Instruct-112K
RealTimeVideo-Instruct-112K A real-time video QA corpus of 112,102 instruction samples. Each question is posed at the moment its answer first becomes visible in the video, so a model must answer from the current scene rather than from the whole clip. It was used, together with offline long-video QA from LLaVA-Video-178K, to train StreamTTT. Annotations only. No video is redistributed. Download each source video set from its original provider (see Video sources).… See the full description on the dataset page: https://huggingface.co/datasets/zeyun-zhong/RealTimeVideo-Instruct-112K.
This repository belongs to zeyun-zhong on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
