zeyun-zhong/RealTimeVideo-Instruct-112K
RealTimeVideo-Instruct-112K A real-time video QA corpus of 112,102 instruction samples. Each question is posed at the moment its answer first becomes visible in the video, so a model must answer from the current scene rather than from the whole clip. It was used, together with offline long-video QA from LLaVA-Video-178K, to train StreamTTT. Annotations only. No video is redistributed. Download each source video set from its original provider (see Video sources).… See the full description on the dataset page: https://huggingface.co/datasets/zeyun-zhong/RealTimeVideo-Instruct-112K.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face