zeyun-zhong/RealTimeVideo-Instruct-112K
RealTimeVideo-Instruct-112K A real-time video QA corpus of 112,102 instruction samples. Each question is posed at the moment its answer first becomes visible in the video, so a model must answer from the current scene rather than from the whole clip. It was used, together with offline long-video QA from LLaVA-Video-178K, to train StreamTTT. Annotations only. No video is redistributed. Download each source video set from its original provider (see Video sources).… See the full description on the dataset page: https://huggingface.co/datasets/zeyun-zhong/RealTimeVideo-Instruct-112K.
RealTimeVideo-Instruct-112K
  
A real-time video QA corpus of 112,102 instruction samples. Each question is posed at the moment its answer first becomes visible in the video, so a model must answer from the current scene rather than from the whole clip. It was used, together with offline long-video QA from LLaVA-Video-178K, to train StreamTTT.
Annotations only. No video is redistributed. Download each source video set from its original provider (see Video sources).
Contents
Format
Each file is a JSON list in LLaVA conversation format, one single-turn QA per entry:
{
"id": "Apartment_release_clean_seq131_M1292_0",
"conversations": [
{"from": "human", "value": "<image>\nIn my current view, where is the bar stool?\nA. In the upper-left of my view\nB. In the upper-right of my view\nC. In the lower-left of my view\nD. In the lower-right of my view\nPlease provide your answer by stating the letter followed by the full option."},
{"from": "gpt", "value": "D. In the lower-right of my view"}
],
"rt_time": "63.79",
"data_source": "ADT",
"video": "ADT_Apartment_release_clean_seq131_M1292_preview_rgb.mp4"
}Types are not uniform across files: rt_time and proact_time are strings in most files and numbers in some (cast with float()), and id may be an integer or a string.
Construction
Built as described in Appendix B of the paper:
- proactive queries from Streamo are relocated to their annotated answer time (
t_q := t_a); - additional action, spatial-reasoning, anticipation, and captioning samples are derived directly from ground-truth annotations of EgoTimeQA, Aria Digital Twin, and Ego4D Short-Term Anticipation.
Usage
huggingface-cli download zeyun-zhong/RealTimeVideo-Instruct-112K \
--repo-type dataset --local-dir $DATA_DIRThen place the videos under the layout described in docs/DATA.md. The StreamTTT dataloader reads these files directly via --dataset_use.
Video sources
License
The annotations are released under CC-BY-4.0.
Exception: ADT/qa_ADT.json derives from Aria Digital Twin ground truth and is released under CC-BY-NC-SA-4.0 (non-commercial, share-alike). Videos are not included; each source keeps its own terms, some of which require a signed agreement or prohibit commercial use.
Citation
@article{chen2026streamttt,
title = {StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs},
author = {Chen, Joya and Zhong, Zeyun and Shou, Mike Zheng},
journal = {arXiv preprint arXiv:2608.13416},
year = {2026}
}Please also cite the source datasets whose videos you use.
