datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Video-T3-QATextual Temporal Understanding Dataset
Temporal Reasoning Transfer from Text to Video, ICLR 2025
Project Page: https://video-t3.github.io/
In each json file, we provide LLaVA-style text QA samples, using the synthesization method described in our paper.
For example:
[
{
"from": "human",
"value": "Based on the following captions describing keyframes of a video, answer the next question.\n\nCaptions:\nThe image displays a circular emblem with a metallic appearance, conveying a… See the full description on the dataset page: https://huggingface.co/datasets/MMInstruction/Video-T3-QA.CapRL-Video-QA-20K
CapRL-Video-QA-20K.jsonl Video Path Setup
Each value is a relative path under the Hugging Face dataset root of lmms-lab/LLaVA-Video-178K.
Example:
"videos": ["0_30_s_youtube_v0_1/videos/liwei_youtube_videos/videos/youtube_video_2024/ytb_khSwLQOthHQ.mp4"]
Required Video Data
Download the original videos from Hugging Face:
Dataset: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K
Required subdirectories for this 20k subset:
0_30_s_youtube_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-Video-QA-20K.Urban_Dynamics_VideoQA_datasetLong-Video-Tuning-QAs
