video-qa
Datasets
All datasets matching “video-qa”TaskMeAnything-v1-videoqa-2024
Dataset Card for TaskMeAnything-v1-videoqa-2024
TaskMeAnything-v1-videoqa-2024 benchmark dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-2024-Videoqa
TaskMeAnything-v1-videoqa-2024 is a benchmark for reflecting the current progress of MLMs by automatically finding tasks that SOTA MLMs struggle with using the TaskMeAnything Top-K queries.
This… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-videoqa-2024.TaskMeAnything-v1-videoqa-random
Dataset Card for TaskMeAnything-v1-videoqa-random
TaskMeAnything-v1-videoqa-random dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-Random
TaskMeAnything-v1-videoqa-random is a dataset which randomly sampled questions from TaskMeAnything-v1, including 2,700 VideoQA questions. The dataset contains 9 splits, while each splits contains 300 questions… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-videoqa-random.Video-T3-QATextual Temporal Understanding Dataset
Temporal Reasoning Transfer from Text to Video, ICLR 2025
Project Page: https://video-t3.github.io/
In each json file, we provide LLaVA-style text QA samples, using the synthesization method described in our paper.
For example:
[
{
"from": "human",
"value": "Based on the following captions describing keyframes of a video, answer the next question.\n\nCaptions:\nThe image displays a circular emblem with a metallic appearance, conveying a… See the full description on the dataset page: https://huggingface.co/datasets/MMInstruction/Video-T3-QA.Adaption-video-qa-diverse-topics
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
video_qa_diverse_topics
This dataset contains question-answer pairs derived from a diverse collection of video clips covering topics such as biology, history, sports, and astronomy. Each entry includes a specific question about the video content and a corresponding factual answer, alongside metadata like duration, tags, and object lists. The samples demonstrate a focus on… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-video-qa-diverse-topics.CapRL-Video-QA-20K
CapRL-Video-QA-20K.jsonl Video Path Setup
Each value is a relative path under the Hugging Face dataset root of lmms-lab/LLaVA-Video-178K.
Example:
"videos": ["0_30_s_youtube_v0_1/videos/liwei_youtube_videos/videos/youtube_video_2024/ytb_khSwLQOthHQ.mp4"]
Required Video Data
Download the original videos from Hugging Face:
Dataset: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K
Required subdirectories for this 20k subset:
0_30_s_youtube_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/internlm/CapRL-Video-QA-20K.nlp-qa-audio-video
NLP QA Audio Video Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for NLP QA work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/leond-u0114/nlp-qa-audio-video.
