datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.video-SALMONN_2_testset
video-SALMONN 2 Benchmark
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.pycharades-stage1_data
VideoSearch-R1 Charades-STA
This repository contains the prepared Charades-STA artifacts used by VideoSearch-R1.
Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Project Page: https://mlvlab.github.io/VideoSearch-R1/
Repository: https://github.com/mlvlab/VideoSearch-R1
Stage 1 Cold Start SFT
Charades_STA_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/charades-stage1_data.videos
CoDaCo - videos dataset
This dataset was created using codaco.app.
Description
All data contributed to this campaign goes to the global CoDaCo datasets.
Labels
This dataset includes the following labels:
Captions
Spoken text
Bounding box texts
Bounding box objects
Contains objects
Performed actions
Tags
Emotions
AI generated
Quality rating
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any… See the full description on the dataset page: https://huggingface.co/datasets/codaco/videos.didemo-stage1_data
VideoSearch-R1 DiDeMo
This repository contains the prepared DiDeMo artifacts used by VideoSearch-R1.
Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Project Page: https://mlvlab.github.io/VideoSearch-R1/
Code: https://github.com/mlvlab/VideoSearch-R1
Stage 1 Cold Start SFT
DiDeMo_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on DiDeMo. Each row keeps the… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/didemo-stage1_data.video-scissors-sessions
Coding agent session traces for kaofelix/video-scissors-sessions
This dataset contains redacted coding agent session traces collected while working on git@github.com:kaofelix/video-scissors.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/kaofelix/video-scissors-sessions.activitynet-stage1_data
VideoSearch-R1 ActivityNet
This repository contains the prepared ActivityNet artifacts used by VideoSearch-R1.
Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Project Page: https://mlvlab.github.io/VideoSearch-R1/
Repository: https://github.com/mlvlab/VideoSearch-R1
Stage 1 Cold Start SFT
ActivityNet_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/activitynet-stage1_data.Sci-VBench-Videos
Sci-VBench Videos
Sci-VBench Videos is the complete set of model outputs behind the Sci-VBench paper: 11,216 videos from 16 text-to-video models, together with the automatic and human scores computed on them. Every video was generated from the verbatim benchmark prompt under the model's default configuration — no rewriting, no prompt expansion — so the released prompts and the released videos correspond exactly.
Prompts and evaluation specifications live in the companion repo… See the full description on the dataset page: https://huggingface.co/datasets/Sci-VBench/Sci-VBench-Videos.smolvlm2-fire-videosazm-archive-20260909-ltx-i2v-videos-1370
ltx_i2v_videos_1370.tar
Backup of an existing dataset archive, preserving its original bytes.
File: ltx_i2v_videos_1370.tar
Size: 2,247,976,960 bytes
SHA256: b46c8369b4d8312ae1bae4c02aca0a8e31cd8d9e4b6a598a9689623f42ca8bbf
Verify the downloaded archive with sha256sum -c SHA256SUMS.
azm-archive-20260909-t2v-needlabel-videos
t2v_data_needlabel_videos.tar
Backup of an existing dataset archive, preserving its original bytes.
File: t2v_data_needlabel_videos.tar
Size: 4,675,983,360 bytes
SHA256: 8672488cef23b55e2a70d3279327a38ddbbd8e2809661600d7c707d59b9e1e18
Verify the downloaded archive with sha256sum -c SHA256SUMS.
VideoSimpleQA
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
📖 Overview
Video SimpleQA is the first comprehensive benchmark specifically designed for evaluating factual grounding capabilities in Large Video Language Models (LVLMs). Unlike existing video benchmarks that often involve subjective speculation or conflate factual grounding with reasoning skills, Video SimpleQA focuses exclusively on objective factuality evaluation through multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/VideoSimpleQA/VideoSimpleQA.VideoScienceBench
VideoScienceBench
A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon).
Dataset Summary
Attribute
Value
Examples
160
Domains
Physics, Chemistry
Format
JSONL (prompt + expected phenomenon + vid)
Data Creation Pipeline
Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.tacwam-videos
tacWAM experiment videos
Generated qualitative evaluation videos for tacWAM. Checkpoints and training data are not stored here.
Layout
V4/
T-Rex/ # T-Rex three-view RGB + fingertip F6 + action experiments
Tujian/ # Tujian glove three-view RGB + Pad30 tactile experiments
The dataset folder is the only directory level below the model version; video files live directly in
that folder. Filenames record whether the model is vanilla or fine-tuned, optimizer… See the full description on the dataset page: https://huggingface.co/datasets/haohw/tacwam-videos.Cricket_Delivery_Videos
🏏 Cricket Delivery Videos Dataset
Metric
Value
Dataset Name
taha1418/Cricket_Delivery_Videos
License
Apache-2.0
Data Type
Video & Text (Video-to-Text)
Tasks
Video Captioning, Visual Question Answering (VQA), Action Recognition
Dataset Description
The Cricket Delivery Videos dataset is a specialized collection of video clips, each featuring a single cricket delivery. It is structured to facilitate Video-to-Text tasks, where a model must analyze the… See the full description on the dataset page: https://huggingface.co/datasets/taha1418/Cricket_Delivery_Videos.islamic_videos_transcribedIslamic videos dataset.
Videos from major islamic theologians like Hamza Yusuf, jonathan Brown or Omar Suleiman.
Transcribed using openai/whisper-large-v3
islamic_videos_transcribedIslamic videos dataset.
Videos from major islamic theologians like Hamza Yusuf, jonathan Brown, Nouman Ali Khan or Omar Suleiman.
Transcribed using openai/whisper-large-v3
elv-halluc-videos
ELV-Halluc — videos + annotations
A self-contained mirror of the ELV-Halluc benchmark
(CVPR 2026), bundling the raw .mp4 files together with the annotations so the benchmark can be
run without sourcing videos separately.
Paper: arXiv:2508.21496
Original annotations: HLSv/ELV-Halluc (no videos)
Project page: https://elv-halluc.github.io/
This is an unofficial mirror. All credit for the benchmark goes to the original authors; please
cite their paper (below) rather than this… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/elv-halluc-videos.legal-videos-rag
Legal Videos RAG Benchmark
Legal Videos is a benchmark for evaluating RAG pipelines on real-world legal videos pulled from two legal proceedings video datasets.
LocalView, the largest known database of local government public meetings as they are captured and uploaded online covering more than 1000 hours of video.
Seattle City meetings from the Council Data Project (CDP), is the Seattle city subset of the CDP data having meeting videos and multiple metadata covering about 1200… See the full description on the dataset page: https://huggingface.co/datasets/aintropy-ai/legal-videos-rag.so101_pp_donuts_v1_reward_videosvideosQualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames
QualityVision Jogging Pose Dataset (61 videos, 14,550 frames) — Sample
This Hugging Face dataset is a compact sample extracted from the full QualityVision Jogging Pose export.
Action label: jogging
Keypoints: 33 landmarks per person (MediaPipe / BlazePose) with x, y, z, visibility
Post-processing (as exported): temporal smoothing + body normalization flags are included in metadata
Use this sample to validate the schema and quality before purchasing larger exports.
Pricing &… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames.hamzayusuf_videos_transcribedTranscribed using openai/whisper-large-v3
videos_for_audiotemporalbench-source-videosAhmedKhan_Videosparenting-videos-dataset
قاعدة بيانات فيديوهات التربية العربية
الوصف
مجموعة بيانات تحتوي عناوين ووصوفات فيديوهات حول تربية الأطفال، مُصممة للاستخدام مع محركات البحث القائمة على الذكاء الاصطناعي.
كيفية الاستخدام
from datasets import load_dataset
dataset = load_dataset("Eng-Bassant/parenting-videos-dataset")
downloaded_videos
