datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.cooking-videos-with-captionsDataset of cooking videos obtained from pexels.com. Captions have been generated using AI.
elv-halluc-videos
ELV-Halluc — videos + annotations
A self-contained mirror of the ELV-Halluc benchmark
(CVPR 2026), bundling the raw .mp4 files together with the annotations so the benchmark can be
run without sourcing videos separately.
Paper: arXiv:2508.21496
Original annotations: HLSv/ELV-Halluc (no videos)
Project page: https://elv-halluc.github.io/
This is an unofficial mirror. All credit for the benchmark goes to the original authors; please
cite their paper (below) rather than this… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/elv-halluc-videos.VideoScienceBench
VideoScienceBench
A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon).
Dataset Summary
Attribute
Value
Examples
160
Domains
Physics, Chemistry
Format
JSONL (prompt + expected phenomenon + vid)
Data Creation Pipeline
Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.Video-STaR
Video-STaR 1M Dataset Card
[🖥️ Website]
[📰 Paper]
[💫 Code]
[🤗 Demo]
🎥 Dataset details
Dataset type:
VSTaR-1M is a 1M instruction tuning dataset, created using Video-STaR, with the source datasets:
Kinetics700
STAR-benchmark
FineDiving
The videos for VSTaR-1M can be found in the links above.
VSTaR-1M is built off of diverse task with the goal of enhancing video-language alignment in Large Video-Language Models (LVLMs).
kinetics700_tune_.json - Instruction tuning… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/Video-STaR.VideoSimpleQA
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
📖 Overview
Video SimpleQA is the first comprehensive benchmark specifically designed for evaluating factual grounding capabilities in Large Video Language Models (LVLMs). Unlike existing video benchmarks that often involve subjective speculation or conflate factual grounding with reasoning skills, Video SimpleQA focuses exclusively on objective factuality evaluation through multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/VideoSimpleQA/VideoSimpleQA.legal-videos-rag
Legal Videos RAG Benchmark
Legal Videos is a benchmark for evaluating RAG pipelines on real-world legal videos pulled from two legal proceedings video datasets.
LocalView, the largest known database of local government public meetings as they are captured and uploaded online covering more than 1000 hours of video.
Seattle City meetings from the Council Data Project (CDP), is the Seattle city subset of the CDP data having meeting videos and multiple metadata covering about 1200… See the full description on the dataset page: https://huggingface.co/datasets/aintropy-ai/legal-videos-rag.
