datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.Sci-VBench-Videos
Sci-VBench Videos
Sci-VBench Videos is the complete set of model outputs behind the Sci-VBench paper: 11,216 videos from 16 text-to-video models, together with the automatic and human scores computed on them. Every video was generated from the verbatim benchmark prompt under the model's default configuration — no rewriting, no prompt expansion — so the released prompts and the released videos correspond exactly.
Prompts and evaluation specifications live in the companion repo… See the full description on the dataset page: https://huggingface.co/datasets/Sci-VBench/Sci-VBench-Videos.video-scissors-sessions
Coding agent session traces for kaofelix/video-scissors-sessions
This dataset contains redacted coding agent session traces collected while working on git@github.com:kaofelix/video-scissors.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/kaofelix/video-scissors-sessions.elv-halluc-videos
ELV-Halluc — videos + annotations
A self-contained mirror of the ELV-Halluc benchmark
(CVPR 2026), bundling the raw .mp4 files together with the annotations so the benchmark can be
run without sourcing videos separately.
Paper: arXiv:2508.21496
Original annotations: HLSv/ELV-Halluc (no videos)
Project page: https://elv-halluc.github.io/
This is an unofficial mirror. All credit for the benchmark goes to the original authors; please
cite their paper (below) rather than this… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/elv-halluc-videos.so101_pp_donuts_v1_reward_videosQualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames
QualityVision Jogging Pose Dataset (61 videos, 14,550 frames) — Sample
This Hugging Face dataset is a compact sample extracted from the full QualityVision Jogging Pose export.
Action label: jogging
Keypoints: 33 landmarks per person (MediaPipe / BlazePose) with x, y, z, visibility
Post-processing (as exported): temporal smoothing + body normalization flags are included in metadata
Use this sample to validate the schema and quality before purchasing larger exports.
Pricing &… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Jogging-Pose-Dataset-61-Videos-14550-Frames.downloaded_videos
