datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVUTBenchmark
Audio-centric Video Understanding Benchmark (AVUT)
This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text Shortcut.
Code Repository: https://github.com/lark-png/AVUT
Paper: https://arxiv.org/pdf/2503.19951
Introduction
The Audio-centric Video Understanding Benchmark (AVUT) aims to evaluate the video comprehension capabilities of multimodal Large Language Models (LLMs), with a particular focus on auditory information. Audio… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/AVUTBenchmark.Allo-AVAAVGen-Bench
AVGen-Bench Generated Videos Data Card
Overview
This data card describes the generated audio-video outputs stored directly in the repository root by model directory.
The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AVGen-Bench.Music-AVQAAV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.avav2.2_train_setAVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.AVSCapBench
AVSCapBench
AVSCapBench contains 1,226 manually annotated omni-modal video clips. Each sample includes a dense caption, visual events, audio events split into speech, music, and sfx, and audio-visual synergistic events.
Download
hf download NJU-LINK/AVSCapBench --repo-type dataset --local-dir AVSCapBench
Structure
videos/
1.mp4
2.mp4
...
metadata.jsonl
metadata/
OmniCaption.json
metadata.jsonl is provided for the Hugging… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/AVSCapBench.AV-Deepfake1M
AV-Deepfake1M
This is the official repository for the paper
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset.
Abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most
advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting
high-quality deepfake images and videos, only a few works address the problem of the localization of small… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M.DREAMS-AVATAR
DREAMS-AVATAR
The DREAMS-Avatar dataset from the DEGAS paper
(3DV 2025), re-registered to pure SMPL-X.
These are the same multiview captures introduced as the DREAMS-Avatar dataset in DEGAS
(Fig. 1b); what is new here is the registration.
32 calibrated, matted camera views of a full-body performance, with one SMPL-X body fitted
to all views at once by our multiview tracker: 300 shape coefficients, 100 expression
coefficients, jaw and both eyes, hands as free 45-dim axis-angle… See the full description on the dataset page: https://huggingface.co/datasets/initialneil/DREAMS-AVATAR.minimax_h3_avatar_500
Watch the full 500-video showcase on YouTube
MiniMax H3 Avatar 500
An image-to-video dataset pairing reference avatar images with detailed generation prompts and generated avatar videos. This release contains 500 curated examples in both a browsable raw layout and a typed Hugging Face dataset.
Version 1.0 · Released August 14, 2026
Dataset contents
Each example contains:
A 1024 × 1024 reference avatar image
A detailed English generation prompt
A generated 640 ×… See the full description on the dataset page: https://huggingface.co/datasets/oakmindai/minimax_h3_avatar_500.alma-avatar-quality-pilot-v1AVE-Dataset
AVE Dataset
The original AVE dataset ported from the GitHub
Notes
One video may contain different audio-visual events, so the total number of videos is not 4143.
annotations.txt: annotations of the AVE dataset. For each sample, you can find its event category, YouTube ID, Quality (all good, meaning that it contains an AVE), start time of an audio-visual event, and the end time of an audio-visual event.
train/val/test-Set.txt: training/validation/testing set used in the… See the full description on the dataset page: https://huggingface.co/datasets/UnFaZeD07/AVE-Dataset.AVE-Compass-v2
AVE-Compass
Project Page | Paper | Evaluation Code
AVE-Compass is a benchmark for evaluating instruction-based audio-video editing abilities. It targets realistic editing requests where visual and audio streams are tightly coupled, such as changing a visible event together with its sound, preserving background ambience during visual edits, or editing speech while maintaining audio-visual consistency.
This dataset release contains the source videos, edit instruction JSON files… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/AVE-Compass-v2.nvidia-av-trajectory-clips-256av_aloha_sim_slot_insertionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 100,
"total_frames": 17741,
"total_tasks": 1,
"total_videos": 600,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iantc104/av_aloha_sim_slot_insertion.svbrd-llm-roadside-video-avAVE-Speech
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
Abstract
AVE Speech is a large-scale Mandarin speech corpus that pairs synchronized audio, lip video and surface electromyography (EMG) recordings. The dataset contains 100 sentences read by 100 native speakers. Each participant repeated the full corpus ten times, yielding over 55 hours of data per modality. These complementary signals enable… See the full description on the dataset page: https://huggingface.co/datasets/MML-Group/AVE-Speech.avhallubench
Dataset Card for AVHalluBench
The dataset is for benchmarking hallucination levels in audio-visual LLMs. It consists of 175 videos and each video has hallucination-free audio and visual descriptions. The statistics are provided in the figure below, and more information can be found in our paper.
Paper: CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
Multimodal Hallucination Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/avhallubench.sign-language-avatar-gloss-dgs
Dataset for German Sign Language Avatar Training
Dataset Summary
This dataset provides curated resources for training data-driven avatars
to perform isolated signs in German Sign Language (Deutsche Gebärdensprache, DGS). It includes videos of individual signs as well as corresponding pose estimation results in a structured and reusable format.
The data is at this moment just sourced from SignDict.org and organized into three primary folders:
videos-raw: Original… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/sign-language-avatar-gloss-dgs.EgoDyn-Bench-ECCV2026
EgoDyn-Bench
A physics-grounded VQA benchmark for evaluating Vision-Language Models on trajectory-based dynamics reasoning in autonomous driving.
Project page | Paper | GitHub
This repository contains the data artifacts for the benchmark. The evaluation harness, baselines, and reference implementations live in the companion GitHub repository.
Note on licensing. The nuScenes-derived portion of this dataset is released under CC BY-NC-SA 4.0 to comply with nuScenes' upstream… See the full description on the dataset page: https://huggingface.co/datasets/TUM-AVS/EgoDyn-Bench-ECCV2026.ave-2
AVE-2: Diagnostic Measurement Layer for Graded Audio-Visual Alignment
Dataset Description
AVE-2 is a diagnostic measurement layer over 570,138 AudioSet-derived 3-second clips. The March 2026 release contains 554,564 train clips and 15,574 eval clips, with zero train/eval overlap at the youtube_id level in the released metadata snapshot.
The field already has many raw video-audio pairs. What it still lacks is a scalable way to say what kind of supervision a clip… See the full description on the dataset page: https://huggingface.co/datasets/ali-vosoughi/ave-2.SO101_AVE_07This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 6854,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zacapa/SO101_AVE_07.nvidia-av-trajectory-clips-384pick_place_avoid_and_not_avoid_calculator_overlay_molmobot
pick_place_avoid_and_not_avoid_calculator_overlay_molmobot
MolmoBot-format dataset for Synthmanip/MolmoBot training.
Layout
dataset_manifest.json
train/valid_trajectory_index.json and val/valid_trajectory_index.json
train/house_*/*.h5 and val/house_*/*.h5
HDF5 video sidecars under the same split/house directories
normalization stats: pick_place_avoid_and_not_avoid_calculator_overlay_molmobot_norm_stats.yaml
Split Summary
split
houses
h5… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/pick_place_avoid_and_not_avoid_calculator_overlay_molmobot.d3il_avoiding_vision_224
D3IL Avoiding Vision 224
This is the original D3IL Avoiding demonstration dataset converted to the
LeRobot v2.1 format and augmented with state-aligned 224 x 224 RGB observations.
It contains all 96 original Avoiding demonstrations (7,305 training frames),
covering the benchmark's 24 avoidance modes with four demonstrations per mode.
The numeric demonstrations are preserved from the original D3IL pickle logs.
Dataset summary
Property
Value
Episodes
96… See the full description on the dataset page: https://huggingface.co/datasets/shivakanthsujit/d3il_avoiding_vision_224.Music-avqa-videoh3-avatar-turboAvaMERGav-skillsAV-Skills
Audio-visual instruction and temporally grounded reasoning data for Nemotron-Labs-Audio-Visual Flamingo
AV-Skills supports joint understanding of video, speech, sound, music, and long-range temporal context in real-world videos.
Dataset Summary
AV-Skills is the audio-visual instruction and reasoning dataset for
Nemotron-Labs-Audio-Visual Flamingo, an open audio-visual language model for long and
complex… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/av-skills.
