CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tsinghua-ee /AVUTBenchmark Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text Shortcut. Code Repository: https://github.com/lark-png/AVUT Paper: https://arxiv.org/pdf/2503.19951 Introduction The Audio-centric Video Understanding Benchmark (AVUT) aims to evaluate the video comprehension capabilities of multimodal Large Language Models (LLMs), with a particular focus on auditory information. Audio… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/AVUTBenchmark.videovideo-text-to-text1K<n<10K2 likes10k downloads1y agoHugging Face02avalab /Allo-AVAaudion>1T3 likes8.3k downloads2y agoHugging Face03microsoft /AVGen-Bench AVGen-Bench Generated Videos Data Card Overview This data card describes the generated audio-video outputs stored directly in the repository root by model directory. The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AVGen-Bench.imagetext-to-video1K<n<10K6 likes6.4k downloads4mo agoHugging Face04UnFaZeD07 /Music-AVQAtabular10K<n<100K0 likes4.2k downloads7mo agoHugging Face05plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes4k downloads10mo agoHugging Face06manh6054 /avav2.2_train_setvideon<1K1 likes3.6k downloads2y agoHugging Face07juyil /AVQA-videos AVQA — Audio-Visual Question Answering (videos + annotations) A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life audio-visual question answering over short in-the-wild clips. The original release ships only the QA annotations and expects users to collect the source videos from VGGSound themselves. This repository bundles the source video clips together with the official train/val annotations, so the dataset is usable without any YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.tabularvisual-question-answering10K<n<100K1 likes2.9k downloads4mo agoHugging Face08NJU-LINK /AVSCapBench AVSCapBench &nbsp; &nbsp; AVSCapBench contains 1,226 manually annotated omni-modal video clips. Each sample includes a dense caption, visual events, audio events split into speech, music, and sfx, and audio-visual synergistic events. Download hf download NJU-LINK/AVSCapBench --repo-type dataset --local-dir AVSCapBench Structure videos/ 1.mp4 2.mp4 ... metadata.jsonl metadata/ OmniCaption.json metadata.jsonl is provided for the Hugging… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/AVSCapBench.text1K<n<10K1 likes2.2k downloads2mo agoHugging Face09ControlNet /AV-Deepfake1Mgated AV-Deepfake1M This is the official repository for the paper AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset. Abstract The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality deepfake images and videos, only a few works address the problem of the localization of small… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M.videovideo-classification1M<n<10M20 likes2.1k downloads1y agoHugging Face10initialneil /DREAMS-AVATAR DREAMS-AVATAR The DREAMS-Avatar dataset from the DEGAS paper (3DV 2025), re-registered to pure SMPL-X. These are the same multiview captures introduced as the DREAMS-Avatar dataset in DEGAS (Fig. 1b); what is new here is the registration. 32 calibrated, matted camera views of a full-body performance, with one SMPL-X body fitted to all views at once by our multiview tracker: 300 shape coefficients, 100 expression coefficients, jaw and both eyes, hands as free 45-dim axis-angle… See the full description on the dataset page: https://huggingface.co/datasets/initialneil/DREAMS-AVATAR.imageimage-to-3dn<1K0 likes793 downloads2mo agoHugging Face11oakmindai /minimax_h3_avatar_500 Watch the full 500-video showcase on YouTube MiniMax H3 Avatar 500 An image-to-video dataset pairing reference avatar images with detailed generation prompts and generated avatar videos. This release contains 500 curated examples in both a browsable raw layout and a typed Hugging Face dataset. Version 1.0 · Released August 14, 2026 Dataset contents Each example contains: A 1024 × 1024 reference avatar image A detailed English generation prompt A generated 640 ×… See the full description on the dataset page: https://huggingface.co/datasets/oakmindai/minimax_h3_avatar_500.imagen<1K3 likes792 downloads1mo agoHugging Face12igorcouto /alma-avatar-quality-pilot-v1video4 likes750 downloads10d agoHugging Face13UnFaZeD07 /AVE-Dataset AVE Dataset The original AVE dataset ported from the GitHub Notes One video may contain different audio-visual events, so the total number of videos is not 4143. annotations.txt: annotations of the AVE dataset. For each sample, you can find its event category, YouTube ID, Quality (all good, meaning that it contains an AVE), start time of an audio-visual event, and the end time of an audio-visual event. train/val/test-Set.txt: training/validation/testing set used in the… See the full description on the dataset page: https://huggingface.co/datasets/UnFaZeD07/AVE-Dataset.video1K<n<10K0 likes559 downloads9mo agoHugging Face14NJU-LINK /AVE-Compass-v2 AVE-Compass Project Page | Paper | Evaluation Code AVE-Compass is a benchmark for evaluating instruction-based audio-video editing abilities. It targets realistic editing requests where visual and audio streams are tightly coupled, such as changing a visible event together with its sound, preserving background ambience during visual edits, or editing speech while maintaining audio-visual consistency. This dataset release contains the source videos, edit instruction JSON files… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/AVE-Compass-v2.textvideo-to-videon<1K2 likes534 downloads2mo agoHugging Face15nexar-ai /nvidia-av-trajectory-clips-256video1K<n<10K0 likes528 downloads3mo agoHugging Face16iantc104 /av_aloha_sim_slot_insertionThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 100, "total_frames": 17741, "total_tasks": 1, "total_videos": 600, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iantc104/av_aloha_sim_slot_insertion.tabularrobotics10K<n<100K0 likes503 downloads1y agoHugging Face17Ryan-xiangyu-ut /svbrd-llm-roadside-video-avvideon<1K0 likes448 downloads6mo agoHugging Face18MML-Group /AVE-Speechgated AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals Abstract AVE Speech is a large-scale Mandarin speech corpus that pairs synchronized audio, lip video and surface electromyography (EMG) recordings. The dataset contains 100 sentences read by 100 native speakers. Each participant repeated the full corpus ten times, yielding over 55 hours of data per modality. These complementary signals enable… See the full description on the dataset page: https://huggingface.co/datasets/MML-Group/AVE-Speech.audion<1K7 likes425 downloads1y agoHugging Face19typhoon-ai /avhallubench Dataset Card for AVHalluBench The dataset is for benchmarking hallucination levels in audio-visual LLMs. It consists of 175 videos and each video has hallucination-free audio and visual descriptions. The statistics are provided in the figure below, and more information can be found in our paper. Paper: CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models Multimodal Hallucination Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/avhallubench.videon<1K7 likes398 downloads2y agoHugging Face20fhswf /sign-language-avatar-gloss-dgs Dataset for German Sign Language Avatar Training Dataset Summary This dataset provides curated resources for training data-driven avatars to perform isolated signs in German Sign Language (Deutsche Gebärdensprache, DGS). It includes videos of individual signs as well as corresponding pose estimation results in a structured and reusable format. The data is at this moment just sourced from SignDict.org and organized into three primary folders: videos-raw: Original… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/sign-language-avatar-gloss-dgs.video1K<n<10K2 likes360 downloads2mo agoHugging Face21TUM-AVS /EgoDyn-Bench-ECCV2026 EgoDyn-Bench A physics-grounded VQA benchmark for evaluating Vision-Language Models on trajectory-based dynamics reasoning in autonomous driving. Project page | Paper | GitHub This repository contains the data artifacts for the benchmark. The evaluation harness, baselines, and reference implementations live in the companion GitHub repository. Note on licensing. The nuScenes-derived portion of this dataset is released under CC BY-NC-SA 4.0 to comply with nuScenes' upstream… See the full description on the dataset page: https://huggingface.co/datasets/TUM-AVS/EgoDyn-Bench-ECCV2026.textvideo-text-to-text1K<n<10K1 likes330 downloads3mo agoHugging Face22ali-vosoughi /ave-2 AVE-2: Diagnostic Measurement Layer for Graded Audio-Visual Alignment Dataset Description AVE-2 is a diagnostic measurement layer over 570,138 AudioSet-derived 3-second clips. The March 2026 release contains 554,564 train clips and 15,574 eval clips, with zero train/eval overlap at the youtube_id level in the released metadata snapshot. The field already has many raw video-audio pairs. What it still lacks is a scalable way to say what kind of supervision a clip… See the full description on the dataset page: https://huggingface.co/datasets/ali-vosoughi/ave-2.document100K<n<1M4 likes320 downloads6mo agoHugging Face23zacapa /SO101_AVE_07This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 6854, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zacapa/SO101_AVE_07.tabularrobotics100K<n<1M0 likes314 downloads1y agoHugging Face24nexar-ai /nvidia-av-trajectory-clips-384video1K<n<10K0 likes310 downloads3mo agoHugging Face25ccwatson /pick_place_avoid_and_not_avoid_calculator_overlay_molmobot pick_place_avoid_and_not_avoid_calculator_overlay_molmobot MolmoBot-format dataset for Synthmanip/MolmoBot training. Layout dataset_manifest.json train/valid_trajectory_index.json and val/valid_trajectory_index.json train/house_*/*.h5 and val/house_*/*.h5 HDF5 video sidecars under the same split/house directories normalization stats: pick_place_avoid_and_not_avoid_calculator_overlay_molmobot_norm_stats.yaml Split Summary split houses h5… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/pick_place_avoid_and_not_avoid_calculator_overlay_molmobot.videon<1K0 likes284 downloads3mo agoHugging Face26shivakanthsujit /d3il_avoiding_vision_224 D3IL Avoiding Vision 224 This is the original D3IL Avoiding demonstration dataset converted to the LeRobot v2.1 format and augmented with state-aligned 224 x 224 RGB observations. It contains all 96 original Avoiding demonstrations (7,305 training frames), covering the benchmark's 24 avoidance modes with four demonstrations per mode. The numeric demonstrations are preserved from the original D3IL pickle logs. Dataset summary Property Value Episodes 96… See the full description on the dataset page: https://huggingface.co/datasets/shivakanthsujit/d3il_avoiding_vision_224.videoroboticsn<1K0 likes279 downloads2mo agoHugging Face27Night-Quiet /Music-avqa-videovideo1K<n<10K0 likes266 downloads1y agoHugging Face28Kirill7c /h3-avatar-turbotextn<1K0 likes240 downloads23d agoHugging Face29ZhangHanXD /AvaMERGaudio4 likes216 downloads1y agoHugging Face30nvidia /av-skillsAV-Skills Audio-visual instruction and temporally grounded reasoning data for Nemotron-Labs-Audio-Visual Flamingo AV-Skills supports joint understanding of video, speech, sound, music, and long-range temporal context in real-world videos. Dataset Summary AV-Skills is the audio-visual instruction and reasoning dataset for Nemotron-Labs-Audio-Visual Flamingo, an open audio-visual language model for long and complex… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/av-skills.textvideo-text-to-text10K<n<100K9 likes216 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.