CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gijs /avqa-processedaudio10K<n<100K0 likes5.5k downloads1y agoHugging Face02UnFaZeD07 /Music-AVQAtabular10K<n<100K0 likes3.1k downloads7mo agoHugging Face03juyil /AVQA-videos AVQA — Audio-Visual Question Answering (videos + annotations) A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life audio-visual question answering over short in-the-wild clips. The original release ships only the QA annotations and expects users to collect the source videos from VGGSound themselves. This repository bundles the source video clips together with the official train/val annotations, so the dataset is usable without any YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.tabularvisual-question-answering10K<n<100K1 likes2.8k downloads4mo agoHugging Face04Joysw909 /AVQA Summary | 摘要 This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.audioquestion-answering10K<n<100K2 likes1.8k downloads10mo agoHugging Face05mteb /MUSIC-AVQA_cls-preprocessedaudio1K<n<10K0 likes237 downloads7mo agoHugging Face06inesriahi /valor32k-avqa-v2 Valor32k-AVQA v2.0 Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position. Links Paper: ACM Digital Library Project page: inesriahi.github.io/valor32k-avqa-2 Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.tabularquestion-answering100K<n<1M0 likes201 downloads3mo agoHugging Face07harryhsing /AVQA-R1-6KThis repository contains data presented in EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning. For training and inference, please refer to the Code: https://github.com/HarryHsing/EchoInk Data Format in AVQA-R1-6K: { "problem_id": 0, "problem": "What is the source of the sound in the video?", "data_type": "image_audio", "problem_type": "multiple choice", "options": [ "A. motorcycle", "B. automobile"… See the full description on the dataset page: https://huggingface.co/datasets/harryhsing/AVQA-R1-6K.audio1K<n<10K3 likes79 downloads1y agoHugging Face08umd-zhou-lab /AVQA-Audio-Rubrics AVQA Audio-Reasoning Rubrics Project Page | Paper | Code Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward). Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.textaudio-classification10K<n<100K1 likes67 downloads1mo agoHugging Face09mteb /AVQA_valaudion<1K0 likes46 downloads7mo agoHugging Face10gwkrsrch2 /avqa_hardtext1K<n<10K0 likes20 downloads1y agoHugging Face11syn-omni-sony /avqatabular10K<n<100K0 likes20 downloads8mo agoHugging Face12gwkrsrch2 /music_avqa_hardtext1K<n<10K0 likes15 downloads1y agoHugging Face13gwkrsrch2 /music_avqatext1K<n<10K0 likes14 downloads1y agoHugging Face14gwkrsrch2 /avqa_2025text1K<n<10K0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.