datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.bird3m
Bird3M Dataset
Dataset Description
Bird3M is the first synchronized, multi-modal, multi-individual dataset designed for comprehensive behavioral analysis of freely interacting birds, specifically zebra finches, in naturalistic settings. It addresses the critical need for benchmark datasets that integrate precisely synchronized multi-modal recordings to support tasks such as 3D pose estimation, multi-animal tracking, sound source localization, and vocalization attribution.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission000/bird3m.bird3m-raw-sourcesSEABED
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
SEABED (SouthEast Asian Benchmark for Evaluating Audio Reasoning) covers
six audio-reasoning tasks: speech emotion recognition, speech affective
interpretation, dialect and language identification, dialectal speech
comprehension, prosodic ambiguity resolution, and long-form audio
reasoning. This repository releases a stratified 10% sample (541 of 5,404
records) of its QA data for anonymous peer review, so reviewers… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-1/SEABED.
