datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-asr-leaderboard-resultsAudioMarathon
🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs
Abstract
AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars:
long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audio-htmlwikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.Nemotron-Content-Safety-Audio-Dataset
Nemotron Content Safety Audio Dataset
Dataset Description
The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories.
LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.leaderboard_longformpao-audio-dataset
🎙️ Pa'O Audio Dataset
ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ
📌 Project Summary
The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ).
Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.AudioMarathon
AudioMarathon
AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings. The release package in this directory is organized around 11 benchmark tasks spanning meeting summarization, automatic speech recognition, reading comprehension, authenticity detection, music genre classification, acoustic scene classification, emotion recognition, spoken named entity reasoning, sound event detection, speaker gender… See the full description on the dataset page: https://huggingface.co/datasets/AudioMarathon/AudioMarathon.ambience-audio
Ambience audio dataset
Overview
This dataset was generated by scraping videos from prominent YouTube channels focused on ambient audio. The dataset includes a collection of videos that feature various ambient sounds, such as nature sounds, relaxing music, and environmental noises. For each video, essential metadata was extracted, and a caption was generated using an AI model to enhance the discoverability of the content. This dataset can be useful in various applications… See the full description on the dataset page: https://huggingface.co/datasets/igorriti/ambience-audio.AudioCoT
AudioCoT
AudioCoT is an audio-visual Chain-of-Thought (CoT) correspondent dataset for multimodal large language models in audio generation and editing.
Homepage: ThinkSound Project
Paper: arXiv:2506.21448
GitHub: FunAudioLLM/ThinkSound
Dataset Overview
Each CSV file contains three fields:
id — Unique identifier for the sample
caption — Simple audio description prompt
caption_cot — Chain-of-Thought prompt for audio generation
This dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/liuhuadai/AudioCoT.mcl-mmcl-audiocapsMongolian_audioshuman-perception-audio-deepfake-2026
Human Audio Deepfake Perception 2026
A large-scale listening study evaluating how well humans detect modern audio
deepfakes. The dataset contains 35,532 deepfake-detection judgments from
1,768 anonymous participants across 138 TTS and voice-conversion systems,
collected via a publicly accessible online listening game in 2025–2026.
This is the successor to the 2021 ASVspoof-2019 perception study
(Müller, Pizzi & Williams, 2022)
and extends the same paradigm to modern systems… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/human-perception-audio-deepfake-2026.audiobook-listener-fit-framework
Audiobook Listener-Fit Framework
The Audiobook Listener-Fit Framework is an open structured taxonomy developed by Recommended Audiobooks for describing characteristics that influence the audiobook listening experience.
Traditional ratings mostly describe whether listeners liked a title. This framework is designed to describe how an audiobook listens and which types of listeners may be better suited to it.
Purpose
The framework organizes audiobook characteristics… See the full description on the dataset page: https://huggingface.co/datasets/recommendedaudiobooks/audiobook-listener-fit-framework.9000hours_voa_burmese_audio
Overview
VOA Burmese radio news archive covering Morning (နံနက် ၅:၃၀ – ၆:၃၀) and Evening (ညပိုင်း ၉:၀၀ – ၁၀:၀၀) programmes for every calendar day from 2012-09-16 → 2025-06-09.
Metric
Value
Hours / rows
9 159
Files per day
2 (morning, evening)
Typical file size
15 – 50 MB
Licence
Public-domain (VOA staff recordings, U.S. 17 U.S.C. § 105)
This dataset upgrades Burmese from low-resource to mid-resource status for speech research, enabling self-supervised… See the full description on the dataset page: https://huggingface.co/datasets/freococo/9000hours_voa_burmese_audio.AYDID-audio
AYDID — Audio (gated access)
Segmented 16 kHz mono WAV audio for the AYDID sub-dialectal Yemeni Arabic corpus,
released for non-commercial academic research under a data-use agreement.
Audio is derived from Yemeni broadcast media. To respect source copyright, it is
not publicly redistributed; access is granted to identified researchers who
accept the terms above. Approved users can reproduce the full AYDID pipeline
end-to-end (feature extraction, segmentation, ASR preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/mansoorSaleh/AYDID-audio.youtube-spotify-audio-features
Spotify–YouTube Audio Features
Tabular librosa audio features for tracks aligned with the Spotify / YouTube pipeline in the viral-content-predictor project. Each row is one Spotify track_id matched to a downloaded YouTube audio clip; features are aggregated statistics (mean / std) computed on the decoded waveform.
Files
File
Description
audio_features.csv
One row per track: track_id, 89 derived feature dimensions (means/stds), extraction_success, error_message.… See the full description on the dataset page: https://huggingface.co/datasets/vancenceho/youtube-spotify-audio-features.Arabic-Emotional-Audio-Dataset-Baved
BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging)
A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits.
Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.CoM_Audio_Image_LLM_Generation
This dataset is a Mixture of DIBT/10k_prompts_ranked, lj_speech and Falah/image_generation_prompts_SDXL
Repartition
Why this dataset ?
Training a multimodal router holds crucial significance in the realm of artificial intelligence. By harmonizing different specialized models within a constellation, the router plays a central role in intelligently orchestrating tasks. This approach not only enables precise classification but also paves the way for diverse… See the full description on the dataset page: https://huggingface.co/datasets/Nielzac/CoM_Audio_Image_LLM_Generation.JavisData-Audioaudioset-tf-boxes
TF-SED AudioSet Time-Frequency Boxes
Weak 2-D time-frequency bounding boxes (time and frequency extent) for a
subset of AudioSet. Annotations only —
no audio is included.
Boxes for a clip labelled Speech / Bicycle / Music / Vehicle, drawn over its
spectrogram. Each box localizes an event in both time and frequency.
Using the data
Load the manifest and parse the per-clip boxes:
from datasets import load_dataset
import json
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Dannynis/audioset-tf-boxes.save_audio_punjabi
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cdactvm/save_audio_punjabi.commonvoice_without_audioaudio_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Reihaneh/audio_dataset.test-audio-datasetSpkRecog_VoxCeleb2audio-survey-resultsdarija-hotel-audioAudioEchoes
AudioEchoes
tags: transcription, speech recognition, deep learning
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The 'AudioEchoes' dataset comprises audio files of various echoing environments. Each file has been transcribed by a state-of-the-art speech recognition system, aiming to assist in the development and evaluation of speech-to-text algorithms, especially in the context of recognizing and handling echoes in speech… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/AudioEchoes.
