datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AudioSet
Dataset Card for AudioSet
Dataset Summary
AudioSet is a dataset of 10-second clips from YouTube, annotated into one or more sound categories, following the AudioSet ontology.
Supported Tasks and Leaderboards
audio-classification: Classify audio clips into categories. The leaderboard is available here
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
Example… See the full description on the dataset page: https://huggingface.co/datasets/agkphysics/AudioSet.open-asr-leaderboard
ESB Test Sets: Parquet & Sorted
This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length.
The format is also changed, from custom loading script (un-safe remote code) to parquet (safe).
Broadly speaking, this dataset was generated with the following code-snippet:
from datasets import load_dataset, get_dataset_config_names
DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from
HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard.avspeech-visual-audio
AVSpeech Video + Audio
This repository is a media-bearing reconstruction of the public AVSpeech
annotations. Each row represents an already-trimmed segment and keeps the
original source-video timing and target-face-center metadata.
Dataset structure
clip_id: identifier derived as
{youtube_id}_{start_sec:.3f}_{end_sec:.3f}.
avspeech_metadata: JSON containing youtube_id, start_sec, end_sec,
x_center, and y_center from the AVSpeech annotation.
video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.LAION-Audio-300Mbig_bench_audio
Artificial Analysis Big Bench Audio
Dataset Summary
Big Bench Audio is an audio version of a subset of Big Bench Hard questions. The dataset can be used for evaluating the reasoning capabilities of models that support audio input.
The dataset includes 1000 audio recordings for all questions from the following Big Bench Hard categories. Descriptions are taken from Suzgun et al. (2022):
Formal Fallacies Syllogisms Negation (Formal Fallacies) - 250 questions
Given a context… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio.NatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.stg-paired-audioaudioset
AudioSet
AudioSet[1] consists of an expanding ontology of 527 audio event classes and a collection of 2M human-labelled 10-second sound clips drawn from YouTube.
Some clips are missing on YouTube, so the number of files downloaded is different from time to time.
This repository contains 20550 / 22160 of the balanced train set, 1913637 / 2041789 of the unbalanced train set (separated into 41 parts), and 18887 / 20371 of the evaluation set.
The pre-process script can be found at… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/audioset.audio_samples_1kAudiobookRu
AudiobookRu
Russian-language audiobook audio (chapter-level), scraped from the web. Each row is one chapter: embedded MP3 bytes + book/chapter metadata.
psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.suno-audio
🎵 Suno Audio Dataset
A comprehensive dataset of 49,698 AI-generated music tracks from Suno, organized in 50 batches of 1000 samples each.
🎧 All audio files are playable directly in the dataset viewer!
Dataset Structure
The dataset is organized into batches (batch_0, batch_1, etc.), each containing up to 1000 audio samples with metadata.
Fields
audio: 🎵 Playable MP3 audio file (click to play in viewer!)
id: Unique track identifier
title: Song title… See the full description on the dataset page: https://huggingface.co/datasets/humair025/suno-audio.Neko_Audio-80K_Short
telegram-audiobook-chizzled
Telegram Persian Audiobook Chizzled
1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release
This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.Multitask-National-Speech-Corpus-v1-extendmultilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.open-asr-leaderboard-resultspolice-scanner-audio
Police Scanner Audio Dataset
A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds.
Dataset Overview
This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications.
Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.laion-audio-previewaudiosnippetshindi_audio_dataset_testlibrispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/librispeech_asr.AudioCapsAudioMarathon
🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs
Abstract
AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars:
long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.Audio-Reasoner-CoTAmultilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.Neko_Audio-30K_Long
