CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /audiofolder_two_configs_in_metadataaudion<1K1 likes110k downloads3y agoHugging Face02hf-internal-testing /audiofolder_single_config_in_metadataaudion<1K0 likes101k downloads3y agoHugging Face03agkphysics /AudioSet Dataset Card for AudioSet Dataset Summary AudioSet is a dataset of 10-second clips from YouTube, annotated into one or more sound categories, following the AudioSet ontology. Supported Tasks and Leaderboards audio-classification: Classify audio clips into categories. The leaderboard is available here Languages The class labels in the dataset are in English. Dataset Structure Data Instances Example… See the full description on the dataset page: https://huggingface.co/datasets/agkphysics/AudioSet.audioaudio-classification1M<n<10M109 likes59k downloads11mo agoHugging Face04hf-internal-testing /audiofolder_no_configs_in_metadataaudion<1K0 likes53k downloads3y agoHugging Face05polinaeterna /audiofolder_two_configs_in_metadataaudion<1K0 likes22k downloads3y agoHugging Face06hf-audio /open-asr-leaderboard ESB Test Sets: Parquet & Sorted This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length. The format is also changed, from custom loading script (un-safe remote code) to parquet (safe). Broadly speaking, this dataset was generated with the following code-snippet: from datasets import load_dataset, get_dataset_config_names DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard.audio100K<n<1M80 likes20k downloads3mo agoHugging Face07ProgramComputer /avspeech-visual-audio AVSpeech Video + Audio This repository is a media-bearing reconstruction of the public AVSpeech annotations. Each row represents an already-trimmed segment and keeps the original source-video timing and target-face-center metadata. Dataset structure clip_id: identifier derived as {youtube_id}_{start_sec:.3f}_{end_sec:.3f}. avspeech_metadata: JSON containing youtube_id, start_sec, end_sec, x_center, and y_center from the AVSpeech annotation. video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.audio1M<n<10M5 likes18k downloads1mo agoHugging Face08laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face09ArtificialAnalysis /big_bench_audio Artificial Analysis Big Bench Audio Dataset Summary Big Bench Audio is an audio version of a subset of Big Bench Hard questions. The dataset can be used for evaluating the reasoning capabilities of models that support audio input. The dataset includes 1000 audio recordings for all questions from the following Big Bench Hard categories. Descriptions are taken from Suzgun et al. (2022): Formal Fallacies Syllogisms Negation (Formal Fallacies) - 250 questions Given a context… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio.audioaudio-to-audio1K<n<10K38 likes15k downloads2y agoHugging Face10EarthSpeciesProject /NatureLM-audio-training Dataset card for NatureLM-audio-training Overview NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.audioaudio-classification10M<n<100M18 likes13k downloads1y agoHugging Face11hf-internal-testing /dummy-audio-samplesaudion<1K0 likes13k downloads2d agoHugging Face12hf-internal-testing /audiofolder_two_configs_in_metadata_with_defaultaudion<1K0 likes12k downloads3y agoHugging Face13ilanashapiro /stg-paired-audioaudio0 likes11k downloads1y agoHugging Face14moonshine-ai /audio_samples_1kaudio0 likes9.1k downloads6mo agoHugging Face15Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes7.7k downloads3mo agoHugging Face16humair025 /suno-audio 🎵 Suno Audio Dataset A comprehensive dataset of 49,698 AI-generated music tracks from Suno, organized in 50 batches of 1000 samples each. 🎧 All audio files are playable directly in the dataset viewer! Dataset Structure The dataset is organized into batches (batch_0, batch_1, etc.), each containing up to 1000 audio samples with metadata. Fields audio: 🎵 Playable MP3 audio file (click to play in viewer!) id: Unique track identifier title: Song title… See the full description on the dataset page: https://huggingface.co/datasets/humair025/suno-audio.audiotext-to-audio10K<n<100K2 likes7.1k downloads8mo agoHugging Face17liumindmind /Neko_Audio-80K_Short audio10K<n<100K30 likes6.8k downloads4mo agoHugging Face18Reza2kn /telegram-audiobook-chizzled Telegram Persian Audiobook Chizzled 1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.audion<1K4 likes6.7k downloads11d agoHugging Face19AudioLLMs /Multitask-National-Speech-Corpus-v1-extendaudio10M<n<100M5 likes6.4k downloads1y agoHugging Face20takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.1k downloads6mo agoHugging Face21nader39 /audio-filesaudion<1K0 likes5.9k downloads17h agoHugging Face22RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes5k downloads3mo agoHugging Face23trentmkelly /police-scanner-audio Police Scanner Audio Dataset A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds. Dataset Overview This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications. Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.audioautomatic-speech-recognition1K<n<10K1 likes4.5k downloads1y agoHugging Face24zaibihassan /Quranic-Translation-Audio-Data Overview Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories. Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.audioaudio-to-audio1K<n<10K1 likes4.5k downloads3mo agoHugging Face25laion /laion-audio-previewaudio1M<n<10M11 likes4.5k downloads2y agoHugging Face26mitermix /audiosnippetsaudio1M<n<10M6 likes4.4k downloads2y agoHugging Face27manavtabbly /hindi_audio_dataset_testaudion<1K0 likes4.3k downloads11mo agoHugging Face28zihan-audio /librispeech_asr Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/librispeech_asr.audioautomatic-speech-recognition100K<n<1M0 likes4.3k downloads26d agoHugging Face29zaibihassan /Quranic-Word-By-Word-Audio-Data 🌟 Overview Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines: Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening. Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow. Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.audioautomatic-speech-recognition100K<n<1M4 likes4.2k downloads4mo agoHugging Face30OpenSound /AudioCapsaudio10K<n<100K26 likes4k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.