CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /audiofolder_two_configs_in_metadataaudion<1K1 likes101k downloads3y agoHugging Face02hf-internal-testing /audiofolder_single_config_in_metadataaudion<1K0 likes93k downloads3y agoHugging Face03agkphysics /AudioSet Dataset Card for AudioSet Dataset Summary AudioSet is a dataset of 10-second clips from YouTube, annotated into one or more sound categories, following the AudioSet ontology. Supported Tasks and Leaderboards audio-classification: Classify audio clips into categories. The leaderboard is available here Languages The class labels in the dataset are in English. Dataset Structure Data Instances Example… See the full description on the dataset page: https://huggingface.co/datasets/agkphysics/AudioSet.audioaudio-classification1M<n<10M109 likes58k downloads11mo agoHugging Face04hf-internal-testing /audiofolder_no_configs_in_metadataaudion<1K0 likes49k downloads3y agoHugging Face05mewmeters /audio-data0 likes33k downloads5mo agoHugging Face06polinaeterna /audiofolder_two_configs_in_metadataaudion<1K0 likes20k downloads3y agoHugging Face07hf-audio /open-asr-leaderboard ESB Test Sets: Parquet & Sorted This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length. The format is also changed, from custom loading script (un-safe remote code) to parquet (safe). Broadly speaking, this dataset was generated with the following code-snippet: from datasets import load_dataset, get_dataset_config_names DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard.audio100K<n<1M80 likes19k downloads3mo agoHugging Face08EarthSpeciesProject /NatureLM-audio-training Dataset card for NatureLM-audio-training Overview NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.audioaudio-classification10M<n<100M18 likes19k downloads1y agoHugging Face09ProgramComputer /avspeech-visual-audio AVSpeech Video + Audio This repository is a media-bearing reconstruction of the public AVSpeech annotations. Each row represents an already-trimmed segment and keeps the original source-video timing and target-face-center metadata. Dataset structure clip_id: identifier derived as {youtube_id}_{start_sec:.3f}_{end_sec:.3f}. avspeech_metadata: JSON containing youtube_id, start_sec, end_sec, x_center, and y_center from the AVSpeech annotation. video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.audio1M<n<10M5 likes18k downloads1mo agoHugging Face10laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face11ArtificialAnalysis /big_bench_audio Artificial Analysis Big Bench Audio Dataset Summary Big Bench Audio is an audio version of a subset of Big Bench Hard questions. The dataset can be used for evaluating the reasoning capabilities of models that support audio input. The dataset includes 1000 audio recordings for all questions from the following Big Bench Hard categories. Descriptions are taken from Suzgun et al. (2022): Formal Fallacies Syllogisms Negation (Formal Fallacies) - 250 questions Given a context… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio.audioaudio-to-audio1K<n<10K38 likes15k downloads2y agoHugging Face12aoxo /audios20 likes13k downloads0m agoHugging Face13hf-internal-testing /dummy-audio-samplesaudion<1K0 likes12k downloads13h agoHugging Face14klingfoley /Kling-Audio-Eval Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation 🌐 Website | 📖 arXiv 📋 Dataset Structure The dataset structure is as follows: Kling-Audio-Eval ├── Folder (first-level label) │ ├── Folder (second-level label) │ │ ├── video │ │ │ └── *.mp4 │ │ ├── audio │ │ │ └── *.wav │ │ └── caption.csv # Header: video, audio, audio_tag, video_caption… See the full description on the dataset page: https://huggingface.co/datasets/klingfoley/Kling-Audio-Eval.1K<n<10K15 likes12k downloads1y agoHugging Face15hf-internal-testing /audiofolder_two_configs_in_metadata_with_defaultaudion<1K0 likes12k downloads3y agoHugging Face16yangwang825 /audioset AudioSet AudioSet[1] consists of an expanding ontology of 527 audio event classes and a collection of 2M human-labelled 10-second sound clips drawn from YouTube. Some clips are missing on YouTube, so the number of files downloaded is different from time to time. This repository contains 20550 / 22160 of the balanced train set, 1913637 / 2041789 of the unbalanced train set (separated into 41 parts), and 18887 / 20371 of the evaluation set. The pre-process script can be found at… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/audioset.textaudio-classification1M<n<10M5 likes12k downloads3y agoHugging Face17ilanashapiro /stg-paired-audioaudio0 likes11k downloads1y agoHugging Face18huggingface-course /audio-course-imagesimagen<1K0 likes11k downloads3y agoHugging Face19LeBeGut /AudioCodecBench2 likes9.7k downloads11mo agoHugging Face20Muncy /AudiobookRu AudiobookRu Russian-language audiobook audio (chapter-level), scraped from the web. Each row is one chapter: embedded MP3 bytes + book/chapter metadata. tabular100K<n<1M8 likes8.7k downloads12d agoHugging Face21moonshine-ai /audio_samples_1kaudio0 likes8.1k downloads6mo agoHugging Face22Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes7.7k downloads2mo agoHugging Face23humair025 /suno-audio 🎵 Suno Audio Dataset A comprehensive dataset of 49,698 AI-generated music tracks from Suno, organized in 50 batches of 1000 samples each. 🎧 All audio files are playable directly in the dataset viewer! Dataset Structure The dataset is organized into batches (batch_0, batch_1, etc.), each containing up to 1000 audio samples with metadata. Fields audio: 🎵 Playable MP3 audio file (click to play in viewer!) id: Unique track identifier title: Song title… See the full description on the dataset page: https://huggingface.co/datasets/humair025/suno-audio.audiotext-to-audio10K<n<100K2 likes7.1k downloads8mo agoHugging Face24liumindmind /Neko_Audio-80K_Short audio10K<n<100K30 likes6.8k downloads4mo agoHugging Face25Reza2kn /telegram-audiobook-chizzled Telegram Persian Audiobook Chizzled 1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.audion<1K4 likes6.7k downloads9d agoHugging Face26AudioLLMs /Multitask-National-Speech-Corpus-v1-extendaudio10M<n<100M5 likes6.4k downloads1y agoHugging Face27takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.3k downloads6mo agoHugging Face28nader39 /audio-filesaudion<1K0 likes5.9k downloads4d agoHugging Face29Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes5.3k downloads2mo agoHugging Face30RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes5k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.