CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Muncy /AudiobookRu AudiobookRu Russian-language audiobook audio (chapter-level), scraped from the web. Each row is one chapter: embedded MP3 bytes + book/chapter metadata. tabular100K<n<1M8 likes8.7k downloads12d agoHugging Face02Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes5.3k downloads2mo agoHugging Face03RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes5k downloads3mo agoHugging Face04hf-audio /open-asr-leaderboard-resultstabularn<1K0 likes4.5k downloads2d agoHugging Face05Hezep /AudioMarathon 🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audioaudio-classification1K<n<10K4 likes3.9k downloads10mo agoHugging Face06akazemian /audio-htmltabular10K<n<100K0 likes2.7k downloads1y agoHugging Face07Rcarvalo /audio_datasettabular1M<n<10M0 likes1.8k downloads6mo agoHugging Face08yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes1.1k downloads6mo agoHugging Face09quinnlue /audioset-opus-fbanks AudioSet Opus — INT8 log-mel filterbanks Precomputed Kaldi log-mel filterbank features for AudioSet, quantized to INT8. Derived from danjacobellis/audioset_opus_24kbps and danjacobellis/audioset_opus_24kbps_balanced, which were already deduplicated by content hash. The point of this dataset is to remove audio decoding and filterbank computation from the training loop. In a masked-autoencoder training step at 64×1008 geometry, the fbank front end costs ~41% of wall-clock at 1M… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset-opus-fbanks.tabularaudio-classification1M<n<10M0 likes777 downloads1mo agoHugging Face10hf-audio /leaderboard_longformtabularn<1K0 likes776 downloads2mo agoHugging Face11voidful /NMSQA_audiotabular10K<n<100K1 likes699 downloads4y agoHugging Face12Muno459 /AudioSet AudioSet Google's AudioSet with the audio: 1,780,876 of its 2,084,320 labelled 10 s YouTube segments as 48 kHz FLAC, each with a record of where its audio came from, measured quality, duplicate and eval-overlap flags, and a reason for every segment that could not be found. clips with audio hours classes format size 1,780,876 of 2,084,320 (85.4%) 4,904 527 48 kHz, 24-bit FLAC 2.25 TiB Two versions of every clip: Opus and AAC YouTube stores the… See the full description on the dataset page: https://huggingface.co/datasets/Muno459/AudioSet.audioaudio-classification1M<n<10M2 likes670 downloads7d agoHugging Face13zihan-audio /FreeSound_selectedtabular1K<n<10K0 likes575 downloads6mo agoHugging Face14ozefe /spotify_audio_features Spotify Tracks & Audio Features Dataset Overview This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research. Data Source The raw data for this dataset was originally gathered and hosted by Anna's Archive. Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/ozefe/spotify_audio_features.tabulartabular-regression100M<n<1B11 likes535 downloads9mo agoHugging Face15matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes481 downloads3y agoHugging Face16matlok /python-audio-copilot-training-using-function-knowledge-graphs Python Copilot Audio Training using Global Functions with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.tabulartext-to-audion<1K1 likes474 downloads3y agoHugging Face17pengyizhou /Audio-IOAI-Practice-1tabular100K<n<1M0 likes444 downloads4mo agoHugging Face18ai-music4you3 /enhanced-audiosnippets-long-2-8M Enhanced Audiosnippets Long 2.8M Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis. Dataset Summary Metric Value Total samples 2,633,037 Total audio hours 4,932 h Duration range 3.0s - 1124.3s Mean duration 6.7s Audio format WAV, 48kHz mono Tar files 1,410 Processing Pipeline Each audio sample was processed through: Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.tabularaudio-classification1M<n<10M1 likes427 downloads6mo agoHugging Face19matlok /python-audio-copilot-training-using-import-knowledge-graphs Python Copilot Audio Training using Imports with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.tabulartext-to-audion<1K0 likes416 downloads3y agoHugging Face20weebnan /spotify_audio_features_partitionedtabular100M<n<1B0 likes408 downloads2mo agoHugging Face21matlok /python-audio-copilot-training-using-inheritance-knowledge-graphs Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.tabulartext-to-audion<1K0 likes374 downloads3y agoHugging Face22matlok /python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-audion<1K0 likes364 downloads3y agoHugging Face23cairocode /IEMO_Audio_Text_Mergedtabular1K<n<10K0 likes336 downloads11mo agoHugging Face24cairocode /MSPI_Audio_Text_Mergedtabular1K<n<10K0 likes305 downloads10mo agoHugging Face25quranlab /quran-audio Dataset Card for QuranLab — Qur'an Recitation Audio A verse-aligned reference layer for Qur'an recitation audio: a unified taxonomy of 245 reciters, per-ayah and per-surah reference manifests that link every verse to its original public source, and word-precise timing released under CC-BY-4.0. This dataset hosts no audio files — it points to each recitation where it already lives and contributes the connective scholarship: taxonomy, alignment, and timing. QuranLab is a… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio.tabularautomatic-speech-recognition100K<n<1M1 likes303 downloads2mo agoHugging Face26nymtheescobar /bengali-talkshow-audio Bengali Talkshow Audio Dataset A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs. Dataset Description This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.audioaudio-classification1K<n<10K0 likes278 downloads8mo agoHugging Face27quranlab /quran-audio-text QuranLab — Verse-Aligned Quran Text + Recitation References This dataset joins QuranLab's canonical Hafs Arabic text to its per-ayah recitation references. Every row is one exact (recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly Simple-Clean transcript, and the corresponding audio_url. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.tabularautomatic-speech-recognition100K<n<1M1 likes268 downloads2mo agoHugging Face28Rcarvalo /audioFRv1tabular100K<n<1M0 likes267 downloads7mo agoHugging Face29P-Arpan /Spotify_Audio_features_2.3Mtabular1M<n<10M0 likes229 downloads1mo agoHugging Face30paodigitalhub /pao-audio-dataset 🎙️ Pa'O Audio Dataset ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ 📌 Project Summary The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ). Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.audioautomatic-speech-recognitionn<1K1 likes226 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.