CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face02Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes8.2k downloads3mo agoHugging Face03xg-chu /UniLSTalkDataset UniLS-Talk Dataset To enable research on unified speaking and listening avatar generation, we curate and construct the UniLS-Talk Dataset, a large-scale collection of high-quality 3D facial motion data. We apply a carefully designed tracking pipeline to extract per-frame FLAME parameters, including expression coefficients, eye-gaze, jaw pose and head pose annotations. The dataset comprises two complementary parts: Paired conversational data sourced from the Seamless Interaction… See the full description on the dataset page: https://huggingface.co/datasets/xg-chu/UniLSTalkDataset.audio2 likes5.9k downloads7mo agoHugging Face04UncovAI /Real_Voiceaudio100K<n<1M1 likes4.2k downloads1y agoHugging Face05QUD-Technologies /quranic-universal-ayahs Qur'anic Universal Ayahs Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset. This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.audioautomatic-speech-recognition100K<n<1M6 likes3.2k downloads10h agoHugging Face06UncovAI /FOR-normaudio0 likes2.4k downloads1y agoHugging Face07syvai /danish-asr-unified Danish ASR Unified Dataset Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours): Source Samples Description VoxPopuli 1,775,578 European Parliament recordings ftspeech 995,677 Danish Parliament (Folketinget) CoRal-v3 read_aloud 299,255 Read-aloud Danish speech nst-da 182,605 NST Danish speech CoRal-v3 conversation 147,249 Conversational Danish speech nota 98,600 Danish broadcast media Common Voice 17 3,484 Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.audioautomatic-speech-recognition1M<n<10M5 likes2.2k downloads2mo agoHugging Face08hzhongresearch /ahead_ds_unmixed Another HEaring AiD DataSet (AHEAD-DS) unmixed Another HEaring AiD DataSet (AHEAD-DS) unmixed is an audio dataset labelled with audiologically relevant scene categories for hearing aids. This dataset contains the environment and speech sounds before they were mixed. The file ahead_ds_unmixed.csv documents the details of every file. Website Paper Code Dataset AHEAD-DS Dataset AHEAD-DS unmixed Models Description of data All files are encoded as single channel WAV, 16 bit… See the full description on the dataset page: https://huggingface.co/datasets/hzhongresearch/ahead_ds_unmixed.audioaudio-classification10K<n<100K0 likes2.1k downloads9mo agoHugging Face09TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes1.8k downloads6mo agoHugging Face10kiarashQ /farsi-asr-unified-cleaned 🎧 Farsi ASR Unified Dataset (Parquet Sharded Edition) Overview The Farsi ASR Unified Dataset is a large-scale, high-quality, and fully standardized collection of Persian (Farsi) speech-to-text data — designed specifically for modern machine learning and ASR (Automatic Speech Recognition) workflows. This dataset consolidates audio–text pairs from multiple open sources, applies a rigorous cleaning and normalization pipeline, and stores everything efficiently in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/kiarashQ/farsi-asr-unified-cleaned.audio1M<n<10M6 likes1.5k downloads11mo agoHugging Face11meituan-longcat /UNO-Bench UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models 🔔News 🔥[2025/12/04] We have released the evaluation scripts uno-eval, a unified evaluation framework for omni-modal benchmarks. More benchmarks will be supported in the future. 🔥[2025/12/04] We have released the scoring model UNO-Scorer-Qwen3-14B. Feel free to use it! 👀 UNO-Bench Overview Multimodal Large Languages models have… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/UNO-Bench.audio1K<n<10K23 likes1.2k downloads10mo agoHugging Face12malerlab /maestro-unidac4-ytsv MAESTRO + ASAP audio and MIDI tokens (U-MusT) Tokenized MAESTRO v3.0.0 for U-MusT: DAC audio tokens and MT3-style MIDI event arrays, covering roughly 199 hours of Disklavier-captured piano performance with precisely aligned MIDI. This repository also contains ASAP-derived data. lmx/ and asap_note_events/ come from the ASAP dataset, whose audio is itself MAESTRO. Both carry the same license, so nothing conflicts, but the repository name mentions only one of the two corpora it… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/maestro-unidac4-ytsv.audio0 likes908 downloads23d agoHugging Face13laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes671 downloads1y agoHugging Face14JKA-NLP /unified-kannada-asr-1.0 Dataset Card for "unified-kannada-asr-1.0" More Information needed audio100K<n<1M1 likes576 downloads3y agoHugging Face15uniiiii /Whisper-fine-tune-2audio1M<n<10M0 likes565 downloads2y agoHugging Face16unlimitedbytes /hailuo-ai-voices Hailuo AI Voices Dataset 🎤 A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks. 📊 Dataset Overview The dataset provides a comprehensive collection of voice samples with the following features: Feature Description Audio Files High-quality WAV format recordings Transcription Accurate transcriptions of each… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-voices.audiotext-to-speech10K<n<100K9 likes556 downloads2y agoHugging Face17JST-SUPERB /MUSAN-speech_unit_part2 Dataset Card for "MUSAN-speech_unit_part2" More Information needed audio10K<n<100K0 likes552 downloads2y agoHugging Face18malerlab /bpsd-unirqvae3-unidac4-ytsv BPSD score-image, audio and notation tokens (U-MusT) Tokenized Beethoven Piano Sonata Dataset v2 for U-MusT — the test-only split, and the only corpus in the collection carrying all four modalities: score-image tokens, audio tokens, and LMX notation. Because it is held out for evaluation, the image tokens here are not shift-augmented: they have shape (1, 1, H, W, 4), a single tokenization. The audio tokens retain the 9-variant stack. BPSD ships no system-level image alignment… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/bpsd-unirqvae3-unidac4-ytsv.audio0 likes546 downloads23d agoHugging Face19dianavdavidson /iv_speaker_disjoint_sociodem_unaware_dsaudio100K<n<1M0 likes444 downloads1mo agoHugging Face20JST-SUPERB /MUSAN-speech_unit_part1 Dataset Card for "MUSAN-speech_unit_part1" More Information needed audio10K<n<100K0 likes437 downloads2y agoHugging Face21JST-SUPERB /MUSAN-noise_unit_part2 Dataset Card for "MUSAN-noise_unit_part2" More Information needed audio10K<n<100K0 likes371 downloads2y agoHugging Face22KlingTeam /UnityShotsBench UnityShots Benchmark A multilingual, multi-cultural k-shot storytelling benchmark for evaluating multi-shot audio-video generation. Each case is a short cinematic story told across several shots, with a consistent cast whose identity, voice, and world must persist across every cut. This is the evaluation benchmark released with UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating. 📄 Paper: arXiv:2606.21661 🌐 Project page:… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/UnityShotsBench.audiotext-to-video1K<n<10K10 likes358 downloads3mo agoHugging Face23JST-SUPERB /MUSAN-music_unit_part1 Dataset Card for "MUSAN-music_unit_part1" More Information needed audio10K<n<100K0 likes350 downloads2y agoHugging Face24UncovAI /ASVSpoof21_PA2audio100K<n<1M1 likes342 downloads1y agoHugging Face25zxc0135 /UnityShotsBench UnityShots Benchmark A multilingual, multi-cultural k-shot storytelling benchmark for evaluating multi-shot audio-video generation. Each case is a short cinematic story told across several shots, with a consistent cast whose identity, voice, and world must persist across every cut. This is the evaluation benchmark released with UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating. 📄 Paper: arXiv:2606.21661 🌐 Project page:… See the full description on the dataset page: https://huggingface.co/datasets/zxc0135/UnityShotsBench.audiotext-to-video1K<n<10K0 likes335 downloads3mo agoHugging Face26Shamus /United-Syn-Medaudio100K<n<1M1 likes304 downloads2y agoHugging Face27universe-team /universebench UniVerseBench The evaluation split of UniVerse (同谣).Training data lives in UniVerseSet. UniVerseBench is a multilingual folk-music understanding benchmark for large audio–language models (LALMs). It asks models to listen, not to guess from language priors. 「诗言志,歌永言,声依永,律和声。」—《尚书·舜典》 Sister dataset (training) universe-team/universeset Live museum demo http://143.89.224.8:8790/ What's here Two views of the same benchmark: Subset… See the full description on the dataset page: https://huggingface.co/datasets/universe-team/universebench.audioquestion-answering1K<n<10K0 likes297 downloads1mo agoHugging Face28unlimitedbytes /hailuo-ai-jokes Hailuo AI Jokes Dataset 🎤 A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks. 🎙️ Dataset Content The dataset contains a diverse set of synthetic voice recordings generated by Hailuo AI Audio. The texts are sourced from a variety of public domain jokes and humorous anecdotes. Each audio sample is accompanied by the… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-jokes.audiotext-to-speech10K<n<100K6 likes294 downloads2y agoHugging Face29JST-SUPERB /MUSAN-music_unitaudio10K<n<100K2 likes293 downloads2y agoHugging Face30voidful /librispeech_unit_speech Dataset Card for "librispeech_unit_speech" More Information needed audio1K<n<10K0 likes282 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.