CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face02QUD-Technologies /quranic-universal-ayahs Qur'anic Universal Ayahs Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset. This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.audioautomatic-speech-recognition100K<n<1M6 likes3.2k downloads16h agoHugging Face03syvai /danish-asr-unified Danish ASR Unified Dataset Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours): Source Samples Description VoxPopuli 1,775,578 European Parliament recordings ftspeech 995,677 Danish Parliament (Folketinget) CoRal-v3 read_aloud 299,255 Read-aloud Danish speech nst-da 182,605 NST Danish speech CoRal-v3 conversation 147,249 Conversational Danish speech nota 98,600 Danish broadcast media Common Voice 17 3,484 Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.audioautomatic-speech-recognition1M<n<10M5 likes2.2k downloads2mo agoHugging Face04TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes1.8k downloads6mo agoHugging Face05unlimitedbytes /hailuo-ai-voices Hailuo AI Voices Dataset 🎤 A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks. 📊 Dataset Overview The dataset provides a comprehensive collection of voice samples with the following features: Feature Description Audio Files High-quality WAV format recordings Transcription Accurate transcriptions of each… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-voices.audiotext-to-speech10K<n<100K9 likes556 downloads2y agoHugging Face06unlimitedbytes /hailuo-ai-jokes Hailuo AI Jokes Dataset 🎤 A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks. 🎙️ Dataset Content The dataset contains a diverse set of synthetic voice recordings generated by Hailuo AI Audio. The texts are sourced from a variety of public domain jokes and humorous anecdotes. Each audio sample is accompanied by the… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-jokes.audiotext-to-speech10K<n<100K6 likes294 downloads2y agoHugging Face07am-pranav /stt-unified-bench am-pranav/stt-unified-bench Private, language/locale-partitioned mini-benchmark for STT models. Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val. Audio is staged at 16 kHz and stored in-repo for reproducibility. Schema audio : Audio(sampling_rate=16000, decode=False) text : reference transcription lang : implied by dataset config name source : upstream dataset tag id : source-stable id ⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.audioautomatic-speech-recognition10K<n<100K0 likes276 downloads1y agoHugging Face08UniDataPro /real-vs-fake-human-voice-deepfake-audio Deepfake Audio Dataset Dataset contains 5,000 audio files, comprising both authentic human recordings and synthetic** AI-generated voice** samples. It designed for advanced research in deepfake detection, focusing on detecting fake voices and generated speech analysis. Specifically engineered to challenge voice authentication systems, it supports the development of robust models for real vs fake human voice recognition. By utilizing this dataset, researchers and developers can… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/real-vs-fake-human-voice-deepfake-audio.audioautomatic-speech-recognitionn<1K5 likes260 downloads1mo agoHugging Face09manassehzw /sna-waxal-annotated-unlabeled Shona WAXAL annotated-unlabeled checkpoint This is a self-contained operational checkpoint for pseudo-labeling Shona ASR data. It contains 90,253 conservatively segmented FLAC clips (441.585 hours), but intentionally contains no transcripts. Fields transcription is intentionally empty. speaker_id is an approximate source-blind EOM cluster or unknown; speaker_clip_count is zero for unknown assignments. gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.audioautomatic-speech-recognition10K<n<100K0 likes242 downloads2mo agoHugging Face10united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes217 downloads7mo agoHugging Face11doof-ferb /VietMed_unlabeled unofficial mirror of VietMed (Vietnamese speech data in medical domain) unlabeled set official announcement: https://arxiv.org/abs/2404.05659 official download: https://huggingface.co/datasets/leduckhai/VietMed this repo contains the unlabeled set: 966h - 230k samples i also gather the metadata: see info.csv my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-unlabeled.py need to do: check misspelling, restore foreign words phonetised to… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_unlabeled.audioautomatic-speech-recognition100K<n<1M1 likes204 downloads2y agoHugging Face12emix-1 /unbound009_2_tran emix-1/unbound009_2_tran This dataset contains transcribed audio files organized in folders for scalability. Dataset Structure The dataset is organized with: Audio files: Stored in audio_XXXXX/ folders (5000 files per folder) Metadata: Stored in data_XXXXX/ folders as parquet files This organization follows Hugging Face best practices for datasets with millions of files. Statistics Total files: 3,174 Total batches: 1409 Audio folders: 2 Files per folder: max… See the full description on the dataset page: https://huggingface.co/datasets/emix-1/unbound009_2_tran.audioautomatic-speech-recognition10K<n<100K0 likes184 downloads11mo agoHugging Face13UniDataPro /speech-emotion-recognition Speech Emotion Recognition Dataset comprises 30,000+ audio recordings featuring 4 distinct emotions: euphoria, joy, sadness, and surprise. This extensive collection is designed for research in emotion recognition, focusing on the nuances of emotional speech and the subtleties of speech signals as individuals vocally express their feelings. By utilizing this dataset, researchers and developers can enhance their understanding of sentiment analysis and improve automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/speech-emotion-recognition.audioautomatic-speech-recognitionn<1K6 likes164 downloads1mo agoHugging Face14Sam04 /unfolded-veil-v9_traaudioautomatic-speech-recognition1K<n<10K0 likes144 downloads10mo agoHugging Face15UniDataPro /spanish-speech-recognition-dataset Spanish Speech Dataset for recognition task Dataset comprises 10 hours of telephone dialogues in Spanish, collected from 10 native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data The dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/spanish-speech-recognition-dataset.audioautomatic-speech-recognitionn<1K2 likes134 downloads1mo agoHugging Face16uncleMehrzad /synthetic-speaker-diarization-dataset-fa-large-3000audioaudio-classification1K<n<10K3 likes108 downloads1y agoHugging Face17UniDataPro /slovenian-speech-recognition Slovenian Speech Dataset Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.audioautomatic-speech-recognitionn<1K3 likes85 downloads1mo agoHugging Face18suleiman2003 /unified-hausa-speech Unified Hausa Speech Dataset v5 Dataset Description A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from 6 open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research on one of Africa's most widely spoken languages. Hausa (ISO 639-1: ha) is a Chadic language spoken by over 80 million people across West and Central Africa — primarily in Nigeria and Niger, and as a trade language… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/unified-hausa-speech.audiotext-to-speech100K<n<1M0 likes84 downloads1mo agoHugging Face19UniDataPro /vietnamese-speech-recognition Vietnamese Speech Dataset Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.audioautomatic-speech-recognitionn<1K4 likes82 downloads1mo agoHugging Face20TTS-AGI /majestrino-unified-detailed-captions-temporal Majestrino Unified Detailed Captions with Temporal Aspects Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects. Stats 4,128,665 samples 826 tar files (~1.1 GB each) ~878 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption with temporal aspects caption_type — always unified_detailed_caption_with_temporal_aspects transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.audioaudio-classification1M<n<10M0 likes81 downloads6mo agoHugging Face21undertheseanlp /uts2025_vietipa Vietnamese IPA Dataset A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications. Dataset Description Dataset Summary This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for: Text-to-speech systems development Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.audiotext-to-speechn<1K1 likes80 downloads1y agoHugging Face22UniDataPro /arabic-speech-recognition Arabic Speech Dataset Dataset comprises over 10 hours of audio featuring 20+ native speakers engaged in telephone-quality dialogues in the Arabic language. It contains high-quality speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. By utilizing this dataset, developers and researchers can advance their work in automatic speech recognition and improve recognition systems. - Get the data The… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/arabic-speech-recognition.audioautomatic-speech-recognitionn<1K1 likes80 downloads1mo agoHugging Face23ufal /parczech4speech-unsegmented ParCzech4Speech (Unsegmented Variant) Dataset Summary ParCzech4Speech (Unsegmented Variant) is a large-scale Czech speech dataset derived from parliamentary recordings and official transcripts. This variant captures continuous speech segments without enforcing sentence boundaries, making it well-suited for real-world streaming ASR scenarios and speech modeling tasks that benefit from natural discourse flow. The dataset is created using a combination of WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-unsegmented.audioautomatic-speech-recognition1M<n<10M1 likes79 downloads1y agoHugging Face24UniDataPro /human-robot-conversation-russian Human-Robot Dataset The dataset comprises 660+ hours of Russian speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-russian.audioautomatic-speech-recognitionn<1K1 likes64 downloads1mo agoHugging Face25baki83 /JuzneVesti-SR-Unsloth-Format Emilia-compatible Serbian speech (JuzneVesti-SR) This is a format conversion of JuzneVesti-SR v1.0 for Hugging Face audio training pipelines. It exposes the same columns as kadirnar/Emilia-DE-B000000 and preserves the original train/dev/test split (with dev named validation). Source Peter Rupnik and Nikola Ljubesic, ASR training dataset for Serbian JuzneVesti-SR v1.0, Jozef Stefan Institute / CLARIN.SI (2022). Persistent identifier:… See the full description on the dataset page: https://huggingface.co/datasets/baki83/JuzneVesti-SR-Unsloth-Format.audioautomatic-speech-recognition10K<n<100K0 likes61 downloads5d agoHugging Face26collectivat /una-fraza-al-diya Una fraza al diya Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total. Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/ Citation If you use this dataset, please cite: Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.audiotext-generationn<1K1 likes59 downloads11mo agoHugging Face27universe-team /universeset UniVerseSet The training split of UniVerse (同谣).Held-out evaluation lives in UniVerseBench. UniVerseSet is the post-training corpus for large audio–language models on world folk music: ASR, captions, and audio-grounded chat, plus automatically transcribed ABC scores. 「诗言志,歌永言,声依永,律和声。」—《尚书·舜典》 Sister dataset (benchmark) universe-team/universebench Live museum demo http://143.89.224.8:8790/ What's here Archives (download and unpack;… See the full description on the dataset page: https://huggingface.co/datasets/universe-team/universeset.audioaudio-text-to-text100K<n<1M0 likes53 downloads1mo agoHugging Face28UniDataPro /call-center-audio Call Center Dataset The datasets contain over 13,000+ hours of real-world conversations and feature 90%+ unique speakers. It is designed for advancing conversational AI and speech technologies. The dataset provides high-quality, time-stamped transcripts for model training in speech recognition and speaker diarization, enabling businesses to build and refine their AI systems. By utilizing this dataset, researchers and developers can focus on analyzing customer interactions… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/call-center-audio.audiotext-to-speechn<1K1 likes51 downloads1mo agoHugging Face29UniDataPro /human-robot-conversation-korean Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes51 downloads1mo agoHugging Face30UniDataPro /human-robot-conversation-german Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.audioautomatic-speech-recognitionn<1K1 likes42 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.