CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tugrulbayrak /Real-TurnTurk Real-TurnTurk English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.tabularaudio-classification100K<n<1M3 likes543 downloads4d agoHugging Face02PrathamOrgAI /ReadNet ReadNet Dataset Description ReadNet is an audio dataset collected from more than 2,00,000 children in the age group of 5-16 years in Hindi and Marathi language. The dataset consists of audio files in the wav format where children read out ASER Samples which consists of letters, words, stories and paragraphs in their native language. This dataset is a subset of the larger dataset which consists of an estimated 2500 hours of data. This dataset consists of ~87 hours… See the full description on the dataset page: https://huggingface.co/datasets/PrathamOrgAI/ReadNet.textautomatic-speech-recognition10K<n<100K0 likes246 downloads2mo agoHugging Face03FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes83 downloads5mo agoHugging Face04UniDataPro /slovenian-speech-recognition Slovenian Speech Dataset Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.audioautomatic-speech-recognitionn<1K3 likes81 downloads1mo agoHugging Face05UniDataPro /vietnamese-speech-recognition Vietnamese Speech Dataset Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.audioautomatic-speech-recognitionn<1K4 likes76 downloads1mo agoHugging Face06uy-rrodriguez /BLURB-synthgated BLURB-synth: Synthetic audio data based on BLURB corpora Dataset Summary Synthetic audio data based on BLURB corpora. More details coming soon... Supported Tasks and Leaderboards Biomedical Language Understanding and Reasoning Benchmark (BLURB) Text-to-Speech Automatic-Speech-Recognition Languages English Data Structure Data Instances Coming soon... Data Fields Coming soon...… See the full description on the dataset page: https://huggingface.co/datasets/uy-rrodriguez/BLURB-synth.tabulartext-to-speech1M<n<10M0 likes65 downloads1d agoHugging Face07ud-nlp /human-robot-conversation-korean Human-Robot Conversation Dataset (Korean) - 660+ Hours Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes49 downloads6mo agoHugging Face08UniDataPro /human-robot-conversation-korean Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes48 downloads1mo agoHugging Face09ivrit-ai /crowd-recital-yigated About This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project. Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read. Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below). The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.audioautomatic-speech-recognition1K<n<10K0 likes46 downloads10mo agoHugging Face10researchaudio /apple-speechanalyzer-vs-whisper-cpp-mac Apple SpeechAnalyzer vs whisper.cpp on Mac Four complete speech-recognition benchmark runs over the same deterministic 40-speaker LibriSpeech test-clean snapshot: Engine Model path WER CER Repeated median post-speech latency Repeated p95 Apple SpeechAnalyzer progressiveTranscription on macOS 26.5 1.98% 1.02% 125–132 ms 194–201 ms whisper.cpp server 1.8.4 · ggml-small.en 4.28% 1.79% 122–125 ms 152–161 ms Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.tabularautomatic-speech-recognitionn<1K0 likes45 downloads2mo agoHugging Face11UniDataPro /human-robot-conversation-german Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.audioautomatic-speech-recognitionn<1K1 likes41 downloads1mo agoHugging Face12TransferRapid /CommonVoices20_ro Common Voices Corpus 20.0 (Romanian) Common Voices is an open-source dataset of speech recordings created by Mozilla to improve speech recognition technologies. It consists of crowdsourced voice samples in multiple languages, contributed by volunteers worldwide. Challenges: The raw dataset included numerous recordings with incorrect transcriptions or those requiring adjustments, such as sampling rate modifications, conversion to .wav format, and other refinements essential… See the full description on the dataset page: https://huggingface.co/datasets/TransferRapid/CommonVoices20_ro.audioautomatic-speech-recognition10K<n<100K4 likes37 downloads2y agoHugging Face13UniDataPro /korean-speech-recognition Korean Speech Dataset Dataset comprises 10+ hours of audio recordings from 20+ speakers, featuring telephone-quality speech data from native korean speakers. It provides a diverse collection of spoken language for automatic speech recognition tasks and serves as essential training data for model training in NLP and speech detection research. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/korean-speech-recognition.audioautomatic-speech-recognitionn<1K1 likes37 downloads1mo agoHugging Face14Speech-data /russian-speech-dataset Russian Speech Dataset The Russian Speech Dataset is a structured speech audio dataset designed to deliver high-quality audio data for machine learning and AI-driven voice systems. It includes 91 hours of audio data distributed across 641 files, provided in MP3 and WAV formats with a total size of 307 MB. This well-organized audio dataset ensures balanced voice data, with 50% female and 50% male speakers, and a broad age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/russian-speech-dataset.audioautomatic-speech-recognitionn<1K0 likes37 downloads6mo agoHugging Face15bengaliAI /ben10-asr-results Ben-10 Regional ASR — public results Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard. Field Meaning model_id Hub id or slug model_url Link to weights / paper wer Corpus Word Error Rate on private ben-10-test (lower better) wer_by_region JSON map region → WER backend Decode stack used by maintainers scorer_commit / decode_commit Git SHAs in BengaliAI/reg-speech-aacl evaluated_at ISO date requested_by Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.tabularautomatic-speech-recognitionn<1K0 likes37 downloads2mo agoHugging Face16UniDataPro /human-robot-conversation-english Human-Robot Dataset The dataset comprises 660+ hours of English speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-english.audioautomatic-speech-recognitionn<1K0 likes34 downloads1mo agoHugging Face17ud-nlp /russian-speech-recognition-dataset Russian Telephone Dialogues Dataset - 338 Hours The Russian speech dataset includes 338 hours of telephone dialogues in Russian from 460 native speakers, offering high-quality audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for building accurate Russian dialogue and audio datasets. - Get the data Dataset characteristics: Characteristic Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/russian-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes33 downloads8mo agoHugging Face18UniDataPro /british-english-speech-recognition-dataset British English Speech Dataset for recognition task Dataset comprises 200 hours of high-quality audio recordings featuring 310 speakers, achieving an impressive 95% Sentence Accuracy Rate. This extensive collection of speech data is designed for NLP tasks such as speech recognition, dialogue systems, and language understanding. By utilizing this dataset, developers and researchers can advance their work in automatic speech recognition and improve recognition systems. - Get the… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/british-english-speech-recognition-dataset.textautomatic-speech-recognitionn<1K1 likes31 downloads1mo agoHugging Face19ud-nlp /hindi-speech-recognition-dataset Hindi Telephone Dialogues Dataset - 760 Hours Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes29 downloads8mo agoHugging Face20ud-nlp /british-english-speech-recognition-dataset British English Telephone Dialogues Dataset - 200 Hours The dataset consists of 200 hours of high-quality telephone dialogues from 310 native speakers in the UK, with detailed annotations (transcriptions, timestamps, speaker ID, gender, and background noise) to support speech recognition systems, NLP tasks, and machine learning models requiring diverse British English audio datasets. - Get the data Dataset characteristics: Characteristic Data Description Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/british-english-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes29 downloads8mo agoHugging Face21ud-nlp /human-robot-conversation-english Human-Robot Conversation Dataset (English) - 660+ Hours Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.audioautomatic-speech-recognitionn<1K1 likes29 downloads6mo agoHugging Face22UniDataPro /hindi-speech-recognition-dataset Hindi Speech Dataset for recognition task Dataset comprises 760 hours of telephone dialogues in Hindi, collected from 1,000+ native speakers across various topics and domains. This dataset boasts an impressive 95% sentence accuracy rate, making it a valuable resource for advancing speech recognition technology. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K1 likes26 downloads1mo agoHugging Face23Speech-data /Romanian-Speech-Dataset 🎧 Romanian Speech Dataset The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes25 downloads6mo agoHugging Face24UniDataPro /french-speech-recognition-dataset French Speech Dataset for recognition task Dataset comprises 547 hours of telephone dialogues in French, collected from 964 native speakers across various topics and domains, with an impressive 98% Word Accuracy Rate. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/french-speech-recognition-dataset.textautomatic-speech-recognitionn<1K1 likes21 downloads1mo agoHugging Face25UniDataPro /german-speech-recognition-dataset German Speech Dataset for recognition task Dataset comprises 431 hours of telephone dialogues in German, collected from 590+ native speakers across various topics and domains, achieving an impressive 95% sentence accuracy rate. It is designed for research in automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural language processing (NLP). - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/german-speech-recognition-dataset.textautomatic-speech-recognitionn<1K1 likes21 downloads1mo agoHugging Face26ud-nlp /french-speech-recognition-dataset French Telephone Dialogues Dataset - 547 Hours his speech recognition dataset comprises 547 hours of telephone dialogues in French from 964 native speakers, providing audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for training and evaluating automatic speech recognition technology. - Get the data Dataset characteristics: Characteristic Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/french-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes20 downloads8mo agoHugging Face27ud-nlp /korean-speech-recognition Korean Speech Recognition Dataset - 10+ hours Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes20 downloads10mo agoHugging Face28ud-nlp /german-speech-recognition-dataset German Telephone Dialogues Dataset - 431 Hours Dataset comprises 431 hours of high-quality audio recordings from 590+ native German speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating German speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in German for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/german-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes18 downloads8mo agoHugging Face29Rizul2159 /WildVid-LIP WildVid-LIP: In-The-Wild Temporal Anchors for Visual Speech Recognition WildVid-LIP is a large-scale, open-source dataset mapping over 100,000 curated temporal segments from unconstrained, real-world YouTube videos. It provides precise timestamp anchors optimized for training Visual Speech Recognition (VSR / Lip-Reading), audio-visual synchronization, and multimodal self-supervised models. Instead of distributing heavy, monolithic video files—which introduces platform friction… See the full description on the dataset page: https://huggingface.co/datasets/Rizul2159/WildVid-LIP.tabularautomatic-speech-recognition100K<n<1M1 likes18 downloads3mo agoHugging Face30ud-nlp /vietnamese-speech-recognition Vietnamese Speech Dataset - 10+ hours Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/vietnamese-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes15 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.