CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.4k downloads10mo agoHugging Face02nmac /lex_fridman_podcast Dataset Card for "lex_fridman_podcast" Dataset Summary This dataset contains transcripts from the Lex Fridman podcast (Episodes 1 to 325). The transcripts were generated using OpenAI Whisper (large model) and made publicly available at: https://karpathy.ai/lexicap/index.html. Languages English Dataset Structure The dataset contains around 803K entries, consisting of audio transcripts generated from episodes 1 to 325 of the Lex Fridman… See the full description on the dataset page: https://huggingface.co/datasets/nmac/lex_fridman_podcast.textautomatic-speech-recognition100K<n<1M9 likes96 downloads4y agoHugging Face03Charif-Ayfarah /Afar-language-text-to-speech-TTS Usage This dataset is designed to support the development of Text-to-Speech (TTS) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition. For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended: Female Voices: Emeli, Hanaawi, Kareera, Laysani, Kulsuma… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-text-to-speech-TTS.text-to-speech1 likes56 downloads10mo agoHugging Face04maristombayeva /lost-in-speechgated Lost in Speech A trilingual benchmark for reference-free classification of synthetically introduced factual and contextual alterations in English, Russian, and Kazakh. It contains 12,013 samples derived from news articles, with text, synthesized speech, and ASR transcript representations used in the study. Altered samples are LLM-generated rewrites with a controlled alteration type—contradiction, fabrication, or context inconsistency—and severity level—mild, moderate, or severe.… See the full description on the dataset page: https://huggingface.co/datasets/maristombayeva/lost-in-speech.audiotext-classification1K<n<10K0 likes38 downloads10d agoHugging Face05uam-wmi-asr-eval-labs /2026-dwesui-g01-neurologia DWESUI 2026 - Grupa 1 - neurologia Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny. Zespol (atrybucja): Grupa 1 (DWESUI 2026) Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia Domena: neurologia Licencja zrodla: nagrania YouTube CC-BY + synteza TTS Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.audioautomatic-speech-recognitionn<1K0 likes34 downloads1mo agoHugging Face06Speech-data /Latvian-Speech-Dataset Latvian Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Latvian (la) 🏷️ Tags Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Latvian 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes29 downloads6mo agoHugging Face07uam-wmi-asr-eval-labs /2026-dwesui-g02-kulinarna DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA) Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny. Zespol (atrybucja): Grupa 2 (DWESUI 2026) Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna Domena: kulinarna Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.audioautomatic-speech-recognitionn<1K0 likes21 downloads1mo agoHugging Face08Codyfederer /last last This is a merged speech dataset containing 345 audio segments from 2 source datasets. Dataset Information Total Segments: 345 Speakers: 7 Languages: en Emotions: neutral, angry, happy, sad Original Datasets: 2 Dataset Structure Each example contains: audio: Audio file (WAV format, 16kHz sampling rate) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged datasets) emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/last.audioautomatic-speech-recognitionn<1K0 likes20 downloads1y agoHugging Face09Speech-data /Lithuanian-Speech-Dataset Lithuanian Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Lithuanian (lt) 🏷️ Tags Lithuanian, Audio, Speech, Speech Recognition, ML, Machine, Machine Learning 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes18 downloads6mo agoHugging Face10Speech-data /Luxembourgish-Speech-Dataset Luxembourgish Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Luxembourgish (lb) 🏷️ Tags Audio, Speech, Speech Recognition, Machine, Machine Learning, ML 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes17 downloads6mo agoHugging Face11Rizul2159 /WildVid-LIP WildVid-LIP: In-The-Wild Temporal Anchors for Visual Speech Recognition WildVid-LIP is a large-scale, open-source dataset mapping over 100,000 curated temporal segments from unconstrained, real-world YouTube videos. It provides precise timestamp anchors optimized for training Visual Speech Recognition (VSR / Lip-Reading), audio-visual synchronization, and multimodal self-supervised models. Instead of distributing heavy, monolithic video files—which introduces platform friction… See the full description on the dataset page: https://huggingface.co/datasets/Rizul2159/WildVid-LIP.tabularautomatic-speech-recognition100K<n<1M1 likes17 downloads3mo agoHugging Face12Luel-ai /luel-multilingual-tts-samplesgated Multilingual TTS Samples (Luel) License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE. A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.audiotext-to-speechn<1K0 likes15 downloads5mo agoHugging Face13Speech-data /Lingala-Speech-Dataset Lingala Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Lingala (ln) 🏷️ Tags Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Lingala 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes13 downloads6mo agoHugging Face14real-recordings /LibriReplay-DOA LibriReplay-DOA (Anonymous Submission) Overview LibriReplay-DOA is a multi-channel multi-speaker replay dataset designed for evaluating robust speech processing systems under realistic room playback conditions. The dataset contains replay recordings captured in real rooms under multiple playback configurations (DOA settings). Each session includes multiple overlapping speakers. This dataset is released for peer-review purposes. Dataset Structure The dataset… See the full description on the dataset page: https://huggingface.co/datasets/real-recordings/LibriReplay-DOA.audioautomatic-speech-recognition1K<n<10K1 likes11 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.