CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ARTPARK-IISc /Vaani-transcription-partgatedThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages. This table represents the audio and transcription duration data for various languages. Language Angami Angika Ao Assamese Awadhi Bajjika Bearybashe Bengali Bhili Bhojpuri Bundeli Chakhesang Chakma Chhattisgarhi English Garhwali Garo Gondi Gujarati Halbi Haryanvi Hindi IduMishmi Kannada Kashmiri Karbi Khariboli Khortha Kokborok Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.audioautomatic-speech-recognition1M<n<10M20 likes1.6k downloads6mo agoHugging Face02ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes758 downloads22d agoHugging Face03ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes596 downloads22d agoHugging Face04openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes443 downloads6mo agoHugging Face05modulate /entity-transcription-benchmark Entity Transcription Benchmark Measures whether a speech recognition system transcribes named entities correctly — as distinct from word error rate. WER weights every token equally. The tokens that matter for redaction, lookup, routing and search are proper nouns, and they are a small fraction of any transcript. A system can improve WER while getting worse at exactly the words a downstream consumer needs, and nothing in the standard evaluation will show it. 2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.audioautomatic-speech-recognition1K<n<10K4 likes336 downloads13d agoHugging Face06POTOMITAN /potomitan-gcf-transcriptiongated Kreyol Guadeloupe Transcription Dataset Ce jeu de données contient des segments audio courts (~5 secondes) en créole guadeloupéen (gcf), extraits d’émissions de radio et de télévision. Il vise à entraîner des modèles de reconnaissance automatique de la parole (ASR) pour une langue vivante mais peu disposant de peu de ressources écrites. Dataset Description Le créole guadeloupéen (Karukéya) est une langue créole à base lexicale française, parlée principalement en… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/potomitan-gcf-transcription.audioautomatic-speech-recognition10K<n<100K4 likes225 downloads10mo agoHugging Face07united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes222 downloads7mo agoHugging Face08AIxBlock /doctor-patient-convers-transcriptions-PII-redactedThis dataset contains real-world transcriptions of doctor–patient conversations in English (USA accent), focused on two medical specialties: ENT (Ear, Nose, Throat) and Dermatology and Orthopaedic. All conversations were originally recorded in clinical settings and transcribed by human experts. To comply with privacy regulations, only the transcription files are released, with all personally identifiable information (PII) fully redacted. Due to regulations, we are only able to publish… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/doctor-patient-convers-transcriptions-PII-redacted.automatic-speech-recognition0 likes161 downloads1y agoHugging Face09tech4humans /Audio-Transcription-Models-Comparison-PT-BR Audio Transcription Models Comparison A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese. About the Dataset This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering: Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.audioautomatic-speech-recognitionn<1K3 likes144 downloads8mo agoHugging Face10serge-wilson /wolof_speech_transcription Wolof Speech Transcription Description Dataset de reconnaissance automatique de la parole (ASR) en wolof, une langue d'Afrique de l'Ouest parlée par plus de 10 millions de locuteurs, principalement au Sénégal. Ce dataset est un miroir HuggingFace du corpus wolof du projet ALFFA hébergé à l'origine sur GitHub par le laboratoire GETALP (Grenoble). Source originale Ce dataset provient du projet ALFFA : Repository : getalp/ALFFA_PUBLIC Laboratoire : GETALP… See the full description on the dataset page: https://huggingface.co/datasets/serge-wilson/wolof_speech_transcription.audioautomatic-speech-recognition10K<n<100K3 likes69 downloads6mo agoHugging Face11thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes62 downloads4mo agoHugging Face12AIxBlock /Eng-Filipino-Accented-audio-with-human-transcription-call-center-topicThis dataset contains 103+ hours of spontaneous English conversations spoken in a Filipino accent, recorded in a studio environment to ensure crystal-clear audio quality. The conversations are designed as role-play scenarios between agents and customers across a variety of call center domains. 🗣️ Speech Style: Natural, unscripted role-playing between native Filipino-accented English speakers, simulating real-world customer interactions. 🎧 Audio Format: High-quality stereo WAV files, recorded… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Eng-Filipino-Accented-audio-with-human-transcription-call-center-topic.audioautomatic-speech-recognitionn<1K5 likes51 downloads1y agoHugging Face13united-nations /transcription-results UN Transcription Benchmark Results Evaluation results for speech-to-text systems on UN Security Council and General Assembly meeting recordings, assessed against official UN verbatim records. See united-nations/transcription-corpus for the underlying audio and ground truth data. Metrics WER: Word Error Rate (reference = verbatim record, no normalization) normalized_wer: WER after lowercasing, punctuation removal, and filler word removal CER: Character Error Rate (same… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-results.automatic-speech-recognitionn<1K0 likes27 downloads6mo agoHugging Face14AWANNABY /French-Medical-Transcription-Benchmark 🩺 French Medical Transcription Evaluation Dataset Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped). 🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local 📊 Présentation du Dataset L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.textautomatic-speech-recognition1K<n<10K1 likes23 downloads6mo agoHugging Face15RobotsMali /transcription-scorer Transcription Scorer Dataset The Transcription Scorer dataset was created to support research in reference-free evaluation of Automatic Speech Recognition (ASR) systems using human feedback. Unlike traditional evaluation metrics such as WER and its derivatives, this dataset reflects judgments of ASR outputs by human raters across multiple criteria, simulating the way a teacher grades students. ⚙️ What’s Inside This dataset contains 1200 audio samples (from diverse sources… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/transcription-scorer.audioautomatic-speech-recognition1K<n<10K1 likes21 downloads1y agoHugging Face16AIxBlock /Thai-H2H-Call-center-audio-with-human-transcriptionThis dataset contains natural Thai-language conversations between human agents and human customers, designed to reflect realistic call center interactions across multiple domains. All conversations are conducted through unscripted role-playing, allowing for spontaneous and dynamic exchanges that closely mirror real-world scenarios. 🗣️ Speech Type: Human-to-human dialogues simulating customer-agent interactions. 🎭 Style: Non-scripted, spontaneous role-playing to capture authentic speech… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-H2H-Call-center-audio-with-human-transcription.automatic-speech-recognition3 likes18 downloads1y agoHugging Face17AIxBlock /English-USA-NY-Boston-AAVE-Audio-with-transcriptionThis dataset captures spontaneous English conversations from native U.S. speakers across distinct regional and cultural accents, including: 🗽 New York English 🎓 Boston English 🎤 African American Vernacular English (AAVE) The recordings span three real-life scenarios: General Conversations – informal, everyday discussions between peers. Call Center Simulations – customer-agent style interactions mimicking real support environments. Media Dialogue – scripted reads and semi-spontaneous… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/English-USA-NY-Boston-AAVE-Audio-with-transcription.text-to-audio2 likes11 downloads1y agoHugging Face18narinzar /massive-audio-transcription-pipeline massive-audio-transcription-pipeline outputs Transcription outputs from the massive-audio-transcription-pipeline, a parallel Whisper pipeline that chunks long audio into overlapping windows, transcribes across a worker pool, merges lightweight speaker diarization, and checkpoints every chunk for crash resume. Generation method Backend: faster-whisper base model (CTranslate2), 1 worker. Audio: real public-domain speech from the Hugging Face LibriSpeech dummy… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/massive-audio-transcription-pipeline.automatic-speech-recognition0 likes6 downloads3mo agoHugging Face19iioos /speech-transcription-samples Speech Transcription Samples Dataset Synthetic speech-to-text transcription samples for ASR research. automatic-speech-recognition0 likes5 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.