CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01undertheseanlp /uts2025_vietipa Vietnamese IPA Dataset A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications. Dataset Description Dataset Summary This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for: Text-to-speech systems development Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.audiotext-to-speechn<1K1 likes80 downloads1y agoHugging Face02maikezu /data-kit-sub-iwslt2025-if-long-constraint Data for KIT’s Instruction Following Submission for IWSLT 2025 This repo contains the data used to train our model for IWSLT 2025's Instruction-Following (IF) Speech Processing track. IWSLT 2025's Instruction-Following (IF) Speech Processing track in the scientific domain aims to benchmark foundation models that can follow natural language instructions—an ability well-established in textbased LLMs but still emerging in speech-based counterparts. Our approach employs an end-to-end… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/data-kit-sub-iwslt2025-if-long-constraint.textautomatic-speech-recognition0 likes39 downloads1y agoHugging Face03uam-wmi-asr-eval-labs /2025-zwesui-g02-medyczna ZWESUI 2025 - Grupa 2 - medyczna (mowa syntetyczna) Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2025 (pierwsza), tryb niestacjonarny. Zespol (atrybucja): Grupa 2 (2025) Zrodlo oryginalne: https://huggingface.co/datasets/yanvoi/med_male_female_r2 Domena: medyczna Opis: 200 zdań z terminologią medyczną, mowa syntetyczna (ElevenLabs), głosy męskie i żeńskie; zbiór użyty… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2025-zwesui-g02-medyczna.audioautomatic-speech-recognitionn<1K0 likes39 downloads1mo agoHugging Face04Trelis /medical-terms-2025 Medical Terms 2025 — Medical ASR Benchmark Entity-aware medical ASR benchmark — 50 hard rows with synthetic TTS audio of 2025 drug/condition terminology. Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here. Source 84 manually curated terms from 2025 FDA/EMA/WHO primary sources. Each term has source_url, source_date, and source_quality. Sentences generated by Gemini 2.5 Flash. Audio by Kokoro TTS via Trelis Studio… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/medical-terms-2025.audioautomatic-speech-recognitionn<1K0 likes22 downloads6mo agoHugging Face05uam-wmi-asr-eval-labs /2025-zwesui-g03-debata-prezydencka ZWESUI 2025 - Grupa 3 - debata prezydencka 2025 Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2025 (pierwsza), tryb niestacjonarny. Zespol (atrybucja): Grupa 3 (2025) Zrodlo oryginalne: https://huggingface.co/datasets/directtt/polish_presidential_debate Domena: debata prezydencka Opis: Debata prezydencka TVP z 12 maja 2025 — 13 kandydatów, 195 wypowiedzi, mowa spontaniczna.… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2025-zwesui-g03-debata-prezydencka.audioautomatic-speech-recognitionn<1K0 likes19 downloads1mo agoHugging Face06GeoPoll /dataset-20250728_102101-swgated GeoPoll Swahili Speech Dataset This dataset contains speech recognition data for Swahili (sw) collected and processed by GeoPoll. Dataset Summary This dataset is designed for fine-tuning speech recognition models on Swahili audio data. It includes high-quality audio segments with corresponding transcriptions. Dataset Statistics Total samples: 11814 Total duration: 20.45 hours Average duration: 6.23 seconds per sample Number of speakers: 6 Language: Swahili… See the full description on the dataset page: https://huggingface.co/datasets/GeoPoll/dataset-20250728_102101-sw.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.