CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M468 likes101k downloads4mo agoHugging Face02google /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.audioautomatic-speech-recognition1M<n<10M286 likes14k downloads22d agoHugging Face03Reza2kn /gooshkon-chunks Gooshkon MP3 chunks A single-column Hugging Face Audio dataset. Every row contains embedded, playable MP3 bytes in the audio column. Chunks target approximately 30 seconds and are selected at detected quiet intervals with a 12-48 second safety range. The source recordings are not transcript-aligned; this release uses acoustic silence boundaries to avoid cutting through words whenever a usable pause is available. Dataset runtime The dataset contains approximately 2… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/gooshkon-chunks.audioautomatic-speech-recognition100K<n<1M2 likes1.9k downloads15d agoHugging Face04bond005 /sberdevices_golos_10h_crowd Dataset Card for sberdevices_golos_10h_crowd Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.audioautomatic-speech-recognition10K<n<100K7 likes681 downloads4y agoHugging Face05adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hokkien Taiwan-Tongues-ASR-CE-dataset-hokkien 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.audioautomatic-speech-recognition10K<n<100K7 likes414 downloads9mo agoHugging Face06bond005 /sberdevices_golos_100h_farfield Dataset Card for sberdevices_golos_100h_farfield Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.audioautomatic-speech-recognition10K<n<100K6 likes326 downloads4y agoHugging Face07HTH-inc /japanese-casual-conversational-speech-golden-dataset-preview Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.audioautomatic-speech-recognitionn<1K2 likes246 downloads22d agoHugging Face08hadamard-2 /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K0 likes221 downloads10d agoHugging Face09hadamard-2 /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes216 downloads10d agoHugging Face10q1805 /german-golden-audio_speech-IPA 🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx) An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions. 📊 Dataset Summary Total Samples: 1,354 high-quality audio recordings. Total Size: ~419 MB (Compressed Parquet format). Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.audioautomatic-speech-recognition1K<n<10K0 likes216 downloads28d agoHugging Face11Reza2kn /visualears-golden-6669 🗂️ visualears-golden-6669 English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Held-out VisualEars6669 / Golden6669 evaluation dataset. مجموعهٔ ارزیابی نگه‌داشته‌شدهٔ Golden6669 با شرایط پاک، دورمیدان و مسدود برای سنجش واقع‌گرایانهٔ ASR فارسی. 🧩 Role evaluation and benchmarking asset مصنوع ارزیابی و بنچمارک 📦 Snapshot 13 files; approximately 916.66 MB 13 فایل؛ حدود… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-golden-6669.audioautomatic-speech-recognition1K<n<10K1 likes212 downloads2mo agoHugging Face12adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-zhtw Taiwan-Tongues-ASR-CE-dataset-zhtw 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.audioautomatic-speech-recognition100K<n<1M2 likes204 downloads9mo agoHugging Face13leyu-amharic /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K1 likes190 downloads2mo agoHugging Face14gheero-Leyu /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes188 downloads2mo agoHugging Face15Reubencf /goan-konkani-speech Goan Konkani Speech (Romi) 47,365 audio clips, 108.5 hours of Goan Konkani (ISO 639-3 gom) speech from Goan television news, transcribed in Romi Konkani - Konkani written in the Roman script. Konkani is a low-resource language with very little public speech data. This is assembled from broadcast news, so it is real spoken Konkani: studio anchors, field reporters, phone interviews, and the Konkani-English code-switching that Goan speakers actually use. Contents… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-speech.audioautomatic-speech-recognition10K<n<100K0 likes151 downloads5d agoHugging Face16leyu-amharic /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K1 likes149 downloads2mo agoHugging Face17google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes132 downloads3y agoHugging Face18InternalCan /gospel-aloe-hera-v6 Gospel Aloe — Hera duplex training set (v6) Everyday-conversation companion to InternalCan/gospel-didactic-hera-v6. Same codes-only Hera schema (32 Mimi codebooks, word-level timestamps, v6 B-channel corruption). Shards were packed on three nodes and published into this one repo: Source Shard prefix Role This 8×H100 box + node A local-*, nodeA-* ~3.1k conversations Node B (63.141.33.128) nodeB-* ~7.3k conversations sample_id is the join key. There is no overlap… See the full description on the dataset page: https://huggingface.co/datasets/InternalCan/gospel-aloe-hera-v6.tabularautomatic-speech-recognition10K<n<100K1 likes128 downloads7d agoHugging Face19adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hakka Taiwan-Tongues-ASR-CE-dataset-hakka 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hakka.audioautomatic-speech-recognition1K<n<10K2 likes120 downloads9mo agoHugging Face20Goekdeniz-Guelmez /mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs. audioautomatic-speech-recognitionn<1K0 likes112 downloads1mo agoHugging Face21oddadmix /arabic-audio-collection-sudanese-ahmed-gobara Ahmed Gobara Sudanese Arabic Speech Dataset Dataset Summary The Ahmed Gobara Sudanese Arabic Speech Dataset is a single-speaker Sudanese Arabic speech corpus containing approximately 19 hours of speech recordings and corresponding transcripts. While compact, the dataset offers a clean, consistent single-speaker resource in Sudanese Arabic — an Arabic variety with very few open speech resources — making it especially valuable for voice cloning, speaker adaptation… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-sudanese-ahmed-gobara.audiotext-to-speech1K<n<10K1 likes97 downloads2mo agoHugging Face22Codyfederer /goodforft goodforft This is a merged speech dataset containing 863 audio segments from 4 source datasets. Dataset Information Total Segments: 863 Speakers: 4 Languages: en Emotions: angry, happy, neutral Original Datasets: 4 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged datasets) emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/goodforft.audioautomatic-speech-recognitionn<1K1 likes94 downloads1y agoHugging Face23djsamseng /khmer-speech-large-english-google-translations Dataset Card for khmer-speech-large-english-google-translation Audio recordings of khmer speech with varying speakers and background noises. English transcriptions were transcribed from the Khmer labels using Google Translate. Based off of seanghay/khmer-speech-large. Dataset Details Dataset Sources Huggingface: seanghay/khmer-speech-large Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/djsamseng/khmer-speech-large-english-google-translations.audioautomatic-speech-recognition10K<n<100K5 likes91 downloads6mo agoHugging Face24chuuhtetnaing /myanmar-speech-dataset-google-fleursPlease visit to the GitHub repository for other Myanmar Langauge datasets. Myanmar Speech Dataset (Google Fleurs) This dataset consists exclusively of Myanmar speech recordings, extracted from the larger multilingual Google Fleurs dataset. For the complete multilingual dataset and additional information, please visit the original dataset repository of Google Fleurs HuggingFace page. Original Source Fleurs is the speech version of the FLoRes machine translation benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-speech-dataset-google-fleurs.audiotext-to-speech1K<n<10K0 likes80 downloads1y agoHugging Face25gheero-Leyu /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K1 likes53 downloads2mo agoHugging Face26adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-en Taiwan-Tongues-ASR-CE-dataset-en 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-en.audioautomatic-speech-recognition10K<n<100K0 likes52 downloads9mo agoHugging Face27Reza2kn /visualears-benchmark-269-gold 🗂️ visualears-benchmark-269-gold English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose 269-record gold/noisy benchmark dataset. معیار طلایی ۲۶۹ نمونه‌ای برای بررسی سریع خطاهای گفتار نویزی و مقایسهٔ نسخه‌های مدل. 🧩 Role evaluation and benchmarking asset مصنوع ارزیابی و بنچمارک 📦 Snapshot 276 files; approximately 44.35 MB 276 فایل؛ حدود 44.35 MB 🧱 Packaging 1 Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-benchmark-269-gold.audioautomatic-speech-recognitionn<1K1 likes45 downloads2mo agoHugging Face28alconost /alconost-multilingual-speech-goldgated Multilingual Speech & Translation Dataset — EN↔JA/AR-EG/PL/RU (10 phrases, dual-take) Description 10 English source phrases with expert human translations into Japanese, Egyptian Arabic (ar-EG), and Polish. Each target phrase is recorded by native speakers (two takes each). Audio files are WAV 48 kHz mono, 16‑bit PCM format. Translations are produced and QA'd by professional linguists; recordings follow consistent orthography/style (AR-EG: Egyptian dialect; JA/PL: standard). All… See the full description on the dataset page: https://huggingface.co/datasets/alconost/alconost-multilingual-speech-gold.audiotranslationn<1K0 likes30 downloads8mo agoHugging Face29Akabi /Luxemburgish_Press_Conferences_Govaudioautomatic-speech-recognitionn<1K3 likes28 downloads2y agoHugging Face30freococo /google_myanmar_asr_voices Google Myanmar ASR Dataset (WebDataset Version) This repository provides a clean, user-friendly, and robust version of the Google Myanmar ASR Dataset, which is derived from the OpenSLR-80 Burmese Speech Corpus. This version has been carefully re-processed into the WebDataset format. Each sample consists of a .wav audio file and a clean .json metadata file, packaged into sharded .tar archives. This format is highly efficient for large-scale training of ASR models.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/google_myanmar_asr_voices.audioautomatic-speech-recognition1K<n<10K0 likes27 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.