CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M468 likes101k downloads4mo agoHugging Face02google /WaxalNLP Waxal Datasets The WAXAL dataset is a large-scale multilingual speech corpus for African languages, introduced in the paper WAXAL: A Large-Scale Multilingual African Language Speech Corpus. Dataset Description The Waxal project provides datasets for both Automated Speech Recognition (ASR) and Text-to-Speech (TTS) for African languages. The goal of this dataset's creation and release is to facilitate research that improves the accuracy and fluency of speech and… See the full description on the dataset page: https://huggingface.co/datasets/google/WaxalNLP.audioautomatic-speech-recognition1M<n<10M286 likes14k downloads21d agoHugging Face03laion /laions_got_talent LAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel. Dataset Composition The dataset includes: Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.audio100K<n<1M41 likes9.8k downloads2y agoHugging Face04laion /laions_got_talent_rawaudio10K<n<100K7 likes3.5k downloads2y agoHugging Face05mazkooleg /0-9up_google_speech_commands_augmented_raw Dataset Card for "google_speech_commands_augmented_raw_fixed" More Information needed audio1M<n<10M0 likes3k downloads4y agoHugging Face06laion /laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants. Updated Composition Voices and Languages English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.audio1M<n<10M6 likes2.2k downloads1y agoHugging Face07goea /arc-voicesamples-generatedaudio100K<n<1M1 likes1.9k downloads11h agoHugging Face08goea /arc-speeches-refinedaudio10K<n<100K1 likes1.8k downloads11h agoHugging Face09Reza2kn /gooshkon-chunks Gooshkon MP3 chunks A single-column Hugging Face Audio dataset. Every row contains embedded, playable MP3 bytes in the audio column. Chunks target approximately 30 seconds and are selected at detected quiet intervals with a 12-48 second safety range. The source recordings are not transcript-aligned; this release uses acoustic silence boundaries to avoid cutting through words whenever a usable pause is available. Dataset runtime The dataset contains approximately 2… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/gooshkon-chunks.audioautomatic-speech-recognition100K<n<1M2 likes1.8k downloads15d agoHugging Face10laion /laions_got_talent_enhanced_no_metadataaudio10K<n<100K0 likes1.2k downloads2y agoHugging Face11laion /laions_got_talent_german_bicodecaudio100K<n<1M0 likes977 downloads2y agoHugging Face12zhaochenyang20 /googletime googletime Audio validation set organized as validation/audio/* plus validation/metadata.jsonl. The metadata contains only file_name and transcription; transcriptions include timestamp and speaker markers. audion<1K0 likes822 downloads2mo agoHugging Face13bond005 /sberdevices_golos_10h_crowd Dataset Card for sberdevices_golos_10h_crowd Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.audioautomatic-speech-recognition10K<n<100K7 likes676 downloads4y agoHugging Face14adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-hokkien Taiwan-Tongues-ASR-CE-dataset-hokkien 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.audioautomatic-speech-recognition10K<n<100K7 likes421 downloads9mo agoHugging Face15govnejri /golos_mfa_punctuation Golos MFA Punctuation Расширенная версия датасета Golos — русскоязычного корпуса речи с краудсорс и студийными записями. Датасет дополнен пунктуацией и word-level временными метками (MFA alignment). Опубликовано и поддерживается Jeti Labs. Описание Параметр Значение Язык Русский (ru) Записей 970,597 Аудио ~1,044 часов Частота дискретизации 16,000 Hz Формат WAV, mono, 16-bit Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.audio100K<n<1M5 likes419 downloads5mo agoHugging Face16laion /laions_got_talent_previewaudio1K<n<10K1 likes380 downloads2y agoHugging Face17bond005 /sberdevices_golos_100h_farfield Dataset Card for sberdevices_golos_100h_farfield Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.audioautomatic-speech-recognition10K<n<100K6 likes337 downloads4y agoHugging Face18gokulbnr /QUT-Event-VTR-Dataset Event-Based Visual Teach-and-Repeat via Fast Fourier-Domain Cross-Correlation Welcome to the official QUT-Event-VTR-Dataset dataset repository attached to the paper Event-Based Visual Teach-and-Repeat via Fast Fourier-Domain Cross-Correlation, to be presented at the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). File Structure Cite us at Event-Based Visual Teach-and-Repeat via Fast Fourier-Domain… See the full description on the dataset page: https://huggingface.co/datasets/gokulbnr/QUT-Event-VTR-Dataset.audion<1K0 likes317 downloads3mo agoHugging Face19ylacombe /google-chilean-spanish Dataset Card for Tamil Speech Dataset Summary This dataset consists of 7 hours of transcribed high-quality audio of Chilean Spanish sentences recorded by 31 volunteers. The dataset is intended for speech technologies. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. Supported Tasks text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS). automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-chilean-spanish.audiotext-to-speech1K<n<10K24 likes314 downloads3y agoHugging Face20vnpost-ai /govvox-100h-v3gated GovVox-100h-v3 — 9 tỉnh đặc trưng × 3 miền đều 33.3h thời lượng 95.37 giờ (chia đều ~33.3h/miền) số đoạn 38,744 người nói 1,056 (2,343 id diarization theo phiên) tỉnh 9 (mỗi miền đúng 3 tỉnh đặc trưng nhất) bản ghi nguồn 267 9 tỉnh được chọn — theo kết quả thí nghiệm thực tế Miền Tỉnh Nguồn khả dụng Accuracy mô hình (3 lớp / 5 lớp) Bắc Hà Nội 21.97h 92–100% Bắc Hải Phòng 30.19h 84–86% Bắc Ninh Bình 58.25h 89–92% Trung Hà… See the full description on the dataset page: https://huggingface.co/datasets/vnpost-ai/govvox-100h-v3.audioaudio-classification10K<n<100K1 likes303 downloads8d agoHugging Face21HTH-inc /japanese-casual-conversational-speech-golden-dataset-preview Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.audioautomatic-speech-recognitionn<1K2 likes238 downloads22d agoHugging Face22hadamard-2 /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K0 likes234 downloads10d agoHugging Face23hadamard-2 /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes227 downloads10d agoHugging Face24ylacombe /google-colombian-spanish Dataset Card for "google-colombian-spanish" More Information needed audio1K<n<10K14 likes224 downloads3y agoHugging Face25ittailup /google-la-voices Dataset Card for "google-la-voices" Speaker Durations Speaker Duration (seconds) 00295 1606.144 00610 7026.261 01208 3284.907 01523 6309.888 02121 4687.445 02436 4654.080 02484 9379.925 02485 130.219 03034 5186.048 03349 5143.381 03397 7852.203 03398 118.101 03853 638.037 04310 8260.437 04311 105.472 04766 590.165 05223 8257.773 05679 846.251 0613610207.707 06592 863.659 07049 7580.715 07060 575.659 07505 1743.531… See the full description on the dataset page: https://huggingface.co/datasets/ittailup/google-la-voices.audio1K<n<10K0 likes222 downloads2y agoHugging Face26Hunzla /simplified_google_speech_commands_wav2vec2_960haudio10K<n<100K1 likes218 downloads3y agoHugging Face27q1805 /german-golden-audio_speech-IPA 🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx) An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions. 📊 Dataset Summary Total Samples: 1,354 high-quality audio recordings. Total Size: ~419 MB (Compressed Parquet format). Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.audioautomatic-speech-recognition1K<n<10K0 likes216 downloads27d agoHugging Face28ylacombe /google-tamil Dataset Card for Tamil Speech Dataset Summary This dataset consists of 7 hours of transcribed high-quality audio of Tamil sentences recorded by 50 volunteers. The dataset is intended for speech technologies. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. Supported Tasks text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS). automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-tamil.audiotext-to-speech1K<n<10K7 likes213 downloads3y agoHugging Face29adi-gov-tw /Taiwan-Tongues-ASR-CE-dataset-zhtw Taiwan-Tongues-ASR-CE-dataset-zhtw 本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。 📂 Dataset 結構 本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放: Training set (WebDataset format) train/train-000000.tar train/train-000001.tar ... Test set (WebDataset format) test/test-000000.tar ... tsv set train.tsv test.tsv ... 每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。 🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.audioautomatic-speech-recognition100K<n<1M2 likes205 downloads9mo agoHugging Face30Reza2kn /visualears-golden-6669 🗂️ visualears-golden-6669 English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Held-out VisualEars6669 / Golden6669 evaluation dataset. مجموعهٔ ارزیابی نگه‌داشته‌شدهٔ Golden6669 با شرایط پاک، دورمیدان و مسدود برای سنجش واقع‌گرایانهٔ ASR فارسی. 🧩 Role evaluation and benchmarking asset مصنوع ارزیابی و بنچمارک 📦 Snapshot 13 files; approximately 916.66 MB 13 فایل؛ حدود… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-golden-6669.audioautomatic-speech-recognition1K<n<10K1 likes205 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.