CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ehabnegm /100-hour-Egyptian-dataset-single-speaker Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts. Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data. Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home. This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.audiotext-to-speech10K<n<100K11 likes4.9k downloads2mo agoHugging Face02parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.1k downloads2y agoHugging Face03treble-technologies /Treble10-Speech Treble10-Speech (16 kHz) The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms. The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s. Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.audioautomatic-speech-recognition1K<n<10K21 likes1.1k downloads11mo agoHugging Face04blanchon /parler-tts_mls_eng_10k_snac_token_old Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tabularautomatic-speech-recognition100K<n<1M1 likes990 downloads2y agoHugging Face05bond005 /sberdevices_golos_10h_crowd Dataset Card for sberdevices_golos_10h_crowd Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test subsets.… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd.audioautomatic-speech-recognition10K<n<100K7 likes681 downloads4y agoHugging Face06doof-ferb /vlsp2020_vinai_100h unofficial mirror of VLSP 2020 - VinAI - ASR challenge dataset official announcement: tiếng việt: https://institute.vinbigdata.org/events/vinbigdata-chia-se-100-gio-du-lieu-tieng-noi-cho-cong-dong/ in eglish: https://institute.vinbigdata.org/en/events/vinbigdata-shares-100-hour-data-for-the-community/ VLSP 2020 workshop: https://vlsp.org.vn/vlsp2020 official download: https://drive.google.com/file/d/1vUSxdORDxk-ePUt-bUVDahpoXiqKchMx/view?usp=sharing contact: info@vinbigdata.org… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vlsp2020_vinai_100h.audioautomatic-speech-recognition10K<n<100K19 likes583 downloads1y agoHugging Face07bond005 /sberdevices_golos_100h_farfield Dataset Card for sberdevices_golos_100h_farfield Dataset Summary Sberdevices Golos is a corpus of approximately 1200 hours of 16kHz Russian speech from crowd (reading speech) and farfield (communication with smart devices) domains, prepared by SberDevices Team (Alexander Denisenko, Angelina Kovalenko, Fedor Minkin, and Nikolay Karpov). The data is derived from the crowd-sourcing platform, and has been manually annotated. Authors divide all dataset into train and test… See the full description on the dataset page: https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield.audioautomatic-speech-recognition10K<n<100K6 likes326 downloads4y agoHugging Face08psdn-ai /bangla-10kgated Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh Bangla-10K is a 10,070.8-hour Bengali speech corpus with 567,323 recordings from India and Bangladesh. It combines scripted single-speaker read speech with natural multi-speaker conversations for Bengali automatic speech recognition (ASR). The paper rounds the corpus scale to 10,000 hours. The corpus and its ASR evaluation are described in the anonymous manuscript… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.audioautomatic-speech-recognition100K<n<1M0 likes324 downloads1d agoHugging Face09KasuleTrevor /Lingala_100hrs Lingala 100hrs 110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research. Composition Counts from a full-pass audit on 2026-07-09: Source Upstream location Rows Splits AfriVoice (Lingala) https://huggingface.co/datasets/DigitalUmuganda/AfriVoice 17,544 train (16,144), validation (915), test (485) LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.audioautomatic-speech-recognition10K<n<100K0 likes267 downloads3mo agoHugging Face10nguyenvulebinh /libris_clean_100 Dataset Card for librispeech_asr Dataset Summary LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Supported Tasks and Leaderboards automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/libris_clean_100.audioautomatic-speech-recognition10K<n<100K3 likes237 downloads4y agoHugging Face11badrex /kinyarwanda-speech-1000h Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~1000 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.audioautomatic-speech-recognition100K<n<1M0 likes211 downloads1y agoHugging Face12ErfanAShams /librispeech-alignments_clean100 librispeech-alignments_clean100 This is a subset of librispeech-alignments (https://huggingface.co/datasets/gilkeyio/librispeech-alignments) which only includes train_clean_100 and test_clean splits for small experiments and tutorials. Cite: @inproceedings{panayotov2015librispeech, title={Librispeech: an ASR corpus based on public domain audio books}, author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev}, booktitle={ICASSP}, year={2015}… See the full description on the dataset page: https://huggingface.co/datasets/ErfanAShams/librispeech-alignments_clean100.audioautomatic-speech-recognition10K<n<100K1 likes187 downloads1y agoHugging Face13ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes172 downloads5mo agoHugging Face14parler-tts /mls-eng-10k-tags_tagged_10k_generated Dataset Card for Annotations of 10K hours of English MLS This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.tabularautomatic-speech-recognition1M<n<10M17 likes139 downloads2y agoHugging Face15Appenlimited /1000h-us-english-smartphone-conversation 📚 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings) This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for: Automatic Speech Recognition (ASR) Speaker Identification and Gender/Age Analysis Dialect and Accent Modeling Multi-speaker Speech Separation 🧾 Dataset Contents The dataset includes: metadata.CSV: Metadata including speaker gender, age… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.audioautomatic-speech-recognitionn<1K3 likes137 downloads1y agoHugging Face16twangodev /radiotalk-us-transcripts-qwen3-100k radiotalk-us-transcripts-qwen3-100k 100,000 synthetic US air-traffic-control transcripts, generated with Qwen/Qwen3-32B-NVFP4 (v1 radiotalk pipeline). First release in the radiotalk transcripts series; the v2 release with higher per-transcript realism lives at twangodev/radiotalk-us-transcripts-qwen3-25k. Renamed from radiotalk-us-transcripts-100k on 2026-08-08 to record the generator model in the dataset name; the old id redirects here. textautomatic-speech-recognition10K<n<100K0 likes104 downloads2mo agoHugging Face17SynDataLab-JA-Refs /Irodori-Ja-Spk4-10k SynDataLab/Irodori-Ja-Spk4-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk4): 30s female, news-anchor mature — 30代女性、ニュースキャスター風の落ち着いた声. How this speaker was made The voice identity for Spk4 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk4-10k.audiotext-to-speech10K<n<100K0 likes104 downloads5mo agoHugging Face18tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes97 downloads1mo agoHugging Face19SynDataLab-JA-Refs /Irodori-Ja-Spk3-10k SynDataLab/Irodori-Ja-Spk3-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk3): 40s male, low calm mature — 40代男性、低めで穏やかな落ち着いた声. How this speaker was made The voice identity for Spk3 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from the… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk3-10k.audiotext-to-speech10K<n<100K0 likes95 downloads5mo agoHugging Face20hypaai /Hypa-Speech-10k A multilingual instruction-tuning dataset covering translation,transcription, and language detection. Dataset Card for Hypa-Speech-10k Dataset Summary Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets. The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Speech-10k.audioautomatic-speech-recognition10K<n<100K0 likes88 downloads3mo agoHugging Face21Vikhrmodels /Ficbook-Audio-Instruct-10K Ficbook Audio Instruct 10K Synthetic audio instruction dataset for training Russian audio-language models. Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks. Dataset Description This dataset was created for training and evaluating audio-language models on Russian fiction content. Each sample contains: Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model Text: Original text from ficbook stories Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.audioautomatic-speech-recognition1K<n<10K0 likes82 downloads9mo agoHugging Face22SynDataLab-JA-Refs /Irodori-Ja-Spk1-10k SynDataLab/Irodori-Ja-Spk1-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk1): 30s male, calm conversational — 30代男性、落ち着いた自然な会話調. How this speaker was made The voice identity for Spk1 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized from… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk1-10k.audiotext-to-speech10K<n<100K0 likes66 downloads5mo agoHugging Face23linq1005 /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/linq1005/fleurs.audioautomatic-speech-recognition100K<n<1M0 likes65 downloads2mo agoHugging Face24pharaouk /mls-eng-10k-tags_tagged_10k_generated Dataset Card for Annotations of 10K hours of English MLS This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/mls-eng-10k-tags_tagged_10k_generated.tabularautomatic-speech-recognition1M<n<10M0 likes61 downloads2y agoHugging Face25Kppwdfgu1 /Hypa-Speech-10k A multilingual instruction-tuning dataset covering translation,transcription, and language detection. Dataset Card for Hypa-Speech-10k Dataset Summary Hypa-Speech-10k is a curated, multilingual speech dataset consisting of 10,000 audio-text pairs spanning 18 languages, including several low-resource African languages that are under-represented in mainstream speech datasets. The source text and base audio for this dataset were drawn from the Mozilla Common… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/Hypa-Speech-10k.audioautomatic-speech-recognition10K<n<100K0 likes55 downloads3mo agoHugging Face26doof-ferb /vais1000 unofficial mirror of VAIS-1000 official announcement: https://vais.vn/vi/tai-ve/hts_for_vietnamese (dead) mirror: https://github.com/undertheseanlp/text_to_speech/tree/run/data/vais1000/raw small only 1h40min audio - 1 speaker (female northern accent) - 1k samples pre-process: none need to do: check misspelling, restore foreign words phonetised to vietnamese usage with HuggingFace: # pip install -q "datasets[audio]" from datasets import load_dataset from torch.utils.data import… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/vais1000.audioautomatic-speech-recognition1K<n<10K0 likes54 downloads2y agoHugging Face27SynDataLab-JA-Refs /Irodori-Ja-Spk2-10k SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker family — see also Spk1, Spk2, Spk3, Spk4 repos. Speaker (Spk2): 30s female, narrator-style natural — 30代女性、ナレーター風の自然な声. How this speaker was made The voice identity for Spk2 was created in two stages: Stage 1 — voice anchor design. Using Irodori-TTS-500M-v2-VoiceDesign, several candidate audio samples were synthesized… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA-Refs/Irodori-Ja-Spk2-10k.audiotext-to-speech10K<n<100K0 likes48 downloads5mo agoHugging Face28Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes44 downloads2y agoHugging Face29tannhoo06 /ViMD_chunked_10s ViMD Chunked 10s — 16kHz Preprocessed from ViMD (Nguyen et al., EMNLP 2024). Preprocessing Resample: 44.1kHz -> 16kHz mono Chunking: each audio is split into consecutive NON-OVERLAPPING segments of at most 10 seconds. ALL segments are kept, including the final remainder (no minimum length filter). 1 original file -> ceil(len/10s) samples. Splits: original ViMD train/valid/test kept (speaker-exclusive). Segments of the same file always stay in the same split.… See the full description on the dataset page: https://huggingface.co/datasets/tannhoo06/ViMD_chunked_10s.audioaudio-classification10K<n<100K0 likes41 downloads3mo agoHugging Face30Ethan615 /taiwan-conversation-context-100-domainsgated Taiwan Conversation Context 100 Domains Dataset Description Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。 本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。 資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於: 語音生成資料前處理 Text-to-Speech, TTS Spoken Dialogue Generation Conversational AI Customer Service Dialogue Modeling Role-play Dialogue Dataset 台灣繁體中文語音模型訓練 生活情境問答模型訓練 對話式 AI 助理訓練 RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.texttext-generation1M<n<10M2 likes37 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.