CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bretagne /Banque_Sonore_Dialectes_Bretons [!NOTE] Dataset origin: http://banque.sonore.breton.free.fr/ Description Issue du site Banque Sonore des Dialectes Bretons Présentation du projet La Banque Sonore des Dialectes Bretons est un projet expérimental qui réunit sur internet un vaste ensemble d'enregistrements d'enquêtes effectuées depuis plus d'une dizaine d'années auprès de locuteurs traditionnels de breton. Alimentées par une équipe de bénévoles partageant un intérêt commun pour… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/Banque_Sonore_Dialectes_Bretons.audioautomatic-speech-recognition1K<n<10K3 likes7.6k downloads2mo agoHugging Face02oddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M28 likes3.7k downloads2mo agoHugging Face03grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.2k downloads7mo agoHugging Face04typhoon-ai /thai-dialect-isan-dataset Dataset Card for Thai Dialect Isan Speech Corpus Dataset Description This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language. The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.textautomatic-speech-recognition10K<n<100K5 likes530 downloads10mo agoHugging Face05leyu-amharic /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K1 likes436 downloads2mo agoHugging Face06islomov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K6 likes389 downloads1y agoHugging Face07leyu-amharic /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes368 downloads2mo agoHugging Face08FraPiz /moldovan-dialectal-romanian-speech-corpus Moldovan Dialectal Romanian Educational Speech Corpus This dataset contains aligned Romanian educational speech with Moldovan dialectal characteristics. It was constructed from publicly accessible lesson videos recorded by teachers from the Republic of Moldova and published through the EducatieOnline platform. The corpus supports research on automatic speech recognition (ASR), text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.audioautomatic-speech-recognition10K<n<100K1 likes316 downloads1mo agoHugging Face09Vinidapooh /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M0 likes308 downloads10d agoHugging Face10leyu-amharic /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes251 downloads2mo agoHugging Face11hadamard-2 /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K0 likes234 downloads10d agoHugging Face12hadamard-2 /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes227 downloads10d agoHugging Face13hadamard-2 /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Wello dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K0 likes218 downloads10d agoHugging Face14Kyrillos2001 /Egyptian_Dialect Egyptian Arabic Speech Dataset Dataset Description This dataset contains 2,438 short audio segments in Egyptian Arabic paired with transcriptions. The dataset was created for fine-tuning Automatic Speech Recognition (ASR) models on conversational Egyptian Arabic. Each example contains: WAV audio Egyptian Arabic transcription Dataset Creation Source The audio was collected from publicly available YouTube videos featuring native… See the full description on the dataset page: https://huggingface.co/datasets/Kyrillos2001/Egyptian_Dialect.audioautomatic-speech-recognition1K<n<10K1 likes199 downloads3mo agoHugging Face15leyu-amharic /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K1 likes191 downloads2mo agoHugging Face16hadamard-2 /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes191 downloads10d agoHugging Face17gheero-Leyu /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes188 downloads2mo agoHugging Face18gheero-Leyu /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes181 downloads2mo agoHugging Face19AhmedEladl /saudi-dialect-speech-female 🌍 Saudi Dialectal Arabic Audio Dataset This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning. 🗂️ Dataset Columns Column Description audio The audio chunk (22,050 Hz, mono WAV) duration Chunk duration in seconds base_transcription Transcript from the base Arabic ASR model dialectal_transcription Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.audioautomatic-speech-recognition1K<n<10K1 likes167 downloads1mo agoHugging Face20hadamard-2 /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Shewa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes164 downloads10d agoHugging Face21leyu-amharic /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K1 likes150 downloads2mo agoHugging Face22gheero-Leyu /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K1 likes133 downloads2mo agoHugging Face23gheero-Leyu /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes131 downloads2mo agoHugging Face24wannaphong /thai-dialect-isan-dataset Dataset Card for Thai Dialect Isan Speech Corpus Dataset Description This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language. The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-dialect-isan-dataset.textautomatic-speech-recognition10K<n<100K0 likes122 downloads6mo agoHugging Face25AhmedEladl /emirates-dialect-speech-male 🌍 Emirates Dialectal Arabic Audio Dataset This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning. 📌 Source Data & Provenance Source Repository: https://github.com/MahaAlBlooki/alsanaa-emirati-dataset Domain & Content: Spoken Emirati dialectal Arabic speech recordings. Dialect Focus: Emirates / Gulf Dialectal Arabic. Standardized Format: 22,050 Hz… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/emirates-dialect-speech-male.audioautomatic-speech-recognition1K<n<10K0 likes113 downloads1mo agoHugging Face26KSE-RESEARCH-Group /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K1 likes89 downloads7mo agoHugging Face27vladsfa /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples… See the full description on the dataset page: https://huggingface.co/datasets/vladsfa/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K0 likes75 downloads12d agoHugging Face28KlangAI /klang-dialects Klang Dialects Klang Dialects is an open benchmark for Swedish speech recognition, created by the Klang Research Team from recordings contributed through Knäck Klang. The Swedish benchmark contains 1,804 recordings, 656 speaker IDs, and 5.15 hours of speech in its main configuration, sv-clean. The benchmark supports research on regional variation in Swedish speech recognition. We plan to expand the dataset with more sentences, speakers, and languages. Configurations… See the full description on the dataset page: https://huggingface.co/datasets/KlangAI/klang-dialects.audioautomatic-speech-recognition1K<n<10K2 likes69 downloads6d agoHugging Face29gheero-Leyu /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K1 likes59 downloads2mo agoHugging Face30niloycste68 /Bangali_local_dialect_ASR_HF_Dataset BanglaMix — Code-Switching ASR in Bangladeshi Regional Dialects BanglaMix is a speech dataset for Automatic Speech Recognition (ASR) on dialectal Bangladeshi Bengali mixed with English (code-switching). It covers 15 regional dialects and the natural Bengali–English code-switching common in informal Bangladeshi speech — a setting not covered by existing Bengali corpora, which address either dialects or code-switching, never both. Clips 41,499 transcribed audio clips… See the full description on the dataset page: https://huggingface.co/datasets/niloycste68/Bangali_local_dialect_ASR_HF_Dataset.audioautomatic-speech-recognition10K<n<100K0 likes46 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.