CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bretagne /Banque_Sonore_Dialectes_Bretons [!NOTE] Dataset origin: http://banque.sonore.breton.free.fr/ Description Issue du site Banque Sonore des Dialectes Bretons Présentation du projet La Banque Sonore des Dialectes Bretons est un projet expérimental qui réunit sur internet un vaste ensemble d'enregistrements d'enquêtes effectuées depuis plus d'une dizaine d'années auprès de locuteurs traditionnels de breton. Alimentées par une équipe de bénévoles partageant un intérêt commun pour… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/Banque_Sonore_Dialectes_Bretons.audioautomatic-speech-recognition1K<n<10K3 likes7.5k downloads2mo agoHugging Face02oddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M29 likes4.5k downloads2mo agoHugging Face03grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.6k downloads8mo agoHugging Face04moaead /dialectal-arabic-voices Dialectal Arabic Voices An expanding collection of Arabic audio from YouTube, SoundCloud, and other sources. Currently labelled Palestinian Arabic (ps). 60,775 recordings · approximately 10,895.2 hours · 571.24 GB Column Description audio Original audio, embedded in the Parquet file transcript_text Empty; ASR transcripts are stored in a separate private dataset language Dialect code: ps (Palestinian) source Original channel or account name Audio retains its… See the full description on the dataset page: https://huggingface.co/datasets/moaead/dialectal-arabic-voices.audioautomatic-speech-recognition10K<n<100K0 likes1.9k downloads2h agoHugging Face05ylacombe /english_dialects Dataset Card for "english_dialects" Dataset Summary This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English. The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.audiotext-to-speech10K<n<100K37 likes1.9k downloads3y agoHugging Face06Suchae /Korea-AIHub-middlesenior-dialect-speech-train-part2audio100K<n<1M0 likes1k downloads2y agoHugging Face07malaysia-ai /pseudolabel-dialects-youtube-whisper-large-v3 malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3 How to prepare the dataset huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \ malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.audio1M<n<10M0 likes880 downloads1y agoHugging Face08rbcurzon /ph_dialect_asraudio10K<n<100K1 likes509 downloads1y agoHugging Face09typhoon-ai /thai-dialect-isan-dataset Dataset Card for Thai Dialect Isan Speech Corpus Dataset Description This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language. The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.textautomatic-speech-recognition10K<n<100K5 likes470 downloads10mo agoHugging Face10leyu-amharic /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K1 likes436 downloads2mo agoHugging Face11amgadhasan /arabic_tweets_dialectstexttext-classification100K<n<1M0 likes404 downloads2y agoHugging Face12islomov /podcasts_tashkent_dialect_youtube_uzbek_speech_dataset Tashkent dialect focused podcasts youtube uzbek speech Dataset Description This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models. Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.audioautomatic-speech-recognition10K<n<100K6 likes377 downloads1y agoHugging Face13FraPiz /moldovan-dialectal-romanian-speech-corpus Moldovan Dialectal Romanian Educational Speech Corpus This dataset contains aligned Romanian educational speech with Moldovan dialectal characteristics. It was constructed from publicly accessible lesson videos recorded by teachers from the Republic of Moldova and published through the EducatieOnline platform. The corpus supports research on automatic speech recognition (ASR), text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.audioautomatic-speech-recognition10K<n<100K1 likes375 downloads1mo agoHugging Face14leyu-amharic /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes362 downloads2mo agoHugging Face15n-order /Thai-dialect-corpusaudio100K<n<1M1 likes341 downloads2y agoHugging Face16Suchae /Korea-AIHub-middlesenior-dialect-speech-validation-part2audio10K<n<100K0 likes324 downloads2y agoHugging Face17Vinidapooh /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M0 likes320 downloads14d agoHugging Face18rjnieto /french-dialectsaudio1K<n<10K0 likes299 downloads2y agoHugging Face19leyu-amharic /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes244 downloads2mo agoHugging Face20QCRI /arabic_pos_dialect Dataset Card for Arabic POS Dialect Dataset Summary This dataset was created to support part of speech (POS) tagging in dialects of Arabic. It contains sets of 350 manually segmented and POS tagged tweets for each of four dialects: Egyptian, Levantine, Gulf, and Maghrebi. Supported Tasks and Leaderboards The dataset can be used to train a model for Arabic token segmentation and part of speech tagging in Arabic dialects. Success on this task is typically… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/arabic_pos_dialect.texttoken-classification1K<n<10K12 likes234 downloads3y agoHugging Face21hadamard-2 /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K0 likes224 downloads13d agoHugging Face22hadamard-2 /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes213 downloads13d agoHugging Face23hadamard-2 /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Wello dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K0 likes197 downloads13d agoHugging Face24leyu-amharic /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K1 likes194 downloads2mo agoHugging Face25gheero-Leyu /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes193 downloads2mo agoHugging Face26gheero-Leyu /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes182 downloads2mo agoHugging Face27Abdelrahman-Rezk /Arabic_Dialect_IdentificationArabic dialects, multi-class-Classification, Tweets. Dataset Card for Arabic_Dialect_Identification Dataset Summary We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset relies on applying multiple filters to identify users who belong to different countries based on their account descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman-Rezk/Arabic_Dialect_Identification.tabular100K<n<1M12 likes173 downloads4y agoHugging Face28MBZUAI /Dialectal-Arabic-MMLU DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models Dataset Summary Dialectal-Arabic-MMLU is a large-scale, human-translated for MMLU. We extend MMLU-Redux into 5 major dialects: Syrian, Egyptian, Emirati, Saudi, and Moroccan. This data covers 21K QA pairs across 32 academic and professional domains. More details, please check our paper on DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Dialectal-Arabic-MMLU.tabularmultiple-choice10K<n<100K1 likes173 downloads4mo agoHugging Face29hadamard-2 /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes169 downloads13d agoHugging Face30TingChen-ppmc /Shanghai_Dialect_Conversational_Speech_Corpus Corpus This dataset is built from Magicdata ASR-CZDIACSC: A CHINESE SHANGHAI DIALECT CONVERSATIONAL SPEECH CORPUS This corpus is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please refer to the license for further information. Modifications: The audio is split in sentences based on the time span on the transcription file. Sentences that span less than 1 second is discarded. Topics of conversation is removed. Usage… See the full description on the dataset page: https://huggingface.co/datasets/TingChen-ppmc/Shanghai_Dialect_Conversational_Speech_Corpus.audio1K<n<10K12 likes158 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.