datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Banque_Sonore_Dialectes_Bretons
[!NOTE]
Dataset origin: http://banque.sonore.breton.free.fr/
Description
Issue du site Banque Sonore des Dialectes Bretons
Présentation du projet
La Banque Sonore des Dialectes Bretons est un projet expérimental qui réunit sur internet un vaste ensemble d'enregistrements d'enquêtes effectuées depuis plus d'une dizaine d'années auprès de locuteurs traditionnels de breton.
Alimentées par une équipe de bénévoles partageant un intérêt commun pour… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/Banque_Sonore_Dialectes_Bretons.dialectal-arabic-lahgtna-v2
Dialectal Arabic Lahgtna v2
Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI.
Dataset Summary
~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech
**13 Arabic dialects **, labeled per utterance
16 kHz mono audio
Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
english_dialects
Dataset Card for "english_dialects"
Dataset Summary
This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English.
The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.Korea-AIHub-middlesenior-dialect-speech-train-part2pseudolabel-dialects-youtube-whisper-large-v3
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.ph_dialect_asrleyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.podcasts_tashkent_dialect_youtube_uzbek_speech_dataset
Tashkent dialect focused podcasts youtube uzbek speech
Dataset Description
This dataset contains audio clips and their corresponding transcriptions in the Uzbek language with mostly tashkent dialects. The data was collected from publicly available podcast videos on YouTube. It is designed for training and evaluating Automatic Speech Recognition (ASR) models.
Most of the content comes from the Jahongir Latipov interviews and Bu podcast (respect authors) YouTube videos. The… See the full description on the dataset page: https://huggingface.co/datasets/islomov/podcasts_tashkent_dialect_youtube_uzbek_speech_dataset.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.Thai-dialect-corpusmoldovan-dialectal-romanian-speech-corpus
Moldovan Dialectal Romanian Educational Speech Corpus
This dataset contains aligned Romanian educational speech with Moldovan
dialectal characteristics. It was constructed from publicly accessible lesson
videos recorded by teachers from the Republic of Moldova and published through
the EducatieOnline platform.
The corpus supports research on automatic speech recognition (ASR),
text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal
speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
french-dialectsKorea-AIHub-middlesenior-dialect-speech-validation-part2magicdata-dialect-sichuanese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/magicdata-dialect-sichuanese-tts-lite.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-shewa-dialect.magicdata-dialect-sichuanese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/magicdata-dialect-sichuanese-tts-lite.magicdata-dialect-wu-chinese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/magicdata-dialect-wu-chinese-tts-lite.leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Wello dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-wello-dialect.magicdata-dialect-cantonese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/magicdata-dialect-cantonese-tts-lite.magicdata-dialect-northeastern-chinese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/magicdata-dialect-northeastern-chinese-tts-lite.Egyptian_Dialect
Egyptian Arabic Speech Dataset
Dataset Description
This dataset contains 2,438 short audio segments in Egyptian Arabic paired with transcriptions.
The dataset was created for fine-tuning Automatic Speech Recognition (ASR) models on conversational Egyptian Arabic.
Each example contains:
WAV audio
Egyptian Arabic transcription
Dataset Creation
Source
The audio was collected from publicly available YouTube videos featuring native… See the full description on the dataset page: https://huggingface.co/datasets/Kyrillos2001/Egyptian_Dialect.magicdata-dialect-henan-dialect-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/magicdata-dialect-henan-dialect-tts-lite.leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gonder-dialect.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.magicdata-dialect-northeastern-chinese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/magicdata-dialect-northeastern-chinese-tts-lite.
